Site Reliability Engineer (SRE)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+17 more
Job description
JIRA & Confluence aiops monitoring tools observability aws cloud Site Reliability Engineering (SRE) DevOps & CI/CD MTTR Google Cloud Platform (GCP) dynatrace, * Lead the design and implementation of full-stack observability solutions with Dynatrace as the primary platform.
- Configure Dynatrace for application performance monitoring (APM), infrastructure monitoring, and intelligent alerting.
- Build advanced dashboards and integrate Dynatrace with event management systems to enable proactive incident prevention and root cause analysis.
- Collaborate with teams to optimize Dynatrace usage for AIOps-driven insights and automated anomaly detection.
- Provide oversight for production operations to maximize reliability and automation.
- Develop and evolve SRE best practices, runbooks, and tooling to ensure high availability and resilience.
- Implement data-driven operational strategies to improve decision-making and reduce MTTR.
- Hands-on experience with Dynatrace, Splunk, ELK, Grafana, Prometheus, and (future) ThousandEyes.
- Build and manage CI/CD pipelines and Infrastructure as Code (IaC) solutions using Terraform, Jenkins, TeamCity, Octopus, Bamboo, and U-Deploy across hybrid/multi-cloud environments.
- Develop and manage DevOps pipelines in AWS, Azure, and GCP using Terraform and cloud-native tooling.
- Strong developer background with the ability to understand application layers and infrastructure interactions.
- Define and document standard operating procedures, architecture diagrams, and system documentation using Jira, Confluence, and UML.
- Identify areas for process and efficiency improvement within Platform Services Operations; recommend and implement solutions.
- Drive automation initiatives across all operational processes.
- Proactively monitor system capacity and health indicators; provide analytics and forecasts for scaling.
Requirements
We are looking for a Site Reliability Engineer with deep expertise in Dynatrace and a strong background in observability, automation, and cloud operations. This role focuses on designing and implementing highly reliable, scalable solutions while driving proactive monitoring and operational excellence., * Expert-level experience with Dynatrace, including dashboard creation, alert configuration, and integration with other observability tools.
- Strong knowledge of AIOps, performance tuning, and proactive incident management.
- Familiarity with hybrid/multi-cloud environments and modern DevOps practices.
- Excellent problem-solving skills and ability to work in a fast-paced, collaborative environment.
Benefits & conditions
The pay range that the employer in good faith reasonably expects to pay for this position is $36.98/hour - $57.79/hour. Our benefits include medical, dental, vision and retirement benefits. Applications will be accepted on an ongoing basis.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Find a Developer Job: 12 Best Job Sites For Developers
Where To Find Software Engineering Jobs
Is Software Engineering Over-Saturated?
Highest Paying Tech Companies for Developers