> Markdown version of [/jobs/ext/604304-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/604304-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** DigiCert, Inc. - **Location:** Lehi, UT, United States - **Experience:** Expert - **Salary:** $125,000.0 - $145,000.0 - **Contract:** Permanent contract - **Skills:** Testing (Software), Amazon Web Services, Microsoft Azure, Bash Shell, Cloud Computing, Cloud Engineering, Code Review, Continuous Integration, Software Debugging, DevOps, Distributed Systems, Fault Tolerance, Python (Programming Language), Software Architecture, Reliability Engineering, Prometheus, Software Engineering, Datadog, Data Logging, Google Cloud, Real Time Systems, Delivery Pipeline, Grafana, Reliability of Systems, Infrastructure as Code (IaC), Build Management, Kubernetes, Deployment Automation, Software Coding, Terraform, Splunk, New Relic (SaaS), Dynatrace, Golang - **Published:** June 24, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=1b9e4dd4d067b785 ## About the Role Do you have experience in Tooling?, * Extensive experience in distributed systems, cloud-native architectures (AWS, GCP, Azure), and DevOps practices. * Proficiency in Kubernetes, Terraform, CI/CD pipelines, and Infrastructure as Code (IaC). * Strong scripting and automation skills in Python, Go, Bash, or similar languages. * Expertise in observability tools such as Prometheus, Grafana, Datadog, Splunk, New Relic, and OpenTelemetry. * Ability to troubleshoot complex production issues and drive scalable, resilient solutions. * Experience reviewing code, debugging applications, and conducting software testing to ensure high reliability and quality. ## Description The Site Reliability Engineer (SRE) collaborates with development teams to embed reliability, scalability, and performance best practices throughout the software development lifecycle. This role bridges software engineering and cloud operations, ensuring mission-critical systems remain highly available and resilient. By integrating reliability early, the SRE fosters a culture of shared responsibility while enabling rapid and safe feature delivery. What you will do * Design and build fault-tolerant, high-performing systems that meet Service Level Objectives (SLOs) and Service Level Agreements (SLAs). * Implement monitoring, alerting, distributed tracing, and logging to ensure real-time system health visibility and proactive issue resolution. * Act as a first responder for production incidents, conduct blameless postmortems, and drive root cause analysis (RCA) and corrective actions. * Develop self-healing, automated deployments, and scaling solutions to minimize toil and improve system efficiency. * Improve continuous integration and deployment pipelines to enable safe, rapid, and reliable feature rollouts. * Review code, debug issues, and perform quality assurance (QA) on software components to enhance system reliability and performance. * Work closely with development teams to ensure best practices in software architecture, coding standards, and operational readiness. * Forecast scalability needs and optimize cloud infrastructure costs while balancing performance and efficiency. * Ensure production environments meet security and compliance requirements, collaborating with teams to mitigate vulnerabilities and enforce best practices. * Work closely with development teams to embed reliability at every stage rather than treating it as an afterthought. * Use error budgets to balance feature velocity with system stability. * Implement observability and automation-first principles to measure system health and drive continuous improvement. * Leverage game days, chaos engineering, and resilience testing to validate system robustness and refine operational processes. ## Related Videos - [Answering the Million Dollar Question: Why did I Break Production?](https://www.wearedevelopers.com/videos/1171-answering-the-million-dollar-question-why-did-i-break-production) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [The 8 Best Code Testing Tools](https://www.wearedevelopers.com/magazine/402-the-8-best-code-testing-tools) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)