> Markdown version of [/jobs/ext/3023314-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3023314-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Colorado Public Employees' Retirement Association - **Location:** Lakewood, CO, United States (Remote available) - **Experience:** Expert - **Salary:** $140,000.0 - $165,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Build Automation, Automation of Tests, Microsoft Azure, Cloud Computing, Configuration Management, Continuous Integration, Linux, DevOps, Distributed Systems, Fault Tolerance, IT Management, Python (Programming Language), Windows PowerShell, Reliability Engineering, Shell Script, Software Engineering, Data Logging, Cloud Platform System, Mttr, Kubernetes, Information Technology, Hashicorp, Hardware Infrastructure, Terraform, Devsecops, Microservices - **Published:** September 21, 2026 - **Apply:** https://www.careerjet.com/job/us587f937f46b5d02fd43772668a92a29e/eaa ## About the Role The ideal candidate is a senior, hands-on reliability engineer who thinks in systems, automates and treats every incident as a chance to remove future value leakage. They are fluent across AWS and Azure cloud, Kubernetes, and IaC from build to configuration management utilizing AI tooling to speed MTTR and improve uptime by consolidating metrics, logs, and distributed trace collection into observability tools with well-defined dashboards and actionable integrations into ITSM Service Now management in an environment primarily on premises and rapidly moving to cloud platform and infrastructure services., * 6+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a closely related discipline. * Bachelor's degree in Computer Science, Information Technology, or related field preferred, or an equivalent combination of education and experience. * Relevant certifications preferred: AWS Certified DevOps Engineer or Solutions Architect, Certified Kubernetes Administrator (CKA), HashiCorp Terraform Associate.(MP4.1)(DS4.2) * Production experience with AWS and Azure across compute, networking, and managed services. * Hands-on experience operating and troubleshooting Kubernetes in a hybrid (on-premises and cloud) environment. * Strong Python, Powershell and Linux (MP5.1)(DS5.2)(BP5.3)(BP5.4)shell scripting/automation skills; ability to build tooling, not just run it. * Experience designing and operating an enterprise observability/monitoring platforms. * Demonstrated experience reducing MTTR/MTTD through tooling, automation, and process not headcount. * Experience with CI/CD tooling and modern release practices. * Working knowledge of microservices architecture and the reliability challenges specific to distributed systems. * Experience defining and reporting on SLOs, error budgets, and uptime commitments (99.9%+ environments). * Experience participating in an on-call rotation and leading through live production incidents. Preferred Leadership & Strategic Experience * Experience leading a vendor evaluation and selection process for enterprise tooling, including business case development. * Experience establishing reliability standards or practices adopted across multiple engineering teams without direct reporting authority. * Experience mentoring engineers on reliability, automation, or observability practices. * Experience presenting reliability and cost metrics to IT leadership or executive audiences. Working Conditions The physical demands described here are representative of those that must be met by an employee to successfully perform the essential functions of this job. Reasonable accommodation may be made to enable individuals with disabilities to perform the essential functions. * Standard office environment with frequent computer operation and use of collaboration/communication tools. * Participation in an on-call rotation, including evenings, weekends, and holidays as needed to support 99.99% uptime commitments. * Ability to remain calm, clear, and decisive under pressure during live production incidents. * Ability to sit for prolonged periods of time and operate standard PC equipment. * Ability to handle stress associated with production incidents, tight deadlines, and competing priorities. ## Description Employees are held accountable for all duties of the job. Individuals must be able to perform these duties with or without reasonable accommodation. * Observability Platform & Automation Tooling + Lead the evaluation, selection, and enterprise rollout of PERA's observability platform, including licensing, architecture, and integration strategy. + Design and maintain monitoring, logging, and alerting standards across AWS, Azure and on-premises infrastructure, Kubernetes workloads, APIs, services and microservices.(DS1.1) + Tune alerting to reduce noise and false positives while improving detection of real degradation. + Build automation with AI-driven integrations to eliminate manual, repetitive operational work, including self-healing and auto-remediation driving improvement of SLO/KPI metrics where appropriate.(DS2.1) + Develop dashboards and reporting that translate technical telemetry into business-relevant reliability and cost metrics for IT leadership. * Incident Management & Reliability Engineering + Serve as a senior technical responder and escalation point for Priority 1/Priority 2 production incidents, driving time-to-resolution down through structured incident command practices. + Own PERA's blameless postmortem process: facilitate root-cause reviews, document findings, and track corrective actions to closure. + Define and maintain Service Level Objectives (SLOs) and error budgets for critical production services. + Build and maintain runbooks, escalation paths, and on-call procedures that reduce reliance on tribal knowledge. + Analyze incident trends to identify systemic reliability risks and prioritize remediation work. * Architecture & Scaling + Partner with Infrastructure, Application Development and Investment Technical Services to design for scalability, resiliency, and fault tolerance across hybrid Kubernetes cluster (MP3.1)(DS3.2)orchestration. + Conduct capacity planning and load/performance analysis to proactively identify scaling risks before they affect uptime. + Contribute reliability and observability requirements into architecture reviews for new services and major changes. * CI/CD & Release Engineering + Assess and scale PERA's CI/CD pipelines to support faster, safer, and more frequent deployments. + Implement methodologies to safeguard modern hybrid DevSecOps, and Kubernetes environments through security controls in CI/CD pipelines. + Integrate observability and automated testing gates into the CI/CD pipeline so reliability issues are caught before production. * Cost Optimization & Vendor/Tooling Governance + Own the business case, budget, and ongoing vendor relationship for the selected observability platform. + Identify and implement cloud cost optimization opportunities (rightsizing, autoscaling, reserved capacity) surfaced through observability data. + Evaluate emerging SRE/observability tooling and recommend investments aligned with PERA's hybrid technology roadmap. ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)