> Markdown version of [/jobs/ext/2283937-principal-site-reliability-engineer-remote](https://www.wearedevelopers.com/jobs/ext/2283937-principal-site-reliability-engineer-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Site Reliability Engineer - Remote - **Company:** Unitedhealth Group Inc - **Location:** Eden Prairie, MN, United States (Remote available) - **Salary:** $134,600.0 - $230,800.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, DevOps, Disaster Recovery, Reliability Engineering, Ansible, Prometheus, Runbook, Software Engineering, Datadog, Pulumi, Large Language Models, Grafana, Kubernetes, Information Technology, Terraform - **Published:** August 28, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3367789217&tx=KR7171FFJ&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * 10+ years of experience in software engineering, platform engineering, DevOps, or SRE roles * 3+ years of experience in a principal, staff, lead, or senior technical leadership role * 5+ years of experience with cloud platforms and container orchestration, preferably Azure or AWS * 3+ years of experience with observability tools such as OpenTelemetry, Prometheus, Grafana, Datadog, or similar platforms * 1+ years of experience designing production automation, tooling, or AI-assisted workflows for incident response or operational decision-making, * Bachelor's degree in Computer Science, Information Technology, Engineering, or related field * Experience with LLM-based systems, AI agents, RAG, tool orchestration, evaluations, or guardrails * Experience with resiliency engineering, disaster recovery, chaos engineering, or recovery validation * Experience with infrastructure as code and automation tools such as Terraform, Pulumi, Ansible, Helm, or Kubernetes operators * Solid background in incident command, runbooks, postmortems, production readiness, and reliability governance *All employees working remotely will be required to adhere to UnitedHealth Group's Telecommuter Policy ## Description * Build AI-assisted SRE capabilities that accelerate incident detection, triage, mitigation, and recovery * Connect observability, deployment, runbook, ownership, and incident data into actionable operational context * Design human-in-the-loop workflows for safe mitigation, approvals, recovery verification, and auditability * Standardize OpenTelemetry, SLIs, SLOs, error budgets, and reliability scorecards across critical services * Improve alert quality by reducing noise, clarifying customer impact, and identifying likely causes faster * Lead resiliency testing, DR exercises, chaos engineering, and automated recovery validation * Mentor engineers and lead cross-functional reliability improvements across Optum Financial You'll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in. ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Best Job Boards for Remote Work for Developers](https://www.wearedevelopers.com/magazine/290-best-job-boards-for-remote-work-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Best Paying Remote Jobs](https://www.wearedevelopers.com/magazine/255-best-paying-remote-jobs)