> Markdown version of [/jobs/ext/1826531-senior-principal-infrastructure-operations-eng-remote-or-hybrid](https://www.wearedevelopers.com/jobs/ext/1826531-senior-principal-infrastructure-operations-eng-remote-or-hybrid). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Principal Infrastructure & Operations Eng - Remote or Hybrid - **Company:** Optum, Inc - **Location:** Minnetonka, MN, United States (Remote available) - **Experience:** Expert - **Salary:** $134,600.0 - $230,800.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Application Portfolio Management, Microsoft Azure, Software as a Service, Cloud Computing, Decision Support Systems, Disaster Recovery, Distributed Systems, Middleware, Intrusion Detection and Prevention, Knowledge Management, Mainframes, Reliability Engineering, Software Engineering, Strategies of Testing, Datadog, Enterprise Software Applications, Cloud Platform System, Large Language Models, Grafana, Mttr, Reliability of Systems, Generative AI, Containerization, ArcSight Event Correlation, Splunk, New Relic (SaaS), Dynatrace, Service Stack, Legacy Systems - **Published:** July 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=dad436fac26e6cea ## About the Role * Undergraduate degree in applicable area of expertise or equivalent experience * 10+ years of experience in technology operations, software engineering, infrastructure engineering, platform engineering, or Site Reliability Engineering * 5+ years leading enterprise-scale SRE, reliability engineering, or production engineering organizations * Demonstrated experience owning reliability outcomes for portfolios exceeding 200+ applications, preferably 500+ * Proven success building and leading high-performing engineering teams * Experience managing direct reports, contractors, managed services providers, and vendor relationships * Deep understanding of modern SRE principles including: + SLOs and SLIs + Error budgets + Reliability engineering + Incident management + Resilience engineering + Capacity management + Observability * Experience supporting diverse technology ecosystems spanning legacy platforms, mainframe, distributed systems, and cloud environments * Proven solid executive communication and stakeholder management skills Preferred Qualifications: * Demonstrated implementation of AI-driven operations, AIOps, or autonomous operations capabilities at enterprise scale * Experience leveraging Generative AI, LLMs, operational copilots, agentic workflows, or predictive analytics to improve operational outcomes * Experience leading large-scale operational transformations with measurable business results. * Demonstrated background in software engineering or platform engineering * Experience with cloud platforms such as AWS, Azure, or GCP * Demonstrated familiarity with observability platforms such as Datadog, Dynatrace, New Relic, Splunk, Grafana, Open Telemetry, or similar technologies * Experience in highly regulated industries such as healthcare, financial services, insurance, or government * All employees working remotely will be required to adhere to UnitedHealth Group's Telecommuter Policy. ## Description This Senior Principal role is accountable for advancing enterprise reliability across a complex, high-scale application portfolio by setting the technical direction, operating model, and leadership approach needed to improve stability, resilience, and operational performance. You'll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week. Primary Responsibilities: Enterprise Reliability Leadership * Establish and execute a comprehensive reliability strategy across a portfolio of 510+ applications supporting critical business operations * Define and govern enterprise reliability standards, Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, resiliency requirements, and operational maturity models * Create a reliability operating model that spans modern cloud-native platforms, legacy systems, mainframe workloads, third-party hosted solutions, and SaaS applications * Serve as the executive leader accountable for enterprise application reliability, stability, recoverability, and operational risk reduction AI-First SRE Transformation * Design and implement an AI-first approach to reliability engineering leveraging generative AI, AIOps, predictive analytics, autonomous remediation, intelligent alert management, and operational copilots * Identify opportunities to eliminate manual operational work through automation and machine-driven decision support * Establish AI-powered workflows for: * Incident detection and triage * Root cause analysis * Event correlation * Capacity forecasting * Reliability risk identification * Automated remediation * Knowledge management * Operational reporting * Deliver measurable reductions in Mean Time to Detect (MTTD), Mean Time to Resolve (MTTR), operational toil, and incident volume Portfolio Reliability Management * Develop a portfolio-wide reliability framework capable of managing highly heterogeneous technology stacks including: + Mainframe platforms + Middleware and integration technologies + Distributed applications + Containerized workloads + Public cloud platforms + Vendor-hosted applications + SaaS ecosystems * Establish application criticality tiers and reliability targets across the portfolio * Implement standardized observability and operational telemetry strategies regardless of technology platform Team Leadership * Build, lead, and mentor a high-performing team of direct reports and contractors * Create an elite SRE organization of five or fewer highly skilled engineers capable of delivering enterprise-scale outcomes through leverage, automation, and platform capabilities * Recruit and develop engineers with solid expertise in software engineering, automation, observability, AI, and systems reliability * Foster a culture of ownership, innovation, operational excellence, and continuous improvement Vendor and Partner Management * Drive reliability accountability across a complex ecosystem of vendors, managed service providers, and third-party technology partners * Establish operational performance expectations, reliability metrics, service level agreements, and governance mechanisms with external partners * Ensure vendors contribute actionable telemetry, operational transparency, and incident management discipline * Lead escalations and executive-level discussions related to service disruptions and reliability concerns Observability and Platform Engineering * Define and implement enterprise observability standards across metrics, logs, traces, events, synthetic monitoring, and user experience monitoring * Drive platform engineering initiatives that simplify operational support and reduce application-specific operational burden * Establish self-service reliability capabilities for application teams Incident and Resilience Management * Lead major incident management, post-incident review processes, and enterprise resilience initiatives * Drive systemic problem elimination through engineering-led root cause analysis and preventive action programs * Develop disaster recovery, business continuity, and resiliency testing strategies. * Ensure reliability practices are embedded throughout the software development lifecycle You'll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in. ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Best Paying Remote Jobs](https://www.wearedevelopers.com/magazine/255-best-paying-remote-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best Job Boards for Remote Work for Developers](https://www.wearedevelopers.com/magazine/290-best-job-boards-for-remote-work-for-developers) - [Remote, Hybrid, or In-Office: What’s Really Best for Developers?](https://www.wearedevelopers.com/magazine/638-remote-hybrid-or-in-office-what-s-really-best-for-developers)