> Markdown version of [/jobs/ext/2017422-lead-applied-ai-site-reliability-engineer-ii-pxe-erm](https://www.wearedevelopers.com/jobs/ext/2017422-lead-applied-ai-site-reliability-engineer-ii-pxe-erm). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Applied AI Site Reliability Engineer II - PxE ERM - **Company:** Deloitte T.T.L. - **Location:** Nashville, TN, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Cloud Engineering, Data Structures, Github, Machine Learning, Reliability Engineering, Site Reliability Engineering Practices, Runbook, Software Engineering, SonarQube, Cloud Platform System, Performance Testing, Autoscaling, Kubernetes, Extreme Programming, Machine Learning Operations, Virtual Agents, Devsecops - **Published:** August 10, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/lead-applied-ai-site-reliability-engineer-ii-pxe-erm-hermitage-tn-usa-58880226 ## About the Role environments; manage SLOs, alerting, drift prevention, and security-conscious control measures * Lead development of runbooks, playbooks, postmortems, and automation; mentor peers to meet reliability KPIs * Collaborate with cross-functional teams to define and satisfy service-level objectives and production readiness * Apply advanced SRE practices to cloud-native, AI-infused workloads, ensuring safe degradation and controlled risk Tasks * Bachelor degree in CS, software engineering, data science, ML, or related field * 6+ years software engineering and SRE experience with large-scale cloud-native systems * 3+ years SRE/production engineering focusing on SLIs/SLOs/SLAs; incident response; observability stacks * 3+ years cloud-native engineering on Azure, AWS, or GCP; container orchestration; IaC; networking; multi-environment mgmt * 1+ year establishing reliability standards (SLO discipline, runbooks, budgets) and mentoring teams * Experience operating AI/ML and agentic workloads; aaaaaaaaat _ with MLOps/LLMOps and AI control planes * Experience with load/performance testing, chaos engineering, capacity planning, autoscaling, and FinOps tooling * Software engineering background with understanding of diagrams, data structures, algorithms, and AI-driven development * Experience with XP/Lean/DevSecOps/SRE/CI tooling (GitHub, SonarQube, MLflow) and agentic AI frameworks Key requirements * ## Description Experteer Overview In this Lead Applied AI SRE role, you ensure reliability, performance, and cost efficiency for high-visibility products and AI-infused workloads. You will lead by example, mentoring teams and setting production standards across cross-functional partners. You'll design observability, SLOs, and automated reliability checks to prevent incidents and enable safe, scalable releases. This role blends cloud platform engineering with applied AI fluency to optimize operations and drive business value. You will partner with engineering leadership to shape resilient systems and codify best practices that scale. Compensation / Benefits * Drive reliability, performance, and cost outcomes using SLOs and error budgets; prioritize toil and incidents to improve production resilience * Act as technical advocate for production reliability; set standards, design observability and resilience tooling, and gate production readiness * Own operational integrity of production and pre-production environments; manage SLOs, alerting, drift prevention, and security-conscious control measures * Lead development of runbooks, playbooks, postmortems, and automation; mentor peers to meet reliability KPIs * Collaborate with cross-functional teams to define and satisfy service-level objectives and production readiness * Apply advanced SRE practices to cloud-native, AI-infused workloads, ensuring safe degradation and controlled risk Tasks * Bachelor degree in CS, software engineering, data science, ML, or related field * 6+ years software engineering and SRE experience with large-scale cloud-native systems * 3+ years SRE/production engineering focusing on SLIs/SLOs/SLAs; incident response; observability stacks * 3+ years cloud-native engineering on Azure, AWS, or GCP; container orchestration; IaC; networking; multi-environment mgmt * 1+ year establishing reliability standards (SLO discipline, runbooks, budgets) and mentoring teams * Experience operating AI/ML and agentic workloads; familiarity with MLOps/LLMOps and AI control planes * Experience with load/performance testing, chaos engineering, capacity planning, autoscaling, and FinOps tooling * Software engineering background with understanding of diagrams, data structures, algorithms, and AI-driven development * Experience with XP/Lean/DevSecOps/SRE/CI tooling (GitHub, SonarQube, MLflow) and agentic AI frameworks Key requirements * ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [DevSecOps: Injecting Security into Mobile CI/CD Pipelines](https://www.wearedevelopers.com/videos/273-devsecops-injecting-security-into-mobile-ci-cd-pipelines) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [DevSecOps: Security in DevOps](https://www.wearedevelopers.com/videos/36-devsecops-security-in-devops) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)