> Markdown version of [/jobs/ext/2013419-lead-applied-ai-site-reliability-engineer-ii-pxe-a-a](https://www.wearedevelopers.com/jobs/ext/2013419-lead-applied-ai-site-reliability-engineer-ii-pxe-a-a). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Applied AI Site Reliability Engineer II - PxE A&A - **Company:** Deloitte T.T.L. - **Location:** Nashville, TN, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), .NET Framework, Artificial Intelligence, Amazon Web Services, Microsoft Azure, Bash Shell, C Sharp (Programming Language), Cloud Computing, Cloud Engineering, Continuous Integration, Data Governance, Data Structures, Python (Programming Language), Load Testing, Network Control, NoSQL, Octopus Deploy, Object-Oriented Software Development, Reliability Engineering, Prometheus, Azure Machine Learning, Runbook, SQL Databases, Datadog, Google Cloud, Cloud Monitoring, Autoscaling, Grafana, Kubernetes, Extreme Programming, Machine Learning Operations, Cloudwatch, Terraform, Dynatrace, Devsecops, Golang - **Published:** August 10, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/lead-applied-ai-site-reliability-engineer-ii-pxe-a-and-a-hermitage-tn-usa-58881367 ## About the Role for production releases. * Foster cross-functional collaboration with engineering, security, data governance, and leadership to align on standards and controls. * Advance production engineering practices including chaos testing, capacity planning, and cloud/AI cost engineering. * Mentor and review peers to ensure reliability KPIs (availability, performance, cost) are met or exceeded. * Communicate complex technical concepts clearly to stakeholders and influence decision making. Tasks * 6+ years of software and site reliability engineering experience with large-scale, cloud-native systems. * Proficiency in Python, Go, Bash, Java, C#/.NET, SQL/NoSQL; strong Kubernetes, Terraform, ArgoCD experience; CI/CD and observability stack. * 3+ years owning SLIs, SLOs, SLAs, incident command, on-call, and production observability tools (OpenTelemetry, Prometheus, Grafana, Datadog, Dynatrace, CloudWatch, Azure Monitor, Google Cloud Operations). * 3+ years cloud-native engineering on Azure/AWS/GCP aaaaaaaw _ AI/ML services and container orchestration; IaC, networking, multi-environment management. * 1+ year establishing reliability standards (SLO discipline, runbooks, performance budgets) and mentoring teams. * Experience operating AI/ML workloads, including MLOps/LLMOps, AI control plane, and production reliability considerations. * Experience with load testing, chaos engineering, capacity planning, autoscaling, and cloud/AI cost attribution tooling. * Software engineering fundamentals (OOP/OOD, data structures, algorithms) and AI-augmented development; familiarity with XP/Lean/DevSecOps/SRE tooling. Key requirements * Discretionary annual incentive program ## Description Experteer Overview As a Lead Applied AI SRE II, you will own reliability, performance, and cost outcomes for high-visibility products and AI-infused systems. You'll lead with strong engineering craft across cloud platforms, observability, and AI-enabled workloads, driving scalable, resilient production. You'll set standards, mentor teams, and advocate for production readiness and blameless learning. This role blends hands-on engineering with cross-functional leadership to deliver value through reliable, cost-aware operations. Compensation / Benefits * Own SLOs, error budgets, and incident trends to improve reliability and reduce toil. * Lead design of observability, performance and resilience testing, and operational tooling; gate production admissions using automated reliability checks. * Maintain production and pre-production environments, build SLO-driven dashboards, and drive runbooks and postmortems. * Co-create service-level objectives with engineering teams and ensure readiness for production releases. * Foster cross-functional collaboration with engineering, security, data governance, and leadership to align on standards and controls. * Advance production engineering practices including chaos testing, capacity planning, and cloud/AI cost engineering. * Mentor and review peers to ensure reliability KPIs (availability, performance, cost) are met or exceeded. * Communicate complex technical concepts clearly to stakeholders and influence decision making. Tasks * 6+ years of software and site reliability engineering experience with large-scale, cloud-native systems. * Proficiency in Python, Go, Bash, Java, C#/.NET, SQL/NoSQL; strong Kubernetes, Terraform, ArgoCD experience; CI/CD and observability stack. * 3+ years owning SLIs, SLOs, SLAs, incident command, on-call, and production observability tools (OpenTelemetry, Prometheus, Grafana, Datadog, Dynatrace, CloudWatch, Azure Monitor, Google Cloud Operations). * 3+ years cloud-native engineering on Azure/AWS/GCP including AI/ML services and container orchestration; IaC, networking, multi-environment management. * 1+ year establishing reliability standards (SLO discipline, runbooks, performance budgets) and mentoring teams. * Experience operating AI/ML workloads, including MLOps/LLMOps, AI control plane, and production reliability considerations. * Experience with load testing, chaos engineering, capacity planning, autoscaling, and cloud/AI cost attribution tooling. * Software engineering fundamentals (OOP/OOD, data structures, algorithms) and AI-augmented development; familiarity with XP/Lean/DevSecOps/SRE tooling. Key requirements * Discretionary annual incentive program ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Navigating the AI Wave in DevOps](https://www.wearedevelopers.com/videos/853-navigating-the-ai-wave-in-devops) - [NoSQL Data Modeling for Front-end Developers](https://www.wearedevelopers.com/videos/297-nosql-data-modeling-for-front-end-developers) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)