> Markdown version of [/jobs/ext/2010925-lead-applied-ai-site-reliability-engineer-ii-pxe-a-a](https://www.wearedevelopers.com/jobs/ext/2010925-lead-applied-ai-site-reliability-engineer-ii-pxe-a-a). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Applied AI Site Reliability Engineer II - PxE A&A - **Company:** Deloitte T.T.L. - **Location:** Morristown, NJ, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Cloud Engineering, Continuous Integration, Data Governance, Data Structures, Fault Tolerance, Github, Load Testing, Object-Oriented Software Development, Reliability Engineering, Site Reliability Engineering Practices, Runbook, Software Engineering, Autoscaling, Kubernetes, Extreme Programming, Machine Learning Operations, Terraform, Devsecops - **Published:** August 10, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/lead-applied-ai-site-reliability-engineer-ii-pxe-a-and-a-morristown-nj-usa-58881368 ## About the Role to mentor engineers to improve reliability and reduce toil * Define service-level objectives with cross-functional teams and verify readiness for production * Collaborate with engineering, security, risk, and data governance to ensure compliant production paths * Champion AI/ML production reliability including drift, skew, latency, and cost anomalies * Develop and maintain runbooks, playbooks, and automation for scalable operational delivery * Engage with product teams pre-, during, and post-delivery to align reliability with business goals * Foster a culture of continuous improvement through chaos testing, capacity planning, and cost engineering Tasks * 6+ years in software engineering and site reliability in large-scale, cloud-native systems * 3+ years owning SLIs/SLOs/SLAs, incident management, and production observability stacks * 3+ years cloud-native engineering on Azure, AWS, or GCP with Kubernetes, Terraform, and CI/CD * 1+ years establishing reliability standards, mentoring aaaav _ and driving SLO discipline * Experience operating AI/ML workloads in production, including MLOps/LLMOps and AI control planes * Proficiency with load testing, chaos engineering, capacity planning, autoscaling, and FinOps tooling * Strong understanding of OOP/OOD, data structures, algorithms, and AI-augmented development * Experience with tools/frameworks like XP, Lean, DevSecOps, SRE, GitHub, MLflow, LangFuse/LangSmith * Location within commutable distance to select Deloitte locations and ability to travel ~10% Key requirements * ## Description Experteer Overview In this role you will lead hands-on reliability, performance, and operational integrity for high-visibility products and AI-infused platforms. You will set production standards, drive observability and SRE practices, and mentor teammates to deliver resilient, cost-aware systems at scale. You'll operate across cloud, AI/agentic workloads, and complex environments, ensuring safe production gating and graceful degradation. This opportunity emphasizes shaping production reliability as a strategic value for Deloitte's engineering investments, with a focus on measurable outcomes and cross-functional collaboration. Compensation / Benefits * Drive reliability, performance, and cost outcomes using SLOs and error budgets * Lead design of observability, resilience testing, and operational tooling; gate production admissions based on budgets and reliability checks * Own SLOs, alerting, runbooks, incident postmortems, and automation to improve KPIs * Lead blameless postmortems and mentor engineers to improve reliability and reduce toil * Define service-level objectives with cross-functional teams and verify readiness for production * Collaborate with engineering, security, risk, and data governance to ensure compliant production paths * Champion AI/ML production reliability including drift, skew, latency, and cost anomalies * Develop and maintain runbooks, playbooks, and automation for scalable operational delivery * Engage with product teams pre-, during, and post-delivery to align reliability with business goals * Foster a culture of continuous improvement through chaos testing, capacity planning, and cost engineering Tasks * 6+ years in software engineering and site reliability in large-scale, cloud-native systems * 3+ years owning SLIs/SLOs/SLAs, incident management, and production observability stacks * 3+ years cloud-native engineering on Azure, AWS, or GCP with Kubernetes, Terraform, and CI/CD * 1+ years establishing reliability standards, mentoring teams, and driving SLO discipline * Experience operating AI/ML workloads in production, including MLOps/LLMOps and AI control planes * Proficiency with load testing, chaos engineering, capacity planning, autoscaling, and FinOps tooling * Strong understanding of OOP/OOD, data structures, algorithms, and AI-augmented development * Experience with tools/frameworks like XP, Lean, DevSecOps, SRE, GitHub, MLflow, LangFuse/LangSmith * Location within commutable distance to select Deloitte locations and ability to travel ~10% Key requirements * ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [DevSecOps: Injecting Security into Mobile CI/CD Pipelines](https://www.wearedevelopers.com/videos/273-devsecops-injecting-security-into-mobile-ci-cd-pipelines) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [Bridging AI and Nomad: a Go-based MCP Server for Cluster Control](https://www.wearedevelopers.com/videos/2063-bridging-ai-and-nomad-a-go-based-mcp-server-for-cluster-control) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)