> Markdown version of [/jobs/ext/2015476-lead-applied-ai-site-reliability-engineer-ii-pxe-a-a](https://www.wearedevelopers.com/jobs/ext/2015476-lead-applied-ai-site-reliability-engineer-ii-pxe-a-a). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Applied AI Site Reliability Engineer II - PxE A&A - **Company:** Deloitte T.T.L. - **Location:** Dallas, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), .NET Framework, Artificial Intelligence, Amazon Web Services, Microsoft Azure, Bash Shell, C Sharp (Programming Language), Cloud Computing, Cloud Engineering, Continuous Integration, HP Loadrunner, Apache JMeter, Python (Programming Language), Network Control, NoSQL, Reliability Engineering, Runbook, Software Engineering, SQL Databases, Autoscaling, Kubernetes, Machine Learning Operations, Terraform, Golang - **Published:** August 10, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/lead-applied-ai-site-reliability-engineer-ii-pxe-a-and-a-dallas-tx-usa-58881376 ## About the Role through rapid experimentation; co-define SLOs and readiness with engineering teams * Collaborate with cross-functional partners to ensure secure, compliant, and reliable delivery; gate admissions into production * Advocate for reliability and operability; ensure systems degrade gracefully and operate within budgets * Lead blameless postmortems and drive systemic fixes; promote continuous improvement of reliability KPIs Tasks * 6+ years of software engineering and SRE experience with large-scale distributed cloud-native systems * Proficiency with Python, Go, Bash, Java, C#/.NET, SQL/NoSQL; Kubernetes; Terraform; ArgoCD; CI/CD and observability stacks * 3+ years defining and owning SLI/SLO/SLA, error budgets, incident command/on-call; building/operating observability stacks * 3+ years cloud-native engineering on Azure/AWS/GCP; container orchestration; IaC; multi-environment management * 1+ years establishing reliability standards (SLO discipline, runbooks, budgets) and mentoring teams * aaaa aaaaC_ operating AI/ML/agentic workloads; MLOps/LLMOps; AI control plane; drift and cost anomalies * Load/perf testing (LoadRunner/k6/JMeter); chaos engineering (Azure Chaos Studio, AWS Fault Injector); capacity planning; autoscaling; FinOps tooling Key requirements * ## Description Experteer Overview In this role you will drive reliability, performance, and cost efficiency for high-visibility products and AI-enabled workloads. You'll lead with hands-on engineering across cloud platforms, observability, and performance/RC engineering while mentoring teams to meet production readiness. You'll partner with cross-functional colleagues to set standards and gate production against SLOs and automated reliability checks. This is a chance to shape end-to-end reliability at scale within Deloitte's Technology Product Engineering team. Compensation / Benefits * Drive reliability, performance, and cost outcomes through SLOs and error budgets; prioritize work based on incident trends and toil * Lead production standards and design observability, performance, and resilience testing; gate systems into production * Own operational integrity of production and pre-production environments; create runbooks, playbooks, and automation; mentor others * Develop lean operational solutions through rapid experimentation; co-define SLOs and readiness with engineering teams * Collaborate with cross-functional partners to ensure secure, compliant, and reliable delivery; gate admissions into production * Advocate for reliability and operability; ensure systems degrade gracefully and operate within budgets * Lead blameless postmortems and drive systemic fixes; promote continuous improvement of reliability KPIs Tasks * 6+ years of software engineering and SRE experience with large-scale distributed cloud-native systems * Proficiency with Python, Go, Bash, Java, C#/.NET, SQL/NoSQL; Kubernetes; Terraform; ArgoCD; CI/CD and observability stacks * 3+ years defining and owning SLI/SLO/SLA, error budgets, incident command/on-call; building/operating observability stacks * 3+ years cloud-native engineering on Azure/AWS/GCP; container orchestration; IaC; multi-environment management * 1+ years establishing reliability standards (SLO discipline, runbooks, budgets) and mentoring teams * Experience operating AI/ML/agentic workloads; MLOps/LLMOps; AI control plane; drift and cost anomalies * Load/perf testing (LoadRunner/k6/JMeter); chaos engineering (Azure Chaos Studio, AWS Fault Injector); capacity planning; autoscaling; FinOps tooling Key requirements * ## Related Videos - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Bridging AI and Nomad: a Go-based MCP Server for Cluster Control](https://www.wearedevelopers.com/videos/2063-bridging-ai-and-nomad-a-go-based-mcp-server-for-cluster-control) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)