> Markdown version of [/jobs/ext/2014353-lead-applied-ai-site-reliability-engineer-ii-pxe-a-a](https://www.wearedevelopers.com/jobs/ext/2014353-lead-applied-ai-site-reliability-engineer-ii-pxe-a-a). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Applied AI Site Reliability Engineer II - PxE A&A - **Company:** Deloitte T.T.L. - **Location:** Tampa, FL, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Bash Shell, Cloud Computing, Cloud Engineering, Continuous Integration, Data Structures, Python (Programming Language), Key Management, Load Testing, Object-Oriented Software Development, Pair Programming, Role-Based Access Control, Reliability Engineering, Prometheus, Software Engineering, Datadog, Scripting, Autoscaling, Grafana, Kubernetes, Machine Learning Operations, Cloudwatch, Dynatrace, Golang - **Published:** August 10, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/lead-applied-ai-site-reliability-engineer-ii-pxe-a-and-a-tampa-fl-usa-58881370 ## About the Role years SLIs/SLOs and ensure readiness * Lead and mentor engineers in reliability practices and continuous improvement * Execute chaos testing, capacity planning, and cost engineering to prove readiness * Maintain hands-on ownership of production and pre-production environments to prevent drift * Communicate technically complex concepts clearly to diverse stakeholders * Foster a collaborative, cross-functional culture focused on quality and trust Tasks * 6+ years of software engineering and SRE experience operating large-scale, cloud-native systems * 3+ years defining/owning SLIs, SLOs, SLAs, and incident management including on-call * 3+ years of cloud-native engineering across Azure/AWS/GCP and Kubernetes, containers, infrastructure-as-code * Experience with CI/CD, observability stacks and production-grade monitoring (Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, CloudWatch, etc.) * Experience with AI/ML workloads in production, MLOps, and AI control planes (guardrails, model aa secrets * Familiar with load testing, chaos engineering, autoscaling, FinOps tooling and GPU/token cost attribution * Strong coding and scripting in languages such as Python, Go, Bash; understanding of OOP/OOD and data structures * Knowledge of security/risk controls including least-privilege, RBAC, secrets management * Excellent communication and leadership abilities; comfortable mentoring and pair programming Key requirements * ## Description Experteer Overview In this role you drive reliability, performance, and cost efficiency for high-visibility products and AI workloads across cloud environments. You will lead production standards, design observability and resilience tooling, and ensure safe, scalable production releases. You mentor teams, partner with cross-functional groups, and push pragmatic improvements that reduce toil and incidents. This position combines SRE craft with applied AI fluency to operate AI-enabled platforms at scale. You will influence both technology and processes to deliver dependable, cost-aware systems. Compensation / Benefits * Drive reliability and cost outcomes using SLOs and error budgets * Set production standards and design observability, performance, and resilience testing * Own admission of systems into production with readiness checks and automated reliability gates * Develop runbooks, playbooks, and automation; lead blameless postmortems * Collaborate with cross-functional teams to co-define SLIs/SLOs and ensure readiness * Lead and mentor engineers in reliability practices and continuous improvement * Execute chaos testing, capacity planning, and cost engineering to prove readiness * Maintain hands-on ownership of production and pre-production environments to prevent drift * Communicate technically complex concepts clearly to diverse stakeholders * Foster a collaborative, cross-functional culture focused on quality and trust Tasks * 6+ years of software engineering and SRE experience operating large-scale, cloud-native systems * 3+ years defining/owning SLIs, SLOs, SLAs, and incident management including on-call * 3+ years of cloud-native engineering across Azure/AWS/GCP and Kubernetes, containers, infrastructure-as-code * Experience with CI/CD, observability stacks and production-grade monitoring (Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, CloudWatch, etc.) * Experience with AI/ML workloads in production, MLOps, and AI control planes (guardrails, model gateways) * Familiar with load testing, chaos engineering, autoscaling, FinOps tooling and GPU/token cost attribution * Strong coding and scripting in languages such as Python, Go, Bash; understanding of OOP/OOD and data structures * Knowledge of security/risk controls including least-privilege, RBAC, secrets management * Excellent communication and leadership abilities; comfortable mentoring and pair programming Key requirements * ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)