Lead Applied AI Site Reliability Engineer II - PxE A&A
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
Experteer Overview In this role you will lead hands-on reliability, performance, and operational integrity for high-visibility products and AI-infused platforms. You will set production standards, drive observability and SRE practices, and mentor teammates to deliver resilient, cost-aware systems at scale. You’ll operate across cloud, AI/agentic workloads, and complex environments, ensuring safe production gating and graceful degradation. This opportunity emphasizes shaping production reliability as a strategic value for Deloitte’s engineering investments, with a focus on measurable outcomes and cross-functional collaboration. Compensation / Benefits * Drive reliability, performance, and cost outcomes using SLOs and error budgets * Lead design of observability, resilience testing, and operational tooling; gate production admissions based on budgets and reliability checks * Own SLOs, alerting, runbooks, incident postmortems, and automation to improve KPIs * Lead blameless postmortems and mentor engineers to improve reliability and reduce toil * Define service-level objectives with cross-functional teams and verify readiness for production * Collaborate with engineering, security, risk, and data governance to ensure compliant production paths * Champion AI/ML production reliability including drift, skew, latency, and cost anomalies * Develop and maintain runbooks, playbooks, and automation for scalable operational delivery * Engage with product teams pre-, during, and post-delivery to align reliability with business goals * Foster a culture of continuous improvement through chaos testing, capacity planning, and cost engineering Tasks * 6+ years in software engineering and site reliability in large-scale, cloud-native systems * 3+ years owning SLIs/SLOs/SLAs, incident management, and production observability stacks * 3+ years cloud-native engineering on Azure, AWS, or GCP with Kubernetes, Terraform, and CI/CD * 1+ years establishing reliability standards, mentoring teams, and driving SLO discipline * Experience operating AI/ML workloads in production, including MLOps/LLMOps and AI control planes * Proficiency with load testing, chaos engineering, capacity planning, autoscaling, and FinOps tooling * Strong understanding of OOP/OOD, data structures, algorithms, and AI-augmented development * Experience with tools/frameworks like XP, Lean, DevSecOps, SRE, GitHub, MLflow, LangFuse/LangSmith * Location within commutable distance to select Deloitte locations and ability to travel ~10% Key requirements *
Requirements
to mentor engineers to improve reliability and reduce toil * Define service-level objectives with cross-functional teams and verify readiness for production * Collaborate with engineering, security, risk, and data governance to ensure compliant production paths * Champion AI/ML production reliability including drift, skew, latency, and cost anomalies * Develop and maintain runbooks, playbooks, and automation for scalable operational delivery * Engage with product teams pre-, during, and post-delivery to align reliability with business goals * Foster a culture of continuous improvement through chaos testing, capacity planning, and cost engineering Tasks * 6+ years in software engineering and site reliability in large-scale, cloud-native systems * 3+ years owning SLIs/SLOs/SLAs, incident management, and production observability stacks * 3+ years cloud-native engineering on Azure, AWS, or GCP with Kubernetes, Terraform, and CI/CD * 1+ years establishing reliability standards, mentoring aaaav _ and driving SLO discipline * Experience operating AI/ML workloads in production, including MLOps/LLMOps and AI control planes * Proficiency with load testing, chaos engineering, capacity planning, autoscaling, and FinOps tooling * Strong understanding of OOP/OOD, data structures, algorithms, and AI-augmented development * Experience with tools/frameworks like XP, Lean, DevSecOps, SRE, GitHub, MLflow, LangFuse/LangSmith * Location within commutable distance to select Deloitte locations and ability to travel ~10% Key requirements *
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Navigating the AI Shift
MLOps And AI Driven Development
Dev Digest 120 - Apple and peers
MLOps – What’s the deal behind it?