Lead Applied AI Site Reliability Engineer II - PxE A&A
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+12 more
Job description
Experteer Overview In this role you will drive reliability, performance, and cost efficiency for high-visibility products and AI-enabled workloads. You’ll lead with hands-on engineering across cloud platforms, observability, and performance/RC engineering while mentoring teams to meet production readiness. You’ll partner with cross-functional colleagues to set standards and gate production against SLOs and automated reliability checks. This is a chance to shape end-to-end reliability at scale within Deloitte’s Technology Product Engineering team. Compensation / Benefits * Drive reliability, performance, and cost outcomes through SLOs and error budgets; prioritize work based on incident trends and toil * Lead production standards and design observability, performance, and resilience testing; gate systems into production * Own operational integrity of production and pre-production environments; create runbooks, playbooks, and automation; mentor others * Develop lean operational solutions through rapid experimentation; co-define SLOs and readiness with engineering teams * Collaborate with cross-functional partners to ensure secure, compliant, and reliable delivery; gate admissions into production * Advocate for reliability and operability; ensure systems degrade gracefully and operate within budgets * Lead blameless postmortems and drive systemic fixes; promote continuous improvement of reliability KPIs Tasks * 6+ years of software engineering and SRE experience with large-scale distributed cloud-native systems * Proficiency with Python, Go, Bash, Java, C#/.NET, SQL/NoSQL; Kubernetes; Terraform; ArgoCD; CI/CD and observability stacks * 3+ years defining and owning SLI/SLO/SLA, error budgets, incident command/on-call; building/operating observability stacks * 3+ years cloud-native engineering on Azure/AWS/GCP; container orchestration; IaC; multi-environment management * 1+ years establishing reliability standards (SLO discipline, runbooks, budgets) and mentoring teams * Experience operating AI/ML/agentic workloads; MLOps/LLMOps; AI control plane; drift and cost anomalies * Load/perf testing (LoadRunner/k6/JMeter); chaos engineering (Azure Chaos Studio, AWS Fault Injector); capacity planning; autoscaling; FinOps tooling Key requirements *
Requirements
through rapid experimentation; co-define SLOs and readiness with engineering teams * Collaborate with cross-functional partners to ensure secure, compliant, and reliable delivery; gate admissions into production * Advocate for reliability and operability; ensure systems degrade gracefully and operate within budgets * Lead blameless postmortems and drive systemic fixes; promote continuous improvement of reliability KPIs Tasks * 6+ years of software engineering and SRE experience with large-scale distributed cloud-native systems * Proficiency with Python, Go, Bash, Java, C#/.NET, SQL/NoSQL; Kubernetes; Terraform; ArgoCD; CI/CD and observability stacks * 3+ years defining and owning SLI/SLO/SLA, error budgets, incident command/on-call; building/operating observability stacks * 3+ years cloud-native engineering on Azure/AWS/GCP; container orchestration; IaC; multi-environment management * 1+ years establishing reliability standards (SLO discipline, runbooks, budgets) and mentoring teams * aaaa aaaaC_ operating AI/ML/agentic workloads; MLOps/LLMOps; AI control plane; drift and cost anomalies * Load/perf testing (LoadRunner/k6/JMeter); chaos engineering (Azure Chaos Studio, AWS Fault Injector); capacity planning; autoscaling; FinOps tooling Key requirements *
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
Dev Digest 120 - Apple and peers
Navigating the AI Shift
Is Software Engineering Over-Saturated?