Lead Applied AI Site Reliability Engineer II - PxE A&A

Deloitte T.T.L.
Morristown, NJ, United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
1 year minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Cloud Engineering Continuous Integration Data Governance Data Structures Fault Tolerance Github Load Testing Object-Oriented Software Development
+10 more
Reliability Engineering Site Reliability Engineering Practices Runbook Software Engineering Autoscaling Kubernetes Extreme Programming Machine Learning Operations Terraform Devsecops

Job description

Experteer Overview In this role you will lead hands-on reliability, performance, and operational integrity for high-visibility products and AI-infused platforms. You will set production standards, drive observability and SRE practices, and mentor teammates to deliver resilient, cost-aware systems at scale. You’ll operate across cloud, AI/agentic workloads, and complex environments, ensuring safe production gating and graceful degradation. This opportunity emphasizes shaping production reliability as a strategic value for Deloitte’s engineering investments, with a focus on measurable outcomes and cross-functional collaboration. Compensation / Benefits * Drive reliability, performance, and cost outcomes using SLOs and error budgets * Lead design of observability, resilience testing, and operational tooling; gate production admissions based on budgets and reliability checks * Own SLOs, alerting, runbooks, incident postmortems, and automation to improve KPIs * Lead blameless postmortems and mentor engineers to improve reliability and reduce toil * Define service-level objectives with cross-functional teams and verify readiness for production * Collaborate with engineering, security, risk, and data governance to ensure compliant production paths * Champion AI/ML production reliability including drift, skew, latency, and cost anomalies * Develop and maintain runbooks, playbooks, and automation for scalable operational delivery * Engage with product teams pre-, during, and post-delivery to align reliability with business goals * Foster a culture of continuous improvement through chaos testing, capacity planning, and cost engineering Tasks * 6+ years in software engineering and site reliability in large-scale, cloud-native systems * 3+ years owning SLIs/SLOs/SLAs, incident management, and production observability stacks * 3+ years cloud-native engineering on Azure, AWS, or GCP with Kubernetes, Terraform, and CI/CD * 1+ years establishing reliability standards, mentoring teams, and driving SLO discipline * Experience operating AI/ML workloads in production, including MLOps/LLMOps and AI control planes * Proficiency with load testing, chaos engineering, capacity planning, autoscaling, and FinOps tooling * Strong understanding of OOP/OOD, data structures, algorithms, and AI-augmented development * Experience with tools/frameworks like XP, Lean, DevSecOps, SRE, GitHub, MLflow, LangFuse/LangSmith * Location within commutable distance to select Deloitte locations and ability to travel ~10% Key requirements *

Requirements

to mentor engineers to improve reliability and reduce toil * Define service-level objectives with cross-functional teams and verify readiness for production * Collaborate with engineering, security, risk, and data governance to ensure compliant production paths * Champion AI/ML production reliability including drift, skew, latency, and cost anomalies * Develop and maintain runbooks, playbooks, and automation for scalable operational delivery * Engage with product teams pre-, during, and post-delivery to align reliability with business goals * Foster a culture of continuous improvement through chaos testing, capacity planning, and cost engineering Tasks * 6+ years in software engineering and site reliability in large-scale, cloud-native systems * 3+ years owning SLIs/SLOs/SLAs, incident management, and production observability stacks * 3+ years cloud-native engineering on Azure, AWS, or GCP with Kubernetes, Terraform, and CI/CD * 1+ years establishing reliability standards, mentoring aaaav _ and driving SLO discipline * Experience operating AI/ML workloads in production, including MLOps/LLMOps and AI control planes * Proficiency with load testing, chaos engineering, capacity planning, autoscaling, and FinOps tooling * Strong understanding of OOP/OOD, data structures, algorithms, and AI-augmented development * Experience with tools/frameworks like XP, Lean, DevSecOps, SRE, GitHub, MLflow, LangFuse/LangSmith * Location within commutable distance to select Deloitte locations and ability to travel ~10% Key requirements *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

1:15 min

Key lessons learned from implementing automated mobile DevSecOps

Moataz Nabil Moataz Nabil · LIVE

2:50 min

Introduction and the value of runbooks

Hila Fish · WWC 2023

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:09 min

Shifting security left using the DevSecOps approach

Aarno Aukia · LIVE

Videos

See all

Related articles

See all