Lead Applied AI Site Reliability Engineer II - PxE ERM

Deloitte T.T.L.
Morristown, NJ, United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
1 year minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Cloud Engineering Data Structures Github Machine Learning Reliability Engineering Site Reliability Engineering Practices Runbook Software Engineering
+9 more
SonarQube Cloud Platform System Performance Testing Autoscaling Kubernetes Extreme Programming Machine Learning Operations Virtual Agents Devsecops

Job description

Experteer Overview In this Lead Applied AI SRE role, you ensure reliability, performance, and cost efficiency for high-visibility products and AI-infused workloads. You will lead by example, mentoring teams and setting production standards across cross-functional partners. You’ll design observability, SLOs, and automated reliability checks to prevent incidents and enable safe, scalable releases. This role blends cloud platform engineering with applied AI fluency to optimize operations and drive business value. You will partner with engineering leadership to shape resilient systems and codify best practices that scale. Compensation / Benefits * Drive reliability, performance, and cost outcomes using SLOs and error budgets; prioritize toil and incidents to improve production resilience * Act as technical advocate for production reliability; set standards, design observability and resilience tooling, and gate production readiness * Own operational integrity of production and pre-production environments; manage SLOs, alerting, drift prevention, and security-conscious control measures * Lead development of runbooks, playbooks, postmortems, and automation; mentor peers to meet reliability KPIs * Collaborate with cross-functional teams to define and satisfy service-level objectives and production readiness * Apply advanced SRE practices to cloud-native, AI-infused workloads, ensuring safe degradation and controlled risk Tasks * Bachelor degree in CS, software engineering, data science, ML, or related field * 6+ years software engineering and SRE experience with large-scale cloud-native systems * 3+ years SRE/production engineering focusing on SLIs/SLOs/SLAs; incident response; observability stacks * 3+ years cloud-native engineering on Azure, AWS, or GCP; container orchestration; IaC; networking; multi-environment mgmt * 1+ year establishing reliability standards (SLO discipline, runbooks, budgets) and mentoring teams * Experience operating AI/ML and agentic workloads; familiarity with MLOps/LLMOps and AI control planes * Experience with load/performance testing, chaos engineering, capacity planning, autoscaling, and FinOps tooling * Software engineering background with understanding of diagrams, data structures, algorithms, and AI-driven development * Experience with XP/Lean/DevSecOps/SRE/CI tooling (GitHub, SonarQube, MLflow) and agentic AI frameworks Key requirements *

Requirements

environments; manage SLOs, alerting, drift prevention, and security-conscious control measures * Lead development of runbooks, playbooks, postmortems, and automation; mentor peers to meet reliability KPIs * Collaborate with cross-functional teams to define and satisfy service-level objectives and production readiness * Apply advanced SRE practices to cloud-native, AI-infused workloads, ensuring safe degradation and controlled risk Tasks * Bachelor degree in CS, software engineering, data science, ML, or related field * 6+ years software engineering and SRE experience with large-scale cloud-native systems * 3+ years SRE/production engineering focusing on SLIs/SLOs/SLAs; incident response; observability stacks * 3+ years cloud-native engineering on Azure, AWS, or GCP; container orchestration; IaC; networking; multi-environment mgmt * 1+ year establishing reliability standards (SLO discipline, runbooks, budgets) and mentoring teams * Experience operating AI/ML and agentic workloads; aaaaaaaaau _ with MLOps/LLMOps and AI control planes * Experience with load/performance testing, chaos engineering, capacity planning, autoscaling, and FinOps tooling * Software engineering background with understanding of diagrams, data structures, algorithms, and AI-driven development * Experience with XP/Lean/DevSecOps/SRE/CI tooling (GitHub, SonarQube, MLflow) and agentic AI frameworks Key requirements *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:15 min

Key lessons learned from implementing automated mobile DevSecOps

Moataz Nabil Moataz Nabil · LIVE

2:50 min

Introduction and the value of runbooks

Hila Fish · WWC 2023

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

1:34 min

Transitioning from traditional software development to artificial intelligence consulting

Patrick Schnell Patrick Schnell · Coffee With Developers

2:09 min

Shifting security left using the DevSecOps approach

Aarno Aukia · LIVE

Videos

See all

Related articles

See all