Lead Applied AI Site Reliability Engineer II - PxE A&A

Deloitte T.T.L.
Nashville, TN, United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
1 year minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) .NET Framework Artificial Intelligence Amazon Web Services Microsoft Azure Bash Shell C Sharp (Programming Language) Cloud Computing Cloud Engineering Continuous Integration Data Governance Data Structures
+24 more
Python (Programming Language) Load Testing Network Control NoSQL Octopus Deploy Object-Oriented Software Development Reliability Engineering Prometheus Azure Machine Learning Runbook SQL Databases Datadog Google Cloud Cloud Monitoring Autoscaling Grafana Kubernetes Extreme Programming Machine Learning Operations Cloudwatch Terraform Dynatrace Devsecops Golang

Job description

Experteer Overview As a Lead Applied AI SRE II, you will own reliability, performance, and cost outcomes for high-visibility products and AI-infused systems. You’ll lead with strong engineering craft across cloud platforms, observability, and AI-enabled workloads, driving scalable, resilient production. You’ll set standards, mentor teams, and advocate for production readiness and blameless learning. This role blends hands-on engineering with cross-functional leadership to deliver value through reliable, cost-aware operations. Compensation / Benefits * Own SLOs, error budgets, and incident trends to improve reliability and reduce toil. * Lead design of observability, performance and resilience testing, and operational tooling; gate production admissions using automated reliability checks. * Maintain production and pre-production environments, build SLO-driven dashboards, and drive runbooks and postmortems. * Co-create service-level objectives with engineering teams and ensure readiness for production releases. * Foster cross-functional collaboration with engineering, security, data governance, and leadership to align on standards and controls. * Advance production engineering practices including chaos testing, capacity planning, and cloud/AI cost engineering. * Mentor and review peers to ensure reliability KPIs (availability, performance, cost) are met or exceeded. * Communicate complex technical concepts clearly to stakeholders and influence decision making. Tasks * 6+ years of software and site reliability engineering experience with large-scale, cloud-native systems. * Proficiency in Python, Go, Bash, Java, C#/.NET, SQL/NoSQL; strong Kubernetes, Terraform, ArgoCD experience; CI/CD and observability stack. * 3+ years owning SLIs, SLOs, SLAs, incident command, on-call, and production observability tools (OpenTelemetry, Prometheus, Grafana, Datadog, Dynatrace, CloudWatch, Azure Monitor, Google Cloud Operations). * 3+ years cloud-native engineering on Azure/AWS/GCP including AI/ML services and container orchestration; IaC, networking, multi-environment management. * 1+ year establishing reliability standards (SLO discipline, runbooks, performance budgets) and mentoring teams. * Experience operating AI/ML workloads, including MLOps/LLMOps, AI control plane, and production reliability considerations. * Experience with load testing, chaos engineering, capacity planning, autoscaling, and cloud/AI cost attribution tooling. * Software engineering fundamentals (OOP/OOD, data structures, algorithms) and AI-augmented development; familiarity with XP/Lean/DevSecOps/SRE tooling. Key requirements * Discretionary annual incentive program

Requirements

for production releases. * Foster cross-functional collaboration with engineering, security, data governance, and leadership to align on standards and controls. * Advance production engineering practices including chaos testing, capacity planning, and cloud/AI cost engineering. * Mentor and review peers to ensure reliability KPIs (availability, performance, cost) are met or exceeded. * Communicate complex technical concepts clearly to stakeholders and influence decision making. Tasks * 6+ years of software and site reliability engineering experience with large-scale, cloud-native systems. * Proficiency in Python, Go, Bash, Java, C#/.NET, SQL/NoSQL; strong Kubernetes, Terraform, ArgoCD experience; CI/CD and observability stack. * 3+ years owning SLIs, SLOs, SLAs, incident command, on-call, and production observability tools (OpenTelemetry, Prometheus, Grafana, Datadog, Dynatrace, CloudWatch, Azure Monitor, Google Cloud Operations). * 3+ years cloud-native engineering on Azure/AWS/GCP aaaaaaaw _ AI/ML services and container orchestration; IaC, networking, multi-environment management. * 1+ year establishing reliability standards (SLO discipline, runbooks, performance budgets) and mentoring teams. * Experience operating AI/ML workloads, including MLOps/LLMOps, AI control plane, and production reliability considerations. * Experience with load testing, chaos engineering, capacity planning, autoscaling, and cloud/AI cost attribution tooling. * Software engineering fundamentals (OOP/OOD, data structures, algorithms) and AI-augmented development; familiarity with XP/Lean/DevSecOps/SRE tooling. Key requirements * Discretionary annual incentive program

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes · LIVE

Videos

See all

Related articles

See all