Lead Applied AI Site Reliability Engineer II - PxE A&A

Deloitte T.T.L.
Dallas, TX, United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
1 year minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) .NET Framework Artificial Intelligence Amazon Web Services Microsoft Azure Bash Shell C Sharp (Programming Language) Cloud Computing Cloud Engineering Continuous Integration HP Loadrunner Apache JMeter
+12 more
Python (Programming Language) Network Control NoSQL Reliability Engineering Runbook Software Engineering SQL Databases Autoscaling Kubernetes Machine Learning Operations Terraform Golang

Job description

Experteer Overview In this role you will drive reliability, performance, and cost efficiency for high-visibility products and AI-enabled workloads. You’ll lead with hands-on engineering across cloud platforms, observability, and performance/RC engineering while mentoring teams to meet production readiness. You’ll partner with cross-functional colleagues to set standards and gate production against SLOs and automated reliability checks. This is a chance to shape end-to-end reliability at scale within Deloitte’s Technology Product Engineering team. Compensation / Benefits * Drive reliability, performance, and cost outcomes through SLOs and error budgets; prioritize work based on incident trends and toil * Lead production standards and design observability, performance, and resilience testing; gate systems into production * Own operational integrity of production and pre-production environments; create runbooks, playbooks, and automation; mentor others * Develop lean operational solutions through rapid experimentation; co-define SLOs and readiness with engineering teams * Collaborate with cross-functional partners to ensure secure, compliant, and reliable delivery; gate admissions into production * Advocate for reliability and operability; ensure systems degrade gracefully and operate within budgets * Lead blameless postmortems and drive systemic fixes; promote continuous improvement of reliability KPIs Tasks * 6+ years of software engineering and SRE experience with large-scale distributed cloud-native systems * Proficiency with Python, Go, Bash, Java, C#/.NET, SQL/NoSQL; Kubernetes; Terraform; ArgoCD; CI/CD and observability stacks * 3+ years defining and owning SLI/SLO/SLA, error budgets, incident command/on-call; building/operating observability stacks * 3+ years cloud-native engineering on Azure/AWS/GCP; container orchestration; IaC; multi-environment management * 1+ years establishing reliability standards (SLO discipline, runbooks, budgets) and mentoring teams * Experience operating AI/ML/agentic workloads; MLOps/LLMOps; AI control plane; drift and cost anomalies * Load/perf testing (LoadRunner/k6/JMeter); chaos engineering (Azure Chaos Studio, AWS Fault Injector); capacity planning; autoscaling; FinOps tooling Key requirements *

Requirements

through rapid experimentation; co-define SLOs and readiness with engineering teams * Collaborate with cross-functional partners to ensure secure, compliant, and reliable delivery; gate admissions into production * Advocate for reliability and operability; ensure systems degrade gracefully and operate within budgets * Lead blameless postmortems and drive systemic fixes; promote continuous improvement of reliability KPIs Tasks * 6+ years of software engineering and SRE experience with large-scale distributed cloud-native systems * Proficiency with Python, Go, Bash, Java, C#/.NET, SQL/NoSQL; Kubernetes; Terraform; ArgoCD; CI/CD and observability stacks * 3+ years defining and owning SLI/SLO/SLA, error budgets, incident command/on-call; building/operating observability stacks * 3+ years cloud-native engineering on Azure/AWS/GCP; container orchestration; IaC; multi-environment management * 1+ years establishing reliability standards (SLO discipline, runbooks, budgets) and mentoring teams * aaaa aaaaC_ operating AI/ML/agentic workloads; MLOps/LLMOps; AI control plane; drift and cost anomalies * Load/perf testing (LoadRunner/k6/JMeter); chaos engineering (Azure Chaos Studio, AWS Fault Injector); capacity planning; autoscaling; FinOps tooling Key requirements *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

2:50 min

Introduction and the value of runbooks

Hila Fish · WWC 2023

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:32 min

Structuring automated incident workflows between runbooks and raw models

Aram Hakobyan Aram Hakobyan +1 · WWC Europe 2026

Videos

See all

Related articles

See all