Lead Applied AI Site Reliability Engineer II - PxE A&A

Deloitte T.T.L.
Tampa, FL, United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Cloud Engineering Continuous Integration Data Structures Python (Programming Language) Key Management Load Testing Object-Oriented Software Development
+14 more
Pair Programming Role-Based Access Control Reliability Engineering Prometheus Software Engineering Datadog Scripting Autoscaling Grafana Kubernetes Machine Learning Operations Cloudwatch Dynatrace Golang

Job description

Experteer Overview In this role you drive reliability, performance, and cost efficiency for high-visibility products and AI workloads across cloud environments. You will lead production standards, design observability and resilience tooling, and ensure safe, scalable production releases. You mentor teams, partner with cross-functional groups, and push pragmatic improvements that reduce toil and incidents. This position combines SRE craft with applied AI fluency to operate AI-enabled platforms at scale. You will influence both technology and processes to deliver dependable, cost-aware systems. Compensation / Benefits * Drive reliability and cost outcomes using SLOs and error budgets * Set production standards and design observability, performance, and resilience testing * Own admission of systems into production with readiness checks and automated reliability gates * Develop runbooks, playbooks, and automation; lead blameless postmortems * Collaborate with cross-functional teams to co-define SLIs/SLOs and ensure readiness * Lead and mentor engineers in reliability practices and continuous improvement * Execute chaos testing, capacity planning, and cost engineering to prove readiness * Maintain hands-on ownership of production and pre-production environments to prevent drift * Communicate technically complex concepts clearly to diverse stakeholders * Foster a collaborative, cross-functional culture focused on quality and trust Tasks * 6+ years of software engineering and SRE experience operating large-scale, cloud-native systems * 3+ years defining/owning SLIs, SLOs, SLAs, and incident management including on-call * 3+ years of cloud-native engineering across Azure/AWS/GCP and Kubernetes, containers, infrastructure-as-code * Experience with CI/CD, observability stacks and production-grade monitoring (Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, CloudWatch, etc.) * Experience with AI/ML workloads in production, MLOps, and AI control planes (guardrails, model gateways) * Familiar with load testing, chaos engineering, autoscaling, FinOps tooling and GPU/token cost attribution * Strong coding and scripting in languages such as Python, Go, Bash; understanding of OOP/OOD and data structures * Knowledge of security/risk controls including least-privilege, RBAC, secrets management * Excellent communication and leadership abilities; comfortable mentoring and pair programming Key requirements *

Requirements

years SLIs/SLOs and ensure readiness * Lead and mentor engineers in reliability practices and continuous improvement * Execute chaos testing, capacity planning, and cost engineering to prove readiness * Maintain hands-on ownership of production and pre-production environments to prevent drift * Communicate technically complex concepts clearly to diverse stakeholders * Foster a collaborative, cross-functional culture focused on quality and trust Tasks * 6+ years of software engineering and SRE experience operating large-scale, cloud-native systems * 3+ years defining/owning SLIs, SLOs, SLAs, and incident management including on-call * 3+ years of cloud-native engineering across Azure/AWS/GCP and Kubernetes, containers, infrastructure-as-code * Experience with CI/CD, observability stacks and production-grade monitoring (Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, CloudWatch, etc.) * Experience with AI/ML workloads in production, MLOps, and AI control planes (guardrails, model aa secrets * Familiar with load testing, chaos engineering, autoscaling, FinOps tooling and GPU/token cost attribution * Strong coding and scripting in languages such as Python, Go, Bash; understanding of OOP/OOD and data structures * Knowledge of security/risk controls including least-privilege, RBAC, secrets management * Excellent communication and leadership abilities; comfortable mentoring and pair programming Key requirements *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · WWC Europe 2026

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all