Lead Applied AI Site Reliability Engineer II - PxE A&A

Deloitte T.T.L.
Jersey City, NJ, United States
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
1 year minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Cloud Engineering Data Governance Github Load Testing Octopus Deploy Reliability Engineering Site Reliability Engineering Practices Runbook
+9 more
Software Engineering SonarQube Cloud Platform System Autoscaling Kubernetes Extreme Programming Machine Learning Operations Terraform Devsecops

Job description

Experteer Overview In this role you will drive reliability, performance, and cost efficiency for high-visibility AI-enabled products and platforms. You will lead reliability practices, set production standards, and mentor teams to admit systems into production with confidence. You will operate across cloud, observability, and AI workloads to keep systems resilient at scale. You’ll collaborate with cross-functional partners to deliver accountable, measurable improvements aligned with business outcomes. This is a hands-on, impact-driven engineering leadership opportunity at Deloitte. Compensation / Benefits * Outcome-driven accountability for reliability, performance, and cost, using SLOs and error budgets * Technical leadership to design observability, resilience testing, and production tooling * Hands-on operational ownership of production and pre-production environments * Customer-centric engineering to define readiness and safeguards with engineering teams * Incremental delivery and progressive resilience improvements with measurable impact * Cross-functional collaboration with engineering, security, data governance, and leadership * Advanced technical proficiency in cloud platform ownership, observability, and AI/agentic workloads * Domain expertise on AI infused platforms, drift, latency, and cost anomalies * Effective communication and influence to align solutions with business goals * Engagement and co-creation with product teams and customers to drive momentum Tasks * 6+ years in software and site reliability engineering for large-scale, cloud-native systems * 3+ years owning SLIs/SLOs/SLAs, error budgets, incident response, and production observability * 3+ years cloud-native engineering on Azure/AWS/GCP with Kubernetes, Terraform, ArgoCD * 1+ years establishing reliability standards, runbooks, and performance budgets * Experience operating AI/ML workloads, MLOps, and AI control plane considerations * Experience with load testing, chaos engineering, capacity planning, autoscaling, and FinOps * Solid software engineering foundation and familiarity with AI frameworks and agentic tools * Experience with XP/Lean/DevSecOps/SRE practices and tools (GitHub, ADO, SonarQube, MLflow) * Willingness to travel up to 10% and be located within commutable distance to select locations Key requirements * broad range of benefits

Requirements

progressive resilience improvements with measurable impact * Cross-functional collaboration with engineering, security, data governance, and leadership * Advanced technical proficiency in cloud platform ownership, observability, and AI/agentic workloads * Domain expertise on AI infused platforms, drift, latency, and cost anomalies * Effective communication and influence to align solutions with business goals * Engagement and co-creation with product teams and customers to drive momentum Tasks * 6+ years in software and site reliability engineering for large-scale, cloud-native systems * 3+ years owning SLIs/SLOs/SLAs, error budgets, incident response, and production observability * 3+ years cloud-native engineering on Azure/AWS/GCP with Kubernetes, Terraform, ArgoCD * 1+ years establishing reliability standards, runbooks, and performance budgets * Experience operating AI/ML workloads, MLOps, and AI control plane considerations * Experience with load testing, chaos engineering, capacity a distance autoscaling, and FinOps * Solid software engineering foundation and familiarity with AI frameworks and agentic tools * Experience with XP/Lean/DevSecOps/SRE practices and tools (GitHub, ADO, SonarQube, MLflow) * Willingness to travel up to 10% and be located within commutable distance to select locations Key requirements * broad range of benefits

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

1:15 min

Key lessons learned from implementing automated mobile DevSecOps

Moataz Nabil Moataz Nabil · LIVE

2:50 min

Introduction and the value of runbooks

Hila Fish · World Congress 2023

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:09 min

Shifting security left using the DevSecOps approach

Aarno Aukia · LIVE

Videos

See all

Related articles

See all