Lead Applied AI Site Reliability Engineer II - PxE A&A
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+24 more
Job description
Experteer Overview As a Lead Applied AI SRE II, you will own reliability, performance, and cost outcomes for high-visibility products and AI-infused systems. You’ll lead with strong engineering craft across cloud platforms, observability, and AI-enabled workloads, driving scalable, resilient production. You’ll set standards, mentor teams, and advocate for production readiness and blameless learning. This role blends hands-on engineering with cross-functional leadership to deliver value through reliable, cost-aware operations. Compensation / Benefits * Own SLOs, error budgets, and incident trends to improve reliability and reduce toil. * Lead design of observability, performance and resilience testing, and operational tooling; gate production admissions using automated reliability checks. * Maintain production and pre-production environments, build SLO-driven dashboards, and drive runbooks and postmortems. * Co-create service-level objectives with engineering teams and ensure readiness for production releases. * Foster cross-functional collaboration with engineering, security, data governance, and leadership to align on standards and controls. * Advance production engineering practices including chaos testing, capacity planning, and cloud/AI cost engineering. * Mentor and review peers to ensure reliability KPIs (availability, performance, cost) are met or exceeded. * Communicate complex technical concepts clearly to stakeholders and influence decision making. Tasks * 6+ years of software and site reliability engineering experience with large-scale, cloud-native systems. * Proficiency in Python, Go, Bash, Java, C#/.NET, SQL/NoSQL; strong Kubernetes, Terraform, ArgoCD experience; CI/CD and observability stack. * 3+ years owning SLIs, SLOs, SLAs, incident command, on-call, and production observability tools (OpenTelemetry, Prometheus, Grafana, Datadog, Dynatrace, CloudWatch, Azure Monitor, Google Cloud Operations). * 3+ years cloud-native engineering on Azure/AWS/GCP including AI/ML services and container orchestration; IaC, networking, multi-environment management. * 1+ year establishing reliability standards (SLO discipline, runbooks, performance budgets) and mentoring teams. * Experience operating AI/ML workloads, including MLOps/LLMOps, AI control plane, and production reliability considerations. * Experience with load testing, chaos engineering, capacity planning, autoscaling, and cloud/AI cost attribution tooling. * Software engineering fundamentals (OOP/OOD, data structures, algorithms) and AI-augmented development; familiarity with XP/Lean/DevSecOps/SRE tooling. Key requirements * Discretionary annual incentive program
Requirements
for production releases. * Foster cross-functional collaboration with engineering, security, data governance, and leadership to align on standards and controls. * Advance production engineering practices including chaos testing, capacity planning, and cloud/AI cost engineering. * Mentor and review peers to ensure reliability KPIs (availability, performance, cost) are met or exceeded. * Communicate complex technical concepts clearly to stakeholders and influence decision making. Tasks * 6+ years of software and site reliability engineering experience with large-scale, cloud-native systems. * Proficiency in Python, Go, Bash, Java, C#/.NET, SQL/NoSQL; strong Kubernetes, Terraform, ArgoCD experience; CI/CD and observability stack. * 3+ years owning SLIs, SLOs, SLAs, incident command, on-call, and production observability tools (OpenTelemetry, Prometheus, Grafana, Datadog, Dynatrace, CloudWatch, Azure Monitor, Google Cloud Operations). * 3+ years cloud-native engineering on Azure/AWS/GCP aaaaaaaw _ AI/ML services and container orchestration; IaC, networking, multi-environment management. * 1+ year establishing reliability standards (SLO discipline, runbooks, performance budgets) and mentoring teams. * Experience operating AI/ML workloads, including MLOps/LLMOps, AI control plane, and production reliability considerations. * Experience with load testing, chaos engineering, capacity planning, autoscaling, and cloud/AI cost attribution tooling. * Software engineering fundamentals (OOP/OOD, data structures, algorithms) and AI-augmented development; familiarity with XP/Lean/DevSecOps/SRE tooling. Key requirements * Discretionary annual incentive program
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
Dev Digest 120 - Apple and peers
Highest Paying Tech Companies for Developers
Dev Digest 121 - AI goes offline