Lead Applied AI Site Reliability Engineer II - PxE ERM
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+9 more
Job description
Experteer Overview In this Lead Applied AI SRE role, you ensure reliability, performance, and cost efficiency for high-visibility products and AI-infused workloads. You will lead by example, mentoring teams and setting production standards across cross-functional partners. You’ll design observability, SLOs, and automated reliability checks to prevent incidents and enable safe, scalable releases. This role blends cloud platform engineering with applied AI fluency to optimize operations and drive business value. You will partner with engineering leadership to shape resilient systems and codify best practices that scale. Compensation / Benefits * Drive reliability, performance, and cost outcomes using SLOs and error budgets; prioritize toil and incidents to improve production resilience * Act as technical advocate for production reliability; set standards, design observability and resilience tooling, and gate production readiness * Own operational integrity of production and pre-production environments; manage SLOs, alerting, drift prevention, and security-conscious control measures * Lead development of runbooks, playbooks, postmortems, and automation; mentor peers to meet reliability KPIs * Collaborate with cross-functional teams to define and satisfy service-level objectives and production readiness * Apply advanced SRE practices to cloud-native, AI-infused workloads, ensuring safe degradation and controlled risk Tasks * Bachelor degree in CS, software engineering, data science, ML, or related field * 6+ years software engineering and SRE experience with large-scale cloud-native systems * 3+ years SRE/production engineering focusing on SLIs/SLOs/SLAs; incident response; observability stacks * 3+ years cloud-native engineering on Azure, AWS, or GCP; container orchestration; IaC; networking; multi-environment mgmt * 1+ year establishing reliability standards (SLO discipline, runbooks, budgets) and mentoring teams * Experience operating AI/ML and agentic workloads; familiarity with MLOps/LLMOps and AI control planes * Experience with load/performance testing, chaos engineering, capacity planning, autoscaling, and FinOps tooling * Software engineering background with understanding of diagrams, data structures, algorithms, and AI-driven development * Experience with XP/Lean/DevSecOps/SRE/CI tooling (GitHub, SonarQube, MLflow) and agentic AI frameworks Key requirements *
Requirements
environments; manage SLOs, alerting, drift prevention, and security-conscious control measures * Lead development of runbooks, playbooks, postmortems, and automation; mentor peers to meet reliability KPIs * Collaborate with cross-functional teams to define and satisfy service-level objectives and production readiness * Apply advanced SRE practices to cloud-native, AI-infused workloads, ensuring safe degradation and controlled risk Tasks * Bachelor degree in CS, software engineering, data science, ML, or related field * 6+ years software engineering and SRE experience with large-scale cloud-native systems * 3+ years SRE/production engineering focusing on SLIs/SLOs/SLAs; incident response; observability stacks * 3+ years cloud-native engineering on Azure, AWS, or GCP; container orchestration; IaC; networking; multi-environment mgmt * 1+ year establishing reliability standards (SLO discipline, runbooks, budgets) and mentoring teams * Experience operating AI/ML and agentic workloads; aaaaaaaaat _ with MLOps/LLMOps and AI control planes * Experience with load/performance testing, chaos engineering, capacity planning, autoscaling, and FinOps tooling * Software engineering background with understanding of diagrams, data structures, algorithms, and AI-driven development * Experience with XP/Lean/DevSecOps/SRE/CI tooling (GitHub, SonarQube, MLflow) and agentic AI frameworks Key requirements *
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 120 - Apple and peers
Navigating the AI Shift
Is Software Engineering Over-Saturated?
MLOps And AI Driven Development