Lead Applied AI Site Reliability Engineer II - PxE A&A
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+14 more
Job description
Experteer Overview In this role you drive reliability, performance, and cost efficiency for high-visibility products and AI workloads across cloud environments. You will lead production standards, design observability and resilience tooling, and ensure safe, scalable production releases. You mentor teams, partner with cross-functional groups, and push pragmatic improvements that reduce toil and incidents. This position combines SRE craft with applied AI fluency to operate AI-enabled platforms at scale. You will influence both technology and processes to deliver dependable, cost-aware systems. Compensation / Benefits * Drive reliability and cost outcomes using SLOs and error budgets * Set production standards and design observability, performance, and resilience testing * Own admission of systems into production with readiness checks and automated reliability gates * Develop runbooks, playbooks, and automation; lead blameless postmortems * Collaborate with cross-functional teams to co-define SLIs/SLOs and ensure readiness * Lead and mentor engineers in reliability practices and continuous improvement * Execute chaos testing, capacity planning, and cost engineering to prove readiness * Maintain hands-on ownership of production and pre-production environments to prevent drift * Communicate technically complex concepts clearly to diverse stakeholders * Foster a collaborative, cross-functional culture focused on quality and trust Tasks * 6+ years of software engineering and SRE experience operating large-scale, cloud-native systems * 3+ years defining/owning SLIs, SLOs, SLAs, and incident management including on-call * 3+ years of cloud-native engineering across Azure/AWS/GCP and Kubernetes, containers, infrastructure-as-code * Experience with CI/CD, observability stacks and production-grade monitoring (Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, CloudWatch, etc.) * Experience with AI/ML workloads in production, MLOps, and AI control planes (guardrails, model gateways) * Familiar with load testing, chaos engineering, autoscaling, FinOps tooling and GPU/token cost attribution * Strong coding and scripting in languages such as Python, Go, Bash; understanding of OOP/OOD and data structures * Knowledge of security/risk controls including least-privilege, RBAC, secrets management * Excellent communication and leadership abilities; comfortable mentoring and pair programming Key requirements *
Requirements
years SLIs/SLOs and ensure readiness * Lead and mentor engineers in reliability practices and continuous improvement * Execute chaos testing, capacity planning, and cost engineering to prove readiness * Maintain hands-on ownership of production and pre-production environments to prevent drift * Communicate technically complex concepts clearly to diverse stakeholders * Foster a collaborative, cross-functional culture focused on quality and trust Tasks * 6+ years of software engineering and SRE experience operating large-scale, cloud-native systems * 3+ years defining/owning SLIs, SLOs, SLAs, and incident management including on-call * 3+ years of cloud-native engineering across Azure/AWS/GCP and Kubernetes, containers, infrastructure-as-code * Experience with CI/CD, observability stacks and production-grade monitoring (Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, CloudWatch, etc.) * Experience with AI/ML workloads in production, MLOps, and AI control planes (guardrails, model aa secrets * Familiar with load testing, chaos engineering, autoscaling, FinOps tooling and GPU/token cost attribution * Strong coding and scripting in languages such as Python, Go, Bash; understanding of OOP/OOD and data structures * Knowledge of security/risk controls including least-privilege, RBAC, secrets management * Excellent communication and leadership abilities; comfortable mentoring and pair programming Key requirements *
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
Navigating the AI Shift
MLOps And AI Driven Development
Highest Paying Tech Companies for Developers