Staff Site Reliability Engineer (Ai Platform)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+2 more
Job description
Experteer Overview In this role you will own the reliability, performance, and cost of Manychat’s AI infrastructure.You’ll shape the AI Platform from the ground up, partnering with the Head of Infrastructure to set standards and roadmaps.You’ll work on scalable AI gateways, inference services, and multi-provider integration, ensuring low latency and strong observability.This is a mission-critical, impact-driven SRE position at a fast-growing AI-first company.Compensaciones / Beneficios* Own reliability and performance of AI infrastructure (AI Gateway, inference services, provider integrations)* Design and evolve the AI Gateway, including routing, failover, rate limiting, caching, and guardrails* Build observability for AI systems with latency/throughput/error SLOs, token metrics, drift signals* Drive FinOps for AI workloads: cost visibility, model right-sizing, caching strategies* Conduct capacity planning and incident response; develop runbooks and postmortems* Scale AI expertise across the org by setting standards and coaching teams on LLM-backed featuresResponsabilidades* 5+ years in SRE / platform / infrastructure engineering with production ownership at scale* Hands-on experience with LLM-backed systems in production (Bedrock, OpenAI, Anthropic or similar)* Deep cloud-native background: AWS, Kubernetes, Terraform/IaC, CI/CD* Strong observability practice (Prometheus/Grafana, OpenTelemetry) and experience with SLOs for non-deterministic systems* Proven cost-optimization experience and ability to make costs visible* Staff-level influence and ability to drive technical direction beyond own teamRequisitos principales* Hybrid onboarding with remote start* Relocation support for you and family* Comprehensive health insurance* Professional development budget* Flexible benefits package* Hybrid work and generous time off
Requirements
- 5+ years in SRE / platform / infrastructure engineering with production ownership at scale
- Hands-on experience with LLM-backed systems in production (Bedrock, OpenAI, Anthropic or similar)
- Deep cloud-native background: AWS, Kubernetes, Terraform/IaC, CI/CD
- Strong observability practice (Prometheus/Grafana, OpenTelemetry) and experience with SLOs for non-deterministic systems
- Proven cost-optimization experience and ability to make costs visible
- Staff-level influence and ability to drive technical direction beyond own team
Benefits & conditions
- Hybrid onboarding with remote start
- Relocation support for you and family
- Comprehensive health insurance
- Professional development budget
- Flexible benefits package
- Hybrid work and generous time off
About the company
Barcelona, España
Experteer Overview In this role you will own the reliability, performance, and cost of Manychat’s AI infrastructure. You’ll shape the AI Platform from the ground up, partnering with the Head of Infrastructure to set standards and roadmaps. You’ll work on scalable AI gateways, inference services, and multi-provider integration, ensuring low latency and strong observability. This is a mission-critical, impact-driven SRE position at a fast-growing AI-first company. Compensaciones / Beneficios
- Own reliability and performance of AI infrastructure (AI Gateway, inference services, provider integrations)
- Design and evolve the AI Gateway, including routing, failover, rate limiting, caching, and guardrails
- Build observability for AI systems with latency/throughput/error SLOs, token metrics, drift signals
- Drive FinOps for AI workloads: cost visibility, model right-sizing, caching strategies
- Conduct capacity planning and incident response; develop runbooks and postmortems
- Scale AI expertise across the org by setting standards and coaching teams on LLM-backed features
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.buscojobs.com.esGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 121 - AI goes offline
Navigating the AI Shift
Dev Digest 137 - AI'm not sure about this
Dev Digest 132 - Binging WADFlix?