> Markdown version of [/jobs/ext/1928116-staff-site-reliability-engineer-ai-platform](https://www.wearedevelopers.com/jobs/ext/1928116-staff-site-reliability-engineer-ai-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Site Reliability Engineer (Ai Platform) - **Company:** Manychat - **Location:** Barcelona, Spain (Remote available) - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Continuous Integration, Failover, Reliability Engineering, Prometheus, AI Infrastructure, Large Language Models, Grafana, Caching, AI Platforms, Kubernetes, Low Latency, Terraform - **Published:** August 5, 2026 - **Apply:** https://www.buscojobs.com.es/staff-site-reliability-engineer-ai-platform-en-barcelona-ID-366046101 ## About the Role * 5+ years in SRE / platform / infrastructure engineering with production ownership at scale * Hands-on experience with LLM-backed systems in production (Bedrock, OpenAI, Anthropic or similar) * Deep cloud-native background: AWS, Kubernetes, Terraform/IaC, CI/CD * Strong observability practice (Prometheus/Grafana, OpenTelemetry) and experience with SLOs for non-deterministic systems * Proven cost-optimization experience and ability to make costs visible * Staff-level influence and ability to drive technical direction beyond own team ## Description Experteer Overview In this role you will own the reliability, performance, and cost of Manychat's AI infrastructure.You'll shape the AI Platform from the ground up, partnering with the Head of Infrastructure to set standards and roadmaps.You'll work on scalable AI gateways, inference services, and multi-provider integration, ensuring low latency and strong observability.This is a mission-critical, impact-driven SRE position at a fast-growing AI-first company.Compensaciones / Beneficios* Own reliability and performance of AI infrastructure (AI Gateway, inference services, provider integrations)* Design and evolve the AI Gateway, including routing, failover, rate limiting, caching, and guardrails* Build observability for AI systems with latency/throughput/error SLOs, token metrics, drift signals* Drive FinOps for AI workloads: cost visibility, model right-sizing, caching strategies* Conduct capacity planning and incident response; develop runbooks and postmortems* Scale AI expertise across the org by setting standards and coaching teams on LLM-backed featuresResponsabilidades* 5+ years in SRE / platform / infrastructure engineering with production ownership at scale* Hands-on experience with LLM-backed systems in production (Bedrock, OpenAI, Anthropic or similar)* Deep cloud-native background: AWS, Kubernetes, Terraform/IaC, CI/CD* Strong observability practice (Prometheus/Grafana, OpenTelemetry) and experience with SLOs for non-deterministic systems* Proven cost-optimization experience and ability to make costs visible* Staff-level influence and ability to drive technical direction beyond own teamRequisitos principales* Hybrid onboarding with remote start* Relocation support for you and family* Comprehensive health insurance* Professional development budget* Flexible benefits package* Hybrid work and generous time off ## Related Videos - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [System Resilience: Surviving the Software Storm](https://www.wearedevelopers.com/videos/874-system-resilience-surviving-the-software-storm) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)