> Markdown version of [/jobs/ext/1878506-classical-sre](https://www.wearedevelopers.com/jobs/ext/1878506-classical-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # classical SRE - **Company:** Manychat Runs Ai - **Location:** Barcelona, Spain (Remote available) - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Continuous Integration, Failover, Prometheus, Delivery Pipeline, Large Language Models, Grafana, Caching, AI Platforms, Kubernetes, Machine Learning Operations, Terraform - **Published:** August 1, 2026 - **Apply:** https://www.jobleads.com/es/job/eafa802a9981872070f121e22d6a7a279 ## About the Role * 5+ years in SRE / platform / infrastructure engineering, including production ownership at significant scale. * Hands-on experience operating LLM-backed systems in production: provider APIs (Bedrock, OpenAI, Anthropic, or similar), inference pipelines, self-hosted or managed model serving. * Deep cloud-native background: AWS, Kubernetes, Terraform/IaC, CI/CD. * Strong observability practice (Prometheus/Grafana, OpenTelemetry, or equivalent) and experience defining SLOs for non-deterministic systems. * Proven cost-optimization work: you can show where you cut cloud or inference spend and how you made cost visible. * Staff-level influence: you've set technical direction beyond your own team and brought others along, * Experience building or operating an LLM gateway/proxy (e.g., LiteLLM, Kong AI Gateway, custom). * Experience with GPU workload optimization, quantization, or serving frameworks (vLLM, TGI, Triton). * Experience with eval pipelines and quality monitoring for LLM outputs. ## Description * Own reliability and performance of our AI infrastructure: AI Gateway, inference services, and integrations with Amazon Bedrock, Azure OpenAI, and other LLM providers. * Design and evolve the AI Gateway: routing, failover between providers, rate limiting, caching, and guardrails. * Build observability for AI systems: latency/throughput/error SLOs per model and provider, token-level metrics, quality and drift signals. * Drive cost optimization and FinOps for AI workloads: per-feature cost visibility, model right-sizing, caching strategies, provider mix. * Run capacity planning and incident response for inference services; write and improve runbooks and postmortems. * Scale AI expertise across the org: set standards, review designs, and coach teams shipping LLM-backed features., * Green-field ownership: the AI Platform is young; you'll shape its architecture, standards, and roadmap. * Real scale and real stakes: AI features sit in the critical path of customer-facing automation. * Direct partnership with the Head of Infrastructure; high autonomy and visibility. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [System Resilience: Surviving the Software Storm](https://www.wearedevelopers.com/videos/874-system-resilience-surviving-the-software-storm) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Building Sovereign AI: Lessons from Deploying Secure RAG Systems using Confidential Computing](https://www.wearedevelopers.com/videos/100108-building-sovereign-ai-lessons-from-deploying-secure-rag-systems-using-confidential-computing) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Dev Digest 138 - Are you secure about this?](https://www.wearedevelopers.com/magazine/486-dev-digest-138-are-you-secure-about-this)