Staff Site Reliability Engineer (Ai Platform)

Manychat
Barcelona, Spain
8 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Continuous Integration Failover Reliability Engineering Prometheus AI Infrastructure Large Language Models Grafana Caching AI Platforms Kubernetes
+2 more
Low Latency Terraform

Job description

Experteer Overview In this role you will own the reliability, performance, and cost of Manychat’s AI infrastructure.You’ll shape the AI Platform from the ground up, partnering with the Head of Infrastructure to set standards and roadmaps.You’ll work on scalable AI gateways, inference services, and multi-provider integration, ensuring low latency and strong observability.This is a mission-critical, impact-driven SRE position at a fast-growing AI-first company.Compensaciones / Beneficios* Own reliability and performance of AI infrastructure (AI Gateway, inference services, provider integrations)* Design and evolve the AI Gateway, including routing, failover, rate limiting, caching, and guardrails* Build observability for AI systems with latency/throughput/error SLOs, token metrics, drift signals* Drive FinOps for AI workloads: cost visibility, model right-sizing, caching strategies* Conduct capacity planning and incident response; develop runbooks and postmortems* Scale AI expertise across the org by setting standards and coaching teams on LLM-backed featuresResponsabilidades* 5+ years in SRE / platform / infrastructure engineering with production ownership at scale* Hands-on experience with LLM-backed systems in production (Bedrock, OpenAI, Anthropic or similar)* Deep cloud-native background: AWS, Kubernetes, Terraform/IaC, CI/CD* Strong observability practice (Prometheus/Grafana, OpenTelemetry) and experience with SLOs for non-deterministic systems* Proven cost-optimization experience and ability to make costs visible* Staff-level influence and ability to drive technical direction beyond own teamRequisitos principales* Hybrid onboarding with remote start* Relocation support for you and family* Comprehensive health insurance* Professional development budget* Flexible benefits package* Hybrid work and generous time off

Requirements

  • 5+ years in SRE / platform / infrastructure engineering with production ownership at scale
  • Hands-on experience with LLM-backed systems in production (Bedrock, OpenAI, Anthropic or similar)
  • Deep cloud-native background: AWS, Kubernetes, Terraform/IaC, CI/CD
  • Strong observability practice (Prometheus/Grafana, OpenTelemetry) and experience with SLOs for non-deterministic systems
  • Proven cost-optimization experience and ability to make costs visible
  • Staff-level influence and ability to drive technical direction beyond own team

Benefits & conditions

  • Hybrid onboarding with remote start
  • Relocation support for you and family
  • Comprehensive health insurance
  • Professional development budget
  • Flexible benefits package
  • Hybrid work and generous time off

About the company

Barcelona, España

Experteer Overview In this role you will own the reliability, performance, and cost of Manychat’s AI infrastructure. You’ll shape the AI Platform from the ground up, partnering with the Head of Infrastructure to set standards and roadmaps. You’ll work on scalable AI gateways, inference services, and multi-provider integration, ensuring low latency and strong observability. This is a mission-critical, impact-driven SRE position at a fast-growing AI-first company. Compensaciones / Beneficios

  • Own reliability and performance of AI infrastructure (AI Gateway, inference services, provider integrations)
  • Design and evolve the AI Gateway, including routing, failover, rate limiting, caching, and guardrails
  • Build observability for AI systems with latency/throughput/error SLOs, token metrics, drift signals
  • Drive FinOps for AI workloads: cost visibility, model right-sizing, caching strategies
  • Conduct capacity planning and incident response; develop runbooks and postmortems
  • Scale AI expertise across the org by setting standards and coaching teams on LLM-backed features

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:06 min

Empowering site reliability engineers with integrated AI agents

Osmar Matos Osmar Matos · World Congress 2026 Europe

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

2:59 min

Scaling clusters and handling automated replica failover

Jürgen Pilz · World Congress 2023

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

Videos

See all

Related articles

See all