Staff Site Reliability Engineer (Ai Platform)

Manychat
Barcelona, Spain
about 2 months ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Continuous Integration Failover Reliability Engineering Prometheus AI Infrastructure Large Language Models Grafana Caching Rate Limiting AI Platforms
+3 more
Kubernetes Low Latency Terraform

Job description

Experteer Overview In this role you will own the reliability, performance, and cost of Manychat’s AI infrastructure.You’ll shape the AI Platform from the ground up, partnering with the Head of Infrastructure to set standards and roadmaps.You’ll work on scalable AI gateways, inference services, and multi-provider integration, ensuring low latency and strong observability.This is a mission-critical, impact-driven SRE position at a fast-growing AI-first company.Compensaciones / Beneficios* Own reliability and performance of AI infrastructure (AI Gateway, inference services, provider integrations)* Design and evolve the AI Gateway, including routing, failover, rate limiting, caching, and guardrails* Build observability for AI systems with latency/throughput/error SLOs, token metrics, drift signals* Drive FinOps for AI workloads: cost visibility, model right-sizing, caching strategies* Conduct capacity planning and incident response; develop runbooks and postmortems* Scale AI expertise across the org by setting standards and coaching teams on LLM-backed featuresResponsabilidades* 5+ years in SRE / platform / infrastructure engineering with production ownership at scale* Hands-on experience with LLM-backed systems in production (Bedrock, OpenAI, Anthropic or similar)* Deep cloud-native background: AWS, Kubernetes, Terraform/IaC, CI/CD* Strong observability practice (Prometheus/Grafana, OpenTelemetry) and experience with SLOs for non-deterministic systems* Proven cost-optimization experience and ability to make costs visible* Staff-level influence and ability to drive technical direction beyond own teamRequisitos principales* Hybrid onboarding with remote start* Relocation support for you and family* Comprehensive health insurance* Professional development budget* Flexible benefits package* Hybrid work and generous time off

Requirements

  • 5+ years in SRE / platform / infrastructure engineering with production ownership at scale
  • Hands-on experience with LLM-backed systems in production (Bedrock, OpenAI, Anthropic or similar)
  • Deep cloud-native background: AWS, Kubernetes, Terraform/IaC, CI/CD
  • Strong observability practice (Prometheus/Grafana, OpenTelemetry) and experience with SLOs for non-deterministic systems
  • Proven cost-optimization experience and ability to make costs visible
  • Staff-level influence and ability to drive technical direction beyond own team

Benefits & conditions

  • Hybrid onboarding with remote start
  • Relocation support for you and family
  • Comprehensive health insurance
  • Professional development budget
  • Flexible benefits package
  • Hybrid work and generous time off

About the company

Barcelona, España

Experteer Overview In this role you will own the reliability, performance, and cost of Manychat’s AI infrastructure. You’ll shape the AI Platform from the ground up, partnering with the Head of Infrastructure to set standards and roadmaps. You’ll work on scalable AI gateways, inference services, and multi-provider integration, ensuring low latency and strong observability. This is a mission-critical, impact-driven SRE position at a fast-growing AI-first company. Compensaciones / Beneficios

  • Own reliability and performance of AI infrastructure (AI Gateway, inference services, provider integrations)
  • Design and evolve the AI Gateway, including routing, failover, rate limiting, caching, and guardrails
  • Build observability for AI systems with latency/throughput/error SLOs, token metrics, drift signals
  • Drive FinOps for AI workloads: cost visibility, model right-sizing, caching strategies
  • Conduct capacity planning and incident response; develop runbooks and postmortems
  • Scale AI expertise across the org by setting standards and coaching teams on LLM-backed features

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

1:07 min

Architecting the availability stack with Prometheus and Grafana

Gabriel Labachelerie · World Congress 2023

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

6:45 min

Enabling rapid storage failover with third-party storage classes

Andrew Pruski · World Congress 2022

1:06 min

Empowering site reliability engineers with integrated AI agents

Osmar Matos Osmar Matos · World Congress 2026 Europe

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

Videos

See all

Related articles

See all