Site Reliability Engineer

Codeway
Barcelona, Spain
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior

Job location

Barcelona, Spain

Tech stack

Cloud Engineering
Computer Networks
Continuous Integration
Linux
DevOps
Disaster Recovery
Mobile Application Software
Key Management
Networking Basics
Role-Based Access Control
Reliability Engineering
Software Vulnerability Management
Alwayson
Google Cloud Platform
Autoscaling
Istio
Containerization
Kubernetes
Cloudflare
Cloud Optimization
Api Gateway
Terraform
Devsecops

Job description

We're looking for a Senior Site Reliability Engineer to own and mature reliability, performance, and security across our growing platform. This role sits at the intersection of Engineering, Infrastructure, and Security, helping design, operate, and continuously improve the systems that keep dozens of consumer apps running for users around the world., * Define, instrument, and report on SLIs, SLOs, and error budgets across critical services, so reliability decisions are driven by data rather than opinion.

  • Own observability end-to-end - metrics, logs, traces, dashboards, and alerting - and drive measurable reductions in detection and resolution times.
  • Reduce alert noise and false positives so on-call engineers can trust what wakes them up.
  • Run reliability reviews and an error-budget policy that shapes how teams prioritize between shipping and stability.

Kubernetes & Platform Operations

  • Operate, scale, and upgrade our multi-cluster Kubernetes (GKE) environment: cluster lifecycle, autoscaling, networking, ingress, and resource management.
  • Act as the deep-expertise escalation point for cluster and platform issues across dozens of services.
  • Own capacity planning, performance, and cloud cost efficiency, balancing spend against reliability targets.
  • Build self-service platform tooling that lets product teams move quickly without needing to become infrastructure experts.

Security & Resilience

  • Embed security into the platform through RBAC and least-privilege, secrets management, image and dependency scanning, network policies, and a disciplined patching cadence.
  • Partner with the security function on vulnerability remediation, audit readiness, and secure-by-default infrastructure.
  • Own disaster recovery: define and regularly validate RTO/RPO targets through DR drills and failure testing.
  • Contribute to architecture and production-readiness reviews so reliability and security are designed in, not bolted on.

Incident Response & Automation

  • Lead the on-call rotation and act as incident commander during production incidents.
  • Run blameless post-mortems, quantify impact, and track corrective actions through to closure so the same failure doesn't recur.
  • Build and maintain Infrastructure as Code (Terraform) and CI/CD pipelines, enforcing GitOps and progressive delivery with automated rollbacks.
  • Systematically identify, measure, and eliminate operational toil through automation, protecting engineering time for high-leverage work., * Kubernetes (GKE) and containerized workloads
  • Google Cloud Platform (GCP)
  • Terraform and Infrastructure as Code
  • CI/CD and GitOps tooling
  • Modern observability (metrics, logs, traces, alerting)
  • Cloudflare CDN and edge

What Success Looks Like

  • Clear SLIs, SLOs, and error budgets live and reported for our most critical services.
  • A measurable reduction in detection and resolution times for production incidents.
  • A consistent, blameless incident response practice, with postmortem actions tracked to closure.
  • A hardened Kubernetes fleet, with key security gaps closed across access control, secrets, scanning, and patching.
  • Expanded Infrastructure as Code and observability coverage across the platform.
  • Disaster recovery drills that reliably pass agreed RTO/RPO targets.
  • Operational toil trending down against an explicit target, with automation replacing manual work.
  • Product teams self-serving on standardized, secure, well-instrumented platform tooling.
  • Shared measurement with the Security team, so security outcomes are tracked with the same rigor as reliability ones: vulnerability remediation times, patching cadence and coverage, access review completion, and security incident detection and response times - reported as real metrics, not status updates.
  • A single, agreed view of operational health across engineering and security, so the two functions work from the same definitions and the same dashboards rather than competing narratives.
  • Visibility into how efficiently we operate as an organization, connecting platform performance to the business it supports - the reliability of the systems our marketing and growth stack depends on, the speed and safety of our path from commit to production, and the infrastructure cost behind serving and acquiring users.
  • Metrics leadership actually uses, giving the wider company a clear, honest picture of where we're strong, where we're exposed, and where the next investment should go.

Requirements

  • Experience operating high-traffic, always-on production systems at meaningful scale, typically gained over 5-8 years in SRE, Platform, or DevOps roles.
  • Hands-on production Kubernetes experience - you've run clusters day to day, through upgrades, autoscaling, and real troubleshooting under load, not just deployed to them.
  • A strong cloud engineering background, along with solid Linux and networking fundamentals.
  • A track record of defining and operating with SLOs and error budgets, and comfort being measured on reliability outcomes.
  • Experience with Infrastructure as Code and CI/CD pipeline design - you treat infrastructure and delivery as code.
  • Depth in observability tooling: instrumentation, dashboarding, and alert design.
  • A genuine security-first mindset, where least-privilege, secrets hygiene, and vulnerability management are habits rather than afterthoughts.
  • Scripting and automation fluency in at least one language, used to build tooling and remove toil.
  • Incident-command experience: owning on-call, running blameless post-mortems, and driving resolution times down over time.
  • Ability to communicate clearly with both engineers and leadership, especially under pressure., * Experience with high-scale consumer or mobile app backends, or with AI/ML inference workloads and their scaling characteristics.
  • Experience with GitOps and progressive-delivery patterns such as canary and blue-green rollouts.
  • Familiarity with service mesh, API gateways, or multi-region and multi-cluster topologies.
  • Cloud cost optimization and FinOps discipline at scale.
  • Exposure to compliance initiatives (SOC 2, ISO 27001, GDPR) and broader DevSecOps practice.
  • Chaos engineering or resilience testing experience.
  • Relevant certifications in Kubernetes, cloud, or DevOps disciplines.
  • Experience supporting many independent services and teams concurrently in a fast-shipping, product-led environment.

Benefits & conditions

  • A Competitive Compensation Package. Long story short, we take care of you.
  • A Meal Compensation that covers a decent and nutritious lunch.
  • Full health benefits, including unlimited private health insurance and HPV vaccine coverage.
  • Pet adoption support covering primary healthcare, parasite vaccinations, and microchip costs in the first year after adoption.
  • State-of-the-art tech: MacBook, iPhone 15 Pro, Magic Mouse, Magic Keyboard, adjustable desk, 4K screen, and any other gadget you need.
  • Sport activities support and gym membership.
  • Flexible schedule that encourages responsibility over tracking every minute.
  • English language course support.
  • Top-notch office in the heart of Barcelona at the iconic Edifici Estel.
  • Codebrew coffee shop with healthy snacks at all hours and free breakfast & lunch daily.
  • No dress code.
  • A dynamic team environment with young, talented, and passionate members.
  • Gaming area with PS5 corner.
  • Software support subscriptions.
  • Public transportation support with additional monthly compensation.

Recruiting Process

  • Application: Send us your CV or LinkedIn profile. You can also write a few words about yourself.
  • Hiring Manager Interview.
  • Talent & Culture Interview.
  • Case Study.
  • Technical Interview.
  • Welcome aboard! You are now part of the team.

About the company

Codeway is a global consumer tech company with more than 400 million users worldwide. Since 2020, we've built and scaled 60+ mobile apps across creativity, productivity, wellness, language learning, and entertainment. Our flagship apps - Retake AI, Cleanup, Learna, and DramaPops - lead their categories globally. In 2024, we became the most downloaded app publisher on iOS. We're a team of 300+ people across İstanbul and Barcelona who bring curiosity, passion, trust, and ownership to everything we build. Recognized as a #1 LinkedIn Top Startup and a Great Place to Work in Europe, Codeway is where ambitious people do their life's best work. We're building the next generation of consumer tech and reimagining what mobile apps can be. This is Codeway. This is our way. Join us.

Apply for this position