Site Reliability Engineer, AI Platform

Algolia
Paris, France
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
€69,768.0 - €96,900.0
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Artificial Intelligence Amazon Web Services Cloud Engineering Software Debugging Distributed Systems Python (Programming Language) Reliability Engineering Graphics Processing Unit (GPU) Kubernetes Machine Learning Operations

Job description

  • Own and evolve production infrastructure supporting AI-related workloads and services at scale
  • Design and operate highly available Kubernetes-based platforms
  • Drive reliability through SLOs, observability, capacity planning and production guardrails
  • Lead complex production investigations and turn findings into durable architectural improvements
  • Improve shared infrastructure across networking, databases, service communication and compute
  • Build better CI/CD, progressive delivery, automation and developer experience
  • Drive cloud infrastructure efficiency and FinOps initiatives
  • Participate in and improve on-call and incident response
  • Mentor engineers and raise the technical bar for reliability and production engineering

Requirements

We are looking for a Senior Site Reliability Engineer who can independently own complex production systems, drive technical decisions across teams, and help shape reliable and efficient infrastructure at scale., * Strong hands-on production experience with at least one major cloud provider: GCP, AWS or Azure

  • Strong experience designing and operating Kubernetes and cloud-native production systems at scale
  • Strong understanding of distributed systems, networking and reliability engineering
  • Experience operating business-critical systems with strong availability, scalability and operational requirements
  • Ability to independently own ambiguous, cross-team technical problems and drive them to measurable outcomes
  • Strong automation mindset and ability to balance reliability, engineering velocity and cost
  • Excellent written and spoken English

NICE TO HAVE:

  • Go and/or Python engineering experience
  • Experience with infrastructure supporting AI/ML workloads, model serving, GPUs or other compute-intensive systems
  • Comfortable working AI-first, using coding agents, agentic development workflows, AI-assisted debugging and automation to accelerate engineering and operations, * GRIT - Problem-solving and perseverance capability in an ever-changing and growing environment.
  • TRUST - Willingness to trust our co-workers and to take ownership.
  • CANDOR - Ability to receive and give constructive feedback.

Benefits & conditions

The annual base salary compensation range for this role reflects market pay data within this location. The exact compensation offered for this role may vary depending on specific location and job-related knowledge, technical skills, and experience; and is only one part of our Total Rewards philosophy to compensate and recognize employees for their work. Base Salary Pay Range €69.768-€96.900 EUR

About the company

AI Platform builds and operates the shared production foundations supporting Algolia’s evolving AI ecosystem.

The team works at the intersection of Site Reliability Engineering, cloud infrastructure, software engineering and AI, helping engineering teams bring AI-powered capabilities to production reliably, securely and efficiently. Our scope includes Kubernetes, cloud infrastructure, CI/CD, networking, databases, observability, reliability, FinOps and production operations.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application