Site Reliability Engineer, AI Platform

Algolia
Paris, France
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
€69,768.0 - €96,900.0
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Software Debugging Distributed Systems Python (Programming Language) Reliability Engineering AI Platforms Kubernetes Deployment Automation Machine Learning Operations

Job description

  • Build and operate production infrastructure supporting AI-related workloads and services
  • Operate and improve highly available Kubernetes-based platforms
  • Improve reliability through SLOs, observability, alerting and capacity management
  • Investigate production issues and turn findings into durable fixes and improvements
  • Work across networking, databases, compute and service infrastructure
  • Improve CI/CD pipelines, deployment automation and developer experience
  • Build and maintain infrastructure using Infrastructure as Code
  • Participate in on-call, incident response and operational improvements
  • Collaborate with experienced engineers across AI Platform and progressively take ownership of broader production areas

Requirements

We are looking for a Site Reliability Engineer with strong production fundamentals who enjoys solving operational problems, automating repetitive work and progressively taking ownership of complex systems at scale., * Solid hands-on Kubernetes knowledge, including workloads, resource management, and production operations

  • Strong experience with Infrastructure as Code, and the lifecycle of cloud infrastructure
  • Solid experience building and operating CI/CD pipelines and automated deployment workflows
  • Hands-on experience with at least one major cloud provider: GCP, AWS or Azure
  • Good understanding of networking, distributed systems and reliability engineering
  • Experience with monitoring, observability and troubleshooting production systems
  • Strong automation mindset and the ability to take ownership of well-defined production systems and progressively tackle more complex problems
  • Excellent written and spoken English

NICE TO HAVE:

  • Go and/or Python engineering experience
  • Exposure to AI/ML infrastructure and inferences
  • Comfortable working AI-first, using coding agents, agentic workflows and AI-assisted debugging to accelerate engineering and operations, * GRIT - Problem-solving and perseverance capability in an ever-changing and growing environment.
  • TRUST - Willingness to trust our co-workers and to take ownership.
  • CANDOR - Ability to receive and give constructive feedback.

Benefits & conditions

The annual base salary compensation range for this role reflects market pay data within this location. The exact compensation offered for this role may vary depending on specific location and job-related knowledge, technical skills, and experience; and is only one part of our Total Rewards philosophy to compensate and recognize employees for their work. Base Salary Pay Range €69.768-€96.900 EUR

About the company

AI Platform builds and operates the shared production foundations supporting Algolia’s evolving AI ecosystem.

The team works at the intersection of Site Reliability Engineering, cloud infrastructure, software engineering and AI, helping engineering teams bring AI-powered capabilities to production reliably, securely and efficiently. Our scope includes Kubernetes, cloud infrastructure, CI/CD, networking, databases, observability, reliability, FinOps and production operations.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

2:36 min

Choosing between managed AI platforms and custom governance

Péter Farkas Péter Farkas · Europe 2026 Virtual

1:06 min

Outline of free tools for Microsoft Azure

Radu Vunvulea Radu Vunvulea · World Congress 2022

1:06 min

Empowering site reliability engineers with integrated AI agents

Osmar Matos Osmar Matos · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · World Congress 2022

Videos

See all

Related articles

See all