Site Reliability Engineer, IaaS

Algolia
Paris, France
4 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Artificial Intelligence Amazon Web Services Audit Trail Cloud Computing Computer Programming Linux Python (Programming Language) Networking Basics Octopus Deploy Software Tools Scripting Cloud Platform System
+4 more
Git Flow Kubernetes Cloud Migration Terraform

Job description

As a Site Reliability Engineer in IaaS, you will help build the next generation of Algolia’s production infrastructure.

You will contribute to the cloud baseline and reliable lifecycle capabilities that enable teams to operate and migrate workloads safely on a cloud-native platform. You will work across cloud foundations, Kubernetes, automation, reliability, and large-scale production operations.

As a P3 engineer, you will be a hands-on contributor. You will build, operate, and improve production infrastructure while developing deep expertise in cloud, Kubernetes, reliability, and automation., * Build and improve Cloud Baseline capabilities, including identity and access, networking, security, resource inventory, tagging, and auditability.

  • Develop and maintain infrastructure as code and automation for cloud environments and Kubernetes infrastructure.
  • Contribute to reliable, repeatable cloud and cluster lifecycle operations.
  • Help build self-service capabilities, reusable modules, and clear documentation that make the safe path the easy path for platform consumers.
  • Reduce manual work and configuration drift through automation, testing, GitOps practices, and standardisation.
  • Use automation and AI-assisted engineering tools where appropriate to improve infrastructure analysis, documentation, and safe, repeatable changes.
  • Improve observability, monitoring, alerting, capacity management, and operational documentation.
  • Investigate production issues, participate in the on-call rotation, and turn lessons learned into lasting improvements.
  • Work with Infrastructure, Security, FinOps, and engineering teams to deliver reliable, secure, and cost-aware platform capabilities.

Requirements

  • Hands-on production knowledge of AWS or GCP.
  • Practical Kubernetes knowledge and an interest in operating it in production.
  • Familiarity with infrastructure as code, ideally Terraform.
  • Programming or scripting skills in Python, Go, or an equivalent language.
  • Strong Linux and networking fundamentals.
  • A strong interest in reliability, automation, and solving production problems.
  • Comfort adopting AI-assisted engineering tools, with sound judgement for critical production systems.
  • The ability to communicate clearly and work effectively with a distributed team.
  • Excellent spoken and written English skills., * Familiarity with more than one public cloud provider.
  • Knowledge of GitOps or policy-as-code tooling, such as Argo CD, Helm, OPA, or Kyverno.
  • Experience with cloud migration, platform engineering, or large-scale infrastructure transformation., * GRIT - Problem-solving and perseverance capability in an ever-changing and growing environment.
  • TRUST - Willingness to trust our co-workers and to take ownership.
  • CANDOR - Ability to receive and give constructive feedback.

About the company

The Infrastructure as a Service team is at the center of one of Algolia’s most consequential engineering transformations.

For years, Algolia has operated a production fleet of approximately 4,000 bare-metal servers to deliver the reliability, low latency, and scalability that our customers expect. We are now building the foundations of a unified cloud and Kubernetes platform designed to support Algolia’s growth for years to come.

. We are now building the foundations of a unified cloud and Kubernetes platform designed to support Algolia’s growth for years to come.

This is not a lift-and-shift project. It is an opportunity to rethink how Algolia provisions, secures, operates, observes, upgrades, and scales production infrastructure and to build it as a platform that engineers can safely consume, rather than a queue of manual requests.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

5:02 min

Mapping Git flow branches to application tester segments

Majid Hajian · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:19 min

Applying code assistant capabilities to infrastructure and cloud operations

Ryan J Salva · Coffee With Developers

1:53 min

Evaluating traditional scripting languages for modern development tasks

Jens Knipper Jens Knipper · Europe 2026 Virtual

Videos

See all

Related articles

See all