> Markdown version of [/jobs/ext/2953229-site-reliability-engineer-iaas](https://www.wearedevelopers.com/jobs/ext/2953229-site-reliability-engineer-iaas). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer, IaaS - **Company:** Algolia - **Location:** Paris, France (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Audit Trail, Cloud Computing, Linux, Distributed Systems, Python (Programming Language), Octopus Deploy, Software Tools, Kubernetes, Cloud Migration, Terraform, Golang - **Published:** September 17, 2026 - **Apply:** https://startup.jobs/senior-site-reliability-engineer-iaas-algolia-2-6137428 ## About the Role * Strong hands-on production expertise with AWS or GCP. * Deep practical Kubernetes knowledge, including its cloud infrastructure dependencies and operational challenges. * The ability to design, build, and operate reliable cloud infrastructure in production. * Strong infrastructure-as-code and automation skills, ideally Terraform and Python, Go, or equivalent. * Solid knowledge of Linux, networking, distributed systems, and production operations. * A platform mindset: you design for the engineers who consume what you build, balancing velocity, correctness, security, and reliability. * Awareness of cloud cost drivers and the ability to treat cost as an engineering constraint. * A track record of leading complex technical initiatives and delivering durable solutions across teams. * Comfort adopting AI-assisted engineering tools, with strong judgement for critical production systems. * Excellent spoken and written English skills., * Familiarity with more than one public cloud provider. * Knowledge of GitOps or policy-as-code tooling, such as Argo CD, Helm, OPA, or Kyverno. * Exposure to cloud migration, platform engineering, FinOps, or large-scale infrastructure transformation., * GRIT - Problem-solving and perseverance capability in an ever-changing and growing environment. * TRUST - Willingness to trust our co-workers and to take ownership. * CANDOR - Ability to receive and give constructive feedback. ## Description As a Senior Site Reliability Engineer in IaaS, you will help shape the next generation of Algolia's production infrastructure. You will lead major parts of the Cloud Baseline and the reliable lifecycle capabilities that enable teams to operate and migrate workloads safely on a cloud-native platform. You will work across cloud foundations, Kubernetes, platform engineering, automation, reliability, and large-scale production operations. This role is for an engineer who enjoys solving infrastructure problems where the answer must work not once, but hundreds or thousands of times: creating repeatable cloud environments, enabling a growing fleet of production clusters, reducing manual operations, and maintaining the reliability and cost efficiency our customers expect throughout the transition., * Lead the design and evolution of Cloud Baseline capabilities across cloud providers, including identity and access, networking, account structure, security, auditability, tagging, inventory, and cost visibility. * Design and automate cloud infrastructure foundations that enable a growing fleet of production Kubernetes clusters. * Lead complex infrastructure initiatives, such as cloud-environment standardisation, cluster lifecycle automation, upgrade strategies, or infrastructure-drift reduction. * Ensure cloud and Kubernetes foundations, lifecycle operations, and operational guardrails are reliable and scalable enough to support large-scale workload migration without compromising customer experience. * Treat the platform as a product: define clear interfaces, reusable modules, self-service workflows, documentation, and reliable operational standards for the engineers who consume it. * Build automated guardrails for security, compliance, reliability, and safe change management, allowing teams to move faster without weakening production protections. * Improve platform efficiency through capacity planning, rightsizing, autoscaling, resource governance, and clear cost visibility. * Use automation and AI-assisted engineering tools where appropriate to improve fleet-scale analysis, infrastructure documentation, and safe, repeatable operational changes. * Mentor engineers, share knowledge, and raise the quality of infrastructure design and operations across the team. * Collaborate with Infrastructure, Security, FinOps, and engineering teams to align technical decisions and deliver high-impact platform capabilities. * Participate in the on-call rotation and lead the resolution of complex production issues. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship)