> Markdown version of [/jobs/ext/3042176-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3042176-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** HeadFirst - **Location:** Hoofddorp, Netherlands - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, Microsoft Azure, Cloud Computing, Cloud Engineering, Data Architecture, Github, Log Analysis, Reliability Engineering, Software Engineering, Data Logging, Cloud Platform System, Cloud Monitoring, Grafana, Reliability of Systems, Kubernetes, Build Tools, Terraform, Databricks - **Published:** September 24, 2026 - **Apply:** https://nl.engineering.jobs/nl/vacature/senior-site-reliability-engineer-7946939 ## About the Role Engineering, Reliability, Site ## Description **Keep our global cloud platform reliable, scalable, and resilient.** Do you enjoy solving complex production challenges before they turn into incidents? Are you the kind of engineer who automates repetitive tasks, improves reliability through engineering, and believes that every outage is an opportunity to build a better system? And are you an experienced, hands-on engineer who wants to stay close to the technology rather than move into people management? We're looking for a Senior Site Reliability Engineer to join the Global Platform Team at HeadFirst x Impellam. In this role, you'll help build and operate the cloud platform that powers our global Workforce-as-a-Service (WaaS) ecosystem. You'll work alongside Cloud Engineers, Platform Engineers, Data Engineers, and AI specialists to improve platform resilience, reduce operational overhead, and ensure that our Azure-based engineering environment remains reliable, scalable, and resilient. This is a hands-on individual contributor role, with no people management responsibilities. **Your impact** As a Senior Site Reliability Engineer, your focus is to keep our platforms healthy, reliable, and easy to operate while continuously improving how we build, run and recover production systems. You'll help build the engineering foundations behind our Headless Data Architecture (HDA), which runs on Azure and Databricks, as well as the Custom Apps Infrastructure (CA) that powers integrations, internal applications, and operational workflows across our international organization. Instead of spending your days reacting to incidents, you'll focus on preventing them through automation, observability, and reliability engineering. You'll reduce operational burden, improve platform resilience, and build systems that scale, recover automatically whenever possible, and provide engineering teams across Cloud, Data, and AI with the visibility they need to run production workloads with confidence. **What you will do** * Improve the reliability, availability, and performance of our Azure platform and production environments. * Build and improve monitoring, logging, and alerting using Grafana, OpenTelemetry, Azure Monitor, and Log Analytics. * Automate operational tasks and eliminate repetitive manual work using Infrastructure as Code and scripting. * Design self-healing capabilities and automated remediation to reduce incidents and improve recovery times. * Investigate production incidents, perform root cause analyses, and implement long-term improvements. * Define, measure, and improve Service Level Indicators (SLIs) and Service Level Objectives (SLOs). * Optimize platform performance, scalability, and operational efficiency. * Work closely with Cloud Engineers to improve platform design, resilience, security, and operational reliability. * Support Data and AI teams by improving the reliability of Azure Databricks environments. * Help shape and raise engineering standards around observability, automation, and operational excellence. * Continuously look for opportunities to reduce operational complexity and improve developer productivity. * Act as a technical point of reference for reliability engineering, helping shape technical decisions while remaining hands-on in the technology. **About the role** As part of the Global Platform Team, you'll work alongside engineers in Cloud, Data, and AI to improve the reliability of our Azure-based platform. Using technologies such as Kubernetes, Terraform, Databricks, GitHub Actions, Grafana, and OpenTelemetry, you'll help ensure that our global Workforce-as-a-Service ecosystem remains reliable, scalable, and resilient. This is a hands-on individual contributor role. You will have significant technical ownership and will be expected to get into the detail when needed, from troubleshooting complex production issues to improving automation, observability and system reliability. You will act as a technical reference point for others, but you will not have people management responsibilities. **About HeadFirst x Impellam** HeadFirst x Impellam is one of Europe's leading providers of workforce and talent solutions. Operating across multiple countries, we are transforming into a cloud-native, AI-powered organization that connects people, technology, and data through a modern digital platform. The Global Platform Team is at the heart of that transformation, enabling engineering teams across Cloud, Data, AI, and Software Engineering to build and operate scalable solutions for the future. ## Related Videos - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Empowering Developer Innovation - Balancing Speed, Security, and Scale](https://www.wearedevelopers.com/videos/1689-empowering-developer-innovation-balancing-speed-security-and-scale) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Best Companies in the Netherlands: Top 25 Companies in 2023 ](https://www.wearedevelopers.com/magazine/193-best-companies-in-the-netherlands-top-25-companies-in-2023)