> Markdown version of [/jobs/ext/3208964-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3208964-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** HeadFirst - **Location:** Hoofddorp, Netherlands (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, Microsoft Azure, Bash Shell, Cloud Computing, Cloud Engineering, DevOps, Distributed Systems, Github, Python (Programming Language), Log Analysis, Reliability Engineering, Data Logging, Scripting, Cloud Monitoring, Grafana, Informatica Cloud, Kubernetes, Integration Frameworks, Terraform, Databricks - **Published:** September 3, 2026 - **Apply:** https://nl.indeed.com/viewjob?jk=b038cad4e0f69597 ## About the Role You're passionate about building reliable systems and solving operational challenges through engineering rather than manual intervention. You enjoy understanding how distributed systems behave, thrive in cloud-native environments, and are always looking for ways to improve automation, resilience, and observability. You stay calm under pressure, take ownership of problems, and enjoy collaborating with others to continuously improve the platform's reliability., * 4+ years of experience as a Site Reliability Engineer, Platform Engineer, DevOps Engineer, or Cloud Engineer; * Strong hands-on experience with Microsoft Azure; * Experience with Infrastructure as Code using Terraform; * Experience building and maintaining CI/CD pipelines using GitHub Actions or Azure DevOps; * Experience with observability tools such as Grafana, OpenTelemetry, Azure Monitor, or Log Analytics; * Strong scripting skills using Python, Bash, or similar languages; * Experience supporting distributed cloud platforms in production; * Experience with incident management, root cause analysis, and post-incident improvements; * Familiarity with GitOps principles and modern deployment practices; * Experience with Azure Databricks is a strong advantage; * Experience with SnapLogic or similar integration platforms is a plus. We know there's no such thing as the perfect candidate. If this role excites you but you don't meet every single requirement, we'd still love to hear from you. We're just as interested in your potential, mindset, and ambition as we are in your experience. ## Description * Improve the reliability, availability, and performance of our Azure platform and production environments; * Build and improve monitoring, logging, and alerting using Grafana, OpenTelemetry, Azure Monitor, and Log Analytics; * Automate operational tasks and eliminate repetitive manual work using Infrastructure as Code and scripting; * Design self-healing capabilities and automated remediation to reduce incidents and improve recovery times; * Investigate production incidents, perform root cause analyses, and implement long-term improvements; * Define, measure, and improve Service Level Indicators (SLIs) and Service Level Objectives (SLOs); * Optimize platform performance, scalability, and operational efficiency; * Work closely with Cloud Engineers to improve platform architecture, resilience, and security; * Support Data and AI teams by improving the reliability of Azure Databricks environments; * Drive engineering best practices in observability, automation, and operational excellence; * Continuously look for opportunities to reduce operational complexity and improve developer productivity., As part of the Global Platform Team, you'll work alongside engineers in Cloud, Data, and AI to improve the reliability of our Azure-based platform. Using technologies such as Kubernetes, Terraform, Databricks, GitHub Actions, Grafana, and OpenTelemetry, you'll help ensure that our global Workforce-as-a-Service ecosystem remains reliable, scalable, and resilient.