> Markdown version of [/jobs/ext/2688142-sre-engineer](https://www.wearedevelopers.com/jobs/ext/2688142-sre-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE Engineer - **Company:** Gazelle Global Consulting - **Location:** York, UK - **Experience:** Experienced - **Contract:** Temporary contract - **Skills:** Kubernetes Security, JavaScript (Programming Language), Amazon Web Services, Application Performance Management, JIRA, Microsoft Azure, Bash Shell, Cloud Computing, Software Debugging, DevOps, Monitoring of Systems, Python (Programming Language), Performance Tuning, Reliability Engineering, Runbook, Security Information and Event Management, Software Engineering, Systems Architecture, Systems Integration, Datadog, Scripting, Multi-Cloud, Kubernetes, Dynatrace, Docker, Pagerduty, Golang, Microservices - **Published:** September 3, 2026 - **Apply:** https://www.totaljobs.com/job/site-reliability-engineer/gazelle-global-consulting-ltd-job107933691 ## About the Role * Experience: 10+ years of hands-on experience in Site Reliability Engineering (SRE), DevOps, or Systems Architecture, with at least 3+ years specializing deeply in Datadog administration and configuration. * Cloud & Container Expertise: Deep professional experience working with AWS, Azure, or GCP, paired with heavy production experience managing Kubernetes clusters. * Instrumentation & Coding: Proficiency in systems or scripting languages (e.g., Python, Go, Bash, or JavaScript) and experience instrumenting applications for APM. * Datadog Mastery: Deep understanding of Datadog's core pillars-Infrastructure, APM, Logs, Metrics, Synthetics, and Security Monitoring. Datadog Certifications are a strong plus. * Problem-Solving Mindset: Demonstrated ability to debug complex, distributed microservices architectures under high-pressure incident response scenarios. * Communication Skills: Excellent interpersonal and stakeholder management skills, with the ability to translate technical telemetry data into actionable business and engineering insights. Essential skills/knowledge/experience: Architecture & Implementation * Platform Ownership: Design, deploy, and manage Datadog agents, integrations, and custom metrics across multi-cloud (AWS/Azure/GCP) and containerized (Kubernetes, Docker) environments. * Observability Pipelines: Architect and scale high-throughput log processing, routing, and transformation systems using Datadog. * APM & Infrastructure Monitoring: Configure and optimize Application Performance Monitoring (APM), Distributed Tracing, Real User Monitoring (RUM), and Infrastructure metrics. ## Description * Standardization: Establish company-wide standards for dashboards, monitors, SLOs/SLIs, and alert routing (integrating with PagerDuty, Jira, Opsgenie, etc.). * Cost & Performance Optimization: Audit and optimize Datadog usage, index management, log retention policies, and custom metric volume to maximize ROI and control licensing costs. * Security & Compliance: Leverage Datadog Security products (CSPM, CWPP, Cloud SIEM, Container Security) to maintain compliance postures and mitigate runtime threats. Collaboration & Enablement * Cross-Functional Mentorship: Act as the go-to escalation point and technical mentor for DevOps, SRE, and Software Engineering teams regarding troubleshooting and instrumentation. * Training & Documentation: Create internal documentation, runbooks, and training modules to elevate organizational proficiency in observability. * Vendor Management: Act as the primary technical point of contact for Datadog account teams ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)