SRE Engineer

Gazelle Global Consulting
York, UK
3 days ago
Apply on www.totaljobs.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Kubernetes Security JavaScript (Programming Language) Amazon Web Services Application Performance Management JIRA Microsoft Azure Bash Shell Cloud Computing Software Debugging DevOps Monitoring of Systems Python (Programming Language)
+16 more
Performance Tuning Reliability Engineering Runbook Security Information and Event Management Software Engineering Systems Architecture Systems Integration Datadog Scripting Multi-Cloud Kubernetes Dynatrace Docker Pagerduty Golang Microservices

Job description

  • Standardization: Establish company-wide standards for dashboards, monitors, SLOs/SLIs, and alert routing (integrating with PagerDuty, Jira, Opsgenie, etc.).
  • Cost & Performance Optimization: Audit and optimize Datadog usage, index management, log retention policies, and custom metric volume to maximize ROI and control licensing costs.
  • Security & Compliance: Leverage Datadog Security products (CSPM, CWPP, Cloud SIEM, Container Security) to maintain compliance postures and mitigate runtime threats.

Collaboration & Enablement

  • Cross-Functional Mentorship: Act as the go-to escalation point and technical mentor for DevOps, SRE, and Software Engineering teams regarding troubleshooting and instrumentation.
  • Training & Documentation: Create internal documentation, runbooks, and training modules to elevate organizational proficiency in observability.
  • Vendor Management: Act as the primary technical point of contact for Datadog account teams

Requirements

  • Experience: 10+ years of hands-on experience in Site Reliability Engineering (SRE), DevOps, or Systems Architecture, with at least 3+ years specializing deeply in Datadog administration and configuration.
  • Cloud & Container Expertise: Deep professional experience working with AWS, Azure, or GCP, paired with heavy production experience managing Kubernetes clusters.
  • Instrumentation & Coding: Proficiency in systems or scripting languages (e.g., Python, Go, Bash, or JavaScript) and experience instrumenting applications for APM.
  • Datadog Mastery: Deep understanding of Datadog’s core pillars-Infrastructure, APM, Logs, Metrics, Synthetics, and Security Monitoring. Datadog Certifications are a strong plus.
  • Problem-Solving Mindset: Demonstrated ability to debug complex, distributed microservices architectures under high-pressure incident response scenarios.
  • Communication Skills: Excellent interpersonal and stakeholder management skills, with the ability to translate technical telemetry data into actionable business and engineering insights.

Essential skills/knowledge/experience:

Architecture & Implementation

  • Platform Ownership: Design, deploy, and manage Datadog agents, integrations, and custom metrics across multi-cloud (AWS/Azure/GCP) and containerized (Kubernetes, Docker) environments.
  • Observability Pipelines: Architect and scale high-throughput log processing, routing, and transformation systems using Datadog.
  • APM & Infrastructure Monitoring: Configure and optimize Application Performance Monitoring (APM), Distributed Tracing, Real User Monitoring (RUM), and Infrastructure metrics.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.totaljobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

5:47 min

Integrating user stories and test automation via Jira tools

Christoph Ruggenthaler · LIVE

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

Videos

See all

Related articles

See all