Lead Site Reliability Engineer (Dynatrace)

SF Partners
UK
16 days ago
Apply on www.totaljobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Cloud Computing Continuous Integration Monitoring of Systems Reliability Engineering Datadog Kubernetes Dynatrace

Job description

This is not a role for someone who has simply used Dynatrace dashboards. We’re looking for an engineer who has been involved in the implementation, configuration and ongoing ownership of Dynatrace, and can operate as a technical SME within complex production environments.

Requirements

  • Strong hands-on Dynatrace implementation and administration experience
  • Experience designing and implementing observability/monitoring solutions end-to-end
  • Strong SRE and production engineering background
  • Experience configuring instrumentation, metrics, alerting and monitoring
  • Understanding of technologies such as OneAgent, ActiveGate, distributed tracing and application/infrastructure monitoring
  • Experience troubleshooting complex production issues using observability tooling
  • Cloud experience across AWS, Azure and/or GCP
  • Exposure to CI/CD, automation and modern software delivery
  • Kubernetes/containerisation experience beneficial but not essential
  • Ability to operate as a technical lead/SME, supporting and mentoring other engineers

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.totaljobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · World Congress 2025

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · World Congress 2022

Videos

See all

Related articles

See all