SRE Developer

Tekshapers Inc
Fort Mill, SC, United States
23 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Cloud Computing Distributed Systems Elasticsearch Python (Programming Language) Reliability Engineering Prometheus Scripting Google Cloud Grafana Reliability of Systems
+5 more
Kubernetes Performance Monitor Dynatrace Elk Stack Microservices

Job description

We’re looking for an experienced Observability Site Reliability Engineer (SRE) to drive the design and implementation of scalable observability solutions across distributed systems.

The ideal candidate will have strong hands-on experience with Dynatrace, Grafana, and ELK Stack, along with expertise in End-to-End (E2E) Tracing, Golden Signals, Real User Monitoring (RUM), and Voice of Customer (VoC) metrics.

You’ll collaborate with cross-functional teams to enhance system reliability, ensure proactive monitoring, and deliver actionable insights through data-driven observability practices.

You will also be responsible for designing and maintaining executive and operational dashboards that provide actionable insights into system health, performance trends, and user experience across environments. Key Responsibilities:

  • Implement and maintain observability platforms using Dynatrace, Grafana, and ELK Stack.

  • Enable E2E tracing and define Golden Signals to monitor system health.

Requirements

  • Drive incident response, root cause analysis, and continuous reliability improvements. Skills & Experience Required:

  • Strong expertise with Dynatrace, Grafana, and ELK.

  • Hands-on experience in E2E tracing and distributed tracing tools.

  • Good understanding of Golden Signals, SLI/SLOs, and APM concepts.

  • Experience with RUM, VoC, microservices, Kubernetes, and cloud platforms (AWS/Azure/Google Cloud Platform).

  • Scripting knowledge (Python, Shell, Go) and CI/CD pipeline integration. Nice to Have / Familiarity:

  • Exposure to Cribl for log and telemetry pipeline management.

  • Knowledge of OpenTelemetry, Prometheus, and AIOps frameworks.

  • Certifications in Dynatrace, Grafana, or Elastic Stack.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:07 min

Architecting the availability stack with Prometheus and Grafana

Gabriel Labachelerie · World Congress 2023

1:31 min

Exploring the core components of the ELK stack

Derek Binkley · LIVE

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · World Congress 2025

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

Videos

See all

Related articles

See all