Hiring: Site Reliability Engineer (SRE) | Location: Remote

Source Inc.
United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) .NET Framework Amazon Web Services Data Analysis Data Integration Extract Transform Load (ETL) Python (Programming Language) Reliability Engineering Power BI Runbook Data Logging Google Cloud
+10 more
Grafana Containerization Kubernetes Information Technology Restful APIs Splunk Appdynamics Dynatrace Servicenow Microservices

Job description

  • Assess application reliability, performance, telemetry coverage, and operational maturity.
  • Design and implement observability solutions using APM, distributed tracing, structured logging, and telemetry pipelines.
  • Define and optimize SLIs, SLOs, alerting strategies, and reliability metrics.
  • Lead incident management, RCA, and operational governance initiatives.
  • Develop executive dashboards, scorecards, and operational analytics.
  • Integrate telemetry across monitoring and cloud platforms.
  • Create monitoring standards, runbooks, and SRE best practices.

Requirements

We are hiring an experienced Site Reliability Engineer (SRE) with 15+ years of IT experience to drive enterprise observability, reliability engineering, and operational excellence for large-scale cloud and microservices environments., * Observability: Splunk, Dynatrace, Grafana, AppDynamics, OpenTelemetry, ServiceNow Performance Analytics

  • Programming: Python, Java, .NET, REST APIs, Automation & Scripting
  • Cloud: AWS, Google Cloud Platform, Kubernetes, Microservices, Container Platforms
  • Analytics: Power BI, ETL/Data Integration, Dashboard Development, KPI & Scorecards
  • SRE: Incident Management, RCA, Availability Engineering, Service Health Monitoring, Operational Governance, * 15+ years of experience in Site Reliability Engineering, Observability Engineering, Infrastructure Operations, or Application Support.
  • Strong experience implementing enterprise observability and telemetry solutions.
  • Expertise in executive reporting, operational analytics, and reliability scorecards.
  • Excellent stakeholder communication and leadership skills.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:50 min

Introduction and the value of runbooks

Hila Fish · WWC 2023

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

1:32 min

Structuring automated incident workflows between runbooks and raw models

Aram Hakobyan Aram Hakobyan +1 · WWC Europe 2026

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

Videos

See all

Related articles

See all