TELECOMMUTE Lead Site Reliability Engineer

Goldenpick Technologies
United States
11 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Test Suite Application Lifecycle Management Application Performance Management Computing Platforms Application Services Microsoft Azure Cloud Computing Cloud Engineering Disaster Recovery Distributed Systems Log Analysis Performance Tuning
+9 more
Regression Testing Release Management Reliability Engineering Site Reliability Engineering Practices Data Logging Reliability of Systems Deployment Automation Terraform Dynatrace

Requirements

  • Strong experience in observability and monitoring, including hands-on expertise with:
  • Dynatrace
  • OpenTelemetry (OTel)
  • Distributed tracing
  • Metrics collection and analysis
  • Centralized logging and log aggregation
  • Alerting and dashboard development
  • Proven experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployments.
  • Strong proficiency in Infrastructure as Code (IaC) using Terraform.
  • Experience with CI/CD pipelines, deployment automation, and operational tooling.
  • Expert knowledge of production systems monitoring, incident management, and operational troubleshooting.
  • Strong understanding of application performance management, distributed systems, and modern cloud-native architectures.
  • Cloud & Platform Expertise
  • Strong experience with Microsoft Azure, including:
  • Azure App Services
  • Resource Groups
  • Azure networking concepts
  • Scaling and performance optimization
  • Deployment and release management
  • Application lifecycle management
  • Experience leveraging Azure-native operational tooling such as:
  • Azure Monitor
  • Application Insights
  • Log Analytics
  • Azure dashboards and alerting
  • Experience supporting cloud-native and hybrid infrastructure environments.
  • Reliability & Engineering Practices
  • Demonstrated experience implementing and operating SRE practices, including:
  • Service Level Objectives (SLOs)
  • Service Level Indicators (SLIs)
  • Error budgets
  • Incident management
  • Problem management
  • Root Cause Analysis (RCA)
  • Reliability automation
  • Ability to improve system reliability through:
  • Performance tuning
  • Capacity planning
  • Observability-driven insights
  • Proactive issue detection
  • Reliability engineering initiatives
  • Experience developing automated recovery mechanisms and self-healing solutions.
  • Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:45 min

Building careers inside distributed technology consulting environments

Oliver Zimmert · LIVE

1:12 min

Automating production test suites using natural language prompts

Jonas Menesklou Jonas Menesklou · WWC 2025

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · WWC 2025

1:05 min

Adapting workplace policies to support distributed technical teams

Tejas Chopra Tejas Chopra · LIVE

2:33 min

Evaluating test suite quality through mutation testing

Ewald Ewald · WWC 2024

Videos

See all

Related articles

See all