TELECOMMUTE Lead Site Reliability Engineer
Goldenpick Technologies
United States
11 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source
Tech stack
Test Suite
Application Lifecycle Management
Application Performance Management
Computing Platforms
Application Services
Microsoft Azure
Cloud Computing
Cloud Engineering
Disaster Recovery
Distributed Systems
Log Analysis
Performance Tuning
+9 more
Regression Testing
Release Management
Reliability Engineering
Site Reliability Engineering Practices
Data Logging
Reliability of Systems
Deployment Automation
Terraform
Dynatrace
Requirements
- Strong experience in observability and monitoring, including hands-on expertise with:
- Dynatrace
- OpenTelemetry (OTel)
- Distributed tracing
- Metrics collection and analysis
- Centralized logging and log aggregation
- Alerting and dashboard development
- Proven experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployments.
- Strong proficiency in Infrastructure as Code (IaC) using Terraform.
- Experience with CI/CD pipelines, deployment automation, and operational tooling.
- Expert knowledge of production systems monitoring, incident management, and operational troubleshooting.
- Strong understanding of application performance management, distributed systems, and modern cloud-native architectures.
- Cloud & Platform Expertise
- Strong experience with Microsoft Azure, including:
- Azure App Services
- Resource Groups
- Azure networking concepts
- Scaling and performance optimization
- Deployment and release management
- Application lifecycle management
- Experience leveraging Azure-native operational tooling such as:
- Azure Monitor
- Application Insights
- Log Analytics
- Azure dashboards and alerting
- Experience supporting cloud-native and hybrid infrastructure environments.
- Reliability & Engineering Practices
- Demonstrated experience implementing and operating SRE practices, including:
- Service Level Objectives (SLOs)
- Service Level Indicators (SLIs)
- Error budgets
- Incident management
- Problem management
- Root Cause Analysis (RCA)
- Reliability automation
- Ability to improve system reliability through:
- Performance tuning
- Capacity planning
- Observability-driven insights
- Proactive issue detection
- Reliability engineering initiatives
- Experience developing automated recovery mechanisms and self-healing solutions.
- Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
about 2 years ago
EM
Eli McGarvie
Find a Developer Job: 12 Best Job Sites For Developers
over 3 years ago
DC
Daniel Cranney
Mastering Remote Work: Tips for Developers
over 1 year ago
EM
Eli McGarvie
Best Job Boards for Remote Work for Developers
almost 3 years ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
EM
Eli McGarvie
The Best Job Search Websites of 2025
over 2 years ago