Lead SRE (AWS CloudWatch

Systems, Inc
Dallas, TX, United States
8 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon Elastic Compute Cloud Application Performance Management Software Design Documents Monitoring of Systems Role-Based Access Control Prometheus Runbook Datadog Grafana Cloudformation Infrastructure Automation Frameworks
+3 more
Cloudwatch Splunk Dynatrace

Job description

Lead SRE observability workstream focusing on AWS-native monitoring solutions, migrating from third-party tools (Datadog/Splunk) to AWS CloudWatch ecosystem. Design and implement comprehensive observability architecture using CloudWatch Metrics V2, AWS AppSignals, AWS Distro for OpenTelemetry (ADOT), and AWS X-Ray. Establish enterprise-grade dashboards, alerting frameworks, and CloudFormation-based infrastructure automation for monitoring stack deployment.

Top Skills:

  • Amazon CloudWatch Metrics V2
  • AWS AppSignals
  • AWS Distro for OpenTelemetry (ADOT)
  • AWS X-Ray, * Observability Migration Leadership: Lead migration from Datadog/Splunk to AWS CloudWatch Metrics V2, ensuring feature parity, cost optimization, and minimal disruption to monitoring capabilities
  • Deep Metrics Integration: Implement advanced CloudWatch metrics collection from ZeHorizon platforms, custom application metrics, and infrastructure telemetry with enhanced resolution and dimensionality
  • Distributed Tracing & APM: Design and deploy AWS X-Ray distributed tracing architecture integrated with ADOT (AWS Distro for OpenTelemetry) for end-to-end application performance monitoring
  • AppSignals Implementation: Configure AWS AppSignals for automatic service-level objective (SLO) tracking, anomaly detection, and application health monitoring
  • Dashboard & Alerting: Create CloudWatch dashboards with cross-service correlation, CloudWatch Alarms with composite alarms, and EventBridge-based alert routing to incident management systems
  • Infrastructure as Code: Develop CloudFormation templates (CFT) for repeatable observability stack deployment across multi-account/multi-region environments
  • Cost Optimization: Optimize CloudWatch costs through metric filtering, log retention policies, and efficient query patterns while maintaining observability coverage

Observability & Monitoring:

  • Amazon CloudWatch: Expert-level proficiency in CloudWatch Metrics V2 (high-resolution metrics, metric math, anomaly detection), CloudWatch Logs (Logs Insights queries, subscription filters, metric filters), CloudWatch Alarms (composite alarms, alarm actions), and CloudWatch Dashboards (cross-account/cross-region dashboards, custom widgets)
  • AWS X-Ray: Distributed tracing architecture, service maps, trace analysis, sampling rules, X-Ray SDK integration, and X-Ray daemon configuration
  • AWS Distro for OpenTelemetry (ADOT): ADOT Collector deployment (ECS/EKS/EC2), OpenTelemetry SDK instrumentation, exporters configuration (CloudWatch, X-Ray, Prometheus), and custom metric pipelines
  • AWS AppSignals: Service-level objective (SLO) configuration, automatic instrumentation, anomaly detection, and application health scoring
  • Amazon Managed Service for Prometheus (AMP): Prometheus-compatible metrics collection, PromQL queries, and Grafana integration (if hybrid monitoring required)
  • Amazon Managed Grafana (AMG): Dashboard creation, data source integration (CloudWatch, AMP, X-Ray), and role-based access control

Requirements

  • 8+ years of relevant hands-on delivery experience at L6 scope
  • Prior delivery in a large regulated enterprise environment, preferably financial services
  • Ability to write architecture decision records, design documents, runbooks and test evidence for client Tech Risk review
  • Strong stakeholder communication across engineering, security, operations, SRE and delivery teams

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Summarizing SRE concepts and prioritizing technical debt backlogs

Maxim Schepelin Maxim Schepelin · World Congress 2026 Europe

2:50 min

Introduction and the value of runbooks

Hila Fish · World Congress 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

1:32 min

Structuring automated incident workflows between runbooks and raw models

Aram Hakobyan Aram Hakobyan +1 · World Congress 2026 Europe

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all