Reliability Monitoring Engineer

Amazon.com, Inc.
Ashburn, VA, United States
14 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$100,000.0 - $150,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) ARM Architecture Software as a Service Continuous Integration Python (Programming Language) Linux Kernel Open Source Technology Prometheus Datadog Data Logging Grafana Containerization
+4 more
Information Technology Splunk New Relic (SaaS) Dynatrace

Job description

We are looking for an Observability Engineer to design and operate the metrics, logging, tracing, and alerting platforms that give engineering teams confidence in the systems they run. The role spans the full observability stack - from collection agents and pipelines to long-term storage, dashboards, and alerting workflows - with a strong focus on usability, signal quality, and operational ROI. The ideal candidate has built and operated observability platforms at scale, understands the trade-offs between open-source and SaaS approaches, and can translate noisy telemetry into actionable insight for both engineers and business stakeholders.

Requirements

  • Bachelor’s degree in Computer Science or a related field.
  • Five or more years of experience in SRE, platform engineering, or observability roles.
  • Deep hands-on experience with Prometheus, Grafana, and at least one major commercial observability platform such as Datadog, New Relic, or Splunk.
  • Strong understanding of OpenTelemetry, distributed tracing, and structured logging.
  • Proficiency in at least one general-purpose language such as Go, Python, or Java.
  • Experience operating high-cardinality, high-throughput metrics and log pipelines.
  • Strong understanding of SLOs, error budgets, and SRE principles.
  • Experience integrating observability with CI/CD and incident management tooling.
  • Solid grasp of Linux internals, networking, and container platforms.
  • Excellent communication and collaboration skills.

Preferred Qualifications

  • Experience with Thanos, Mimir, Cortex, Loki, or Tempo at scale.
  • Contributions to OpenTelemetry or observability open-source projects.
  • Familiarity with eBPF-based observability tooling.
  • Experience driving observability cost optimization initiatives.
  • Exposure to regulated environments with audit-grade logging requirements.

About the company

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential., Vantor

  • Herndon, VA
  • $116,000-169,400 per year Vantor is forging the new frontier of spatial intelligence, helping decision makers and operators navigate what’s happening now and shape what’s coming next. Vantor is a place for …

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

Videos

See all

Related articles

See all