> Markdown version of [/jobs/ext/1501474-reliability-monitoring-engineer](https://www.wearedevelopers.com/jobs/ext/1501474-reliability-monitoring-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Reliability Monitoring Engineer - **Company:** Amazon.com, Inc. - **Location:** Ashburn, VA, United States (Remote available) - **Experience:** Expert - **Salary:** $100,000.0 - $150,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), ARM Architecture, Software as a Service, Continuous Integration, Python (Programming Language), Linux Kernel, Open Source Technology, Prometheus, Datadog, Data Logging, Grafana, Containerization, Information Technology, Splunk, New Relic (SaaS), Dynatrace - **Published:** July 30, 2026 - **Apply:** https://www.careerjet.com/jobad/us2f119502e69b3c0048e698246ae34306 ## About the Role * Bachelor's degree in Computer Science or a related field. * Five or more years of experience in SRE, platform engineering, or observability roles. * Deep hands-on experience with Prometheus, Grafana, and at least one major commercial observability platform such as Datadog, New Relic, or Splunk. * Strong understanding of OpenTelemetry, distributed tracing, and structured logging. * Proficiency in at least one general-purpose language such as Go, Python, or Java. * Experience operating high-cardinality, high-throughput metrics and log pipelines. * Strong understanding of SLOs, error budgets, and SRE principles. * Experience integrating observability with CI/CD and incident management tooling. * Solid grasp of Linux internals, networking, and container platforms. * Excellent communication and collaboration skills. Preferred Qualifications * Experience with Thanos, Mimir, Cortex, Loki, or Tempo at scale. * Contributions to OpenTelemetry or observability open-source projects. * Familiarity with eBPF-based observability tooling. * Experience driving observability cost optimization initiatives. * Exposure to regulated environments with audit-grade logging requirements. ## Description We are looking for an Observability Engineer to design and operate the metrics, logging, tracing, and alerting platforms that give engineering teams confidence in the systems they run. The role spans the full observability stack - from collection agents and pipelines to long-term storage, dashboards, and alerting workflows - with a strong focus on usability, signal quality, and operational ROI. The ideal candidate has built and operated observability platforms at scale, understands the trade-offs between open-source and SaaS approaches, and can translate noisy telemetry into actionable insight for both engineers and business stakeholders. ## Related Videos - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Effortlessly Scale Prometheus With The Telemetry Data Platform – And Keep your Grafana Dashboards, Too!](https://www.wearedevelopers.com/magazine/3-effortlessly-scale-prometheus-with-the-telemetry-data-platform-and-keep-your-grafana-dashboards-too) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read)