> Markdown version of [/jobs/ext/2713896-reliability-observability-engineer-2](https://www.wearedevelopers.com/jobs/ext/2713896-reliability-observability-engineer-2). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Reliability Observability Engineer 2 - **Company:** U.S. Bank, National Association - **Location:** Atlanta, GA, United States - **Experience:** Experienced - **Salary:** $86,360.0 - $101,600.0 - **Contract:** Permanent contract - **Skills:** Application Performance Management, Cloud Computing, Distributed Systems, Reliability Engineering, Prometheus, Software Engineering, Datadog, Data Logging, Enterprise Software Applications, Grafana, Kubernetes, Splunk, New Relic (SaaS), Dynatrace, Microservices - **Published:** September 4, 2026 - **Apply:** https://usbank.wd1.myworkdayjobs.com/US_Bank_Careers/job/Atlanta-GA/Reliability-Observability-Engineer-2_2026-0027640 ## About the Role Bachelor's degree, or equivalent work experience - Four to five years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development Preferred Skills/Experience * Expertise in Observability Engineering, Site Reliability Engineering (SRE), or Reliability Engineering. * Strong knowledge of SLIs, SLOs, Error Budgets, and Customer Journey Monitoring. * Demonstrated ability to understand stakeholder needs and guide the development of reliability requirements for large, complex multi-system products. * Hands-on experience with APM, RUM, synthetics, monitoring, logging, tracing, and telemetry frameworks. * Proficiency with Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, or OpenTelemetry. * Experience building, standardizing, and tuning operational dashboards and actionable alerts that communicate service health, customer impact, dependency health, performance trends, failure conditions, severity, ownership, routing, and runbook linkage. * Strong understanding of distributed systems, microservices, cloud platforms, and Kubernetes. * Ability to leverage incident analysis, RCA, and performance data to drive reliability improvements. * Excellent stakeholder management, communication, and technical leadership skills. ## Description * Lead Observability Strategy across critical customer journeys, aligning monitoring capabilities with business outcomes, reliability goals, and customer experience. * Define, implement, and govern SLIs, SLOs, Error Budgets, and Reliability Metrics for enterprise applications and services. * Design and maintain scalable Observability Architectures, including telemetry instrumentation, monitoring frameworks, tagging standards, and alerting models. * Establish Observability Governance for dashboards, alerts, synthetic monitoring, telemetry standards, and lifecycle management of monitoring assets. * Partner with Product, Engineering, SRE, and Operations teams to ensure production readiness, application instrumentation, and reliability measurement. * Develop and optimize Service Health Dashboards and Reporting that provide visibility into availability, latency, customer impact, dependency performance, and SLO compliance. * Analyze telemetry data, incidents, alert history, problem records, and performance trends to identify gaps, reduce alert fatigue, and improve detection accuracy. * Provide technical leadership and mentorship on Distributed Tracing, Logging, Metrics, Synthetic Monitoring, Application Performance Monitoring (APM), Real User Monitoring (RUM), and Alert Governance best practices. ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Effortlessly Scale Prometheus With The Telemetry Data Platform – And Keep your Grafana Dashboards, Too!](https://www.wearedevelopers.com/magazine/3-effortlessly-scale-prometheus-with-the-telemetry-data-platform-and-keep-your-grafana-dashboards-too)