> Markdown version of [/jobs/ext/2746221-reliability-observability-engineer-2](https://www.wearedevelopers.com/jobs/ext/2746221-reliability-observability-engineer-2). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Reliability Observability Engineer 2 - **Company:** U.S. Bank - **Location:** Charlotte, NC, United States - **Experience:** Experienced - **Salary:** $86,360.0 - $101,600.0 - **Contract:** Permanent contract - **Skills:** Application Performance Management, Cloud Computing, Distributed Systems, Reliability Engineering, Prometheus, Software Engineering, Datadog, Data Logging, Enterprise Software Applications, Grafana, Kubernetes, Splunk, New Relic (SaaS), Dynatrace, Microservices - **Published:** September 6, 2026 - **Apply:** https://www.beyondcharlotte.com/job.asp?id=3380058336&tx=VT656TYI&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * Bachelor's degree, or equivalent work experience * Four to five years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development Preferred Skills/Experience * Expertise in Observability Engineering , Site Reliability Engineering (SRE) , or Reliability Engineering . * Strong knowledge of SLIs, SLOs, Error Budgets, and Customer Journey Monitoring . * Demonstrated ability to understand stakeholder needs and guide the development of reliability requirements for large, complex multi-system products. * Hands-on experience with APM, RUM, synthetics, monitoring, logging, tracing, and telemetry frameworks . * Proficiency with Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, or OpenTelemetry . * Experience building, standardizing, and tuning operational dashboards and actionable alerts that communicate service health, customer impact, dependency health, performance trends, failure conditions, severity, ownership, routing, and runbook linkage. * Strong understanding of distributed systems, microservices, cloud platforms, and Kubernetes . * Ability to leverage incident analysis, RCA, and performance data to drive reliability improvements. * Excellent stakeholder management, communication, and technical leadership skills. ## Description * Lead Observability Strategy across critical customer journeys, aligning monitoring capabilities with business outcomes, reliability goals, and customer experience. * Define, implement, and govern SLIs, SLOs, Error Budgets, and Reliability Metrics for enterprise applications and services. * Design and maintain scalable Observability Architectures , including telemetry instrumentation, monitoring frameworks, tagging standards, and alerting models. * Establish Observability Governance for dashboards, alerts, synthetic monitoring, telemetry standards, and lifecycle management of monitoring assets. * Partner with Product, Engineering, SRE, and Operations teams to ensure production readiness, application instrumentation, and reliability measurement. * Develop and optimize Service Health Dashboards and Reporting that provide visibility into availability, latency, customer impact, dependency performance, and SLO compliance. * Analyze telemetry data, incidents, alert history, problem records, and performance trends to identify gaps, reduce alert fatigue, and improve detection accuracy. * Provide technical leadership and mentorship on Distributed Tracing, Logging, Metrics, Synthetic Monitoring, Application Performance Monitoring (APM), Real User Monitoring (RUM), and Alert Governance best practices. ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london)