> Markdown version of [/jobs/ext/3111451-reliability-engineer-3-sre-monitoring-observability](https://www.wearedevelopers.com/jobs/ext/3111451-reliability-engineer-3-sre-monitoring-observability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Reliability Engineer 3 (SRE, Monitoring & Observability) - **Company:** U.S. Bank - **Location:** Gresham, OR, United States - **Experience:** Expert - **Salary:** $105,400.0 - $124,000.0 - **Contract:** Permanent contract - **Skills:** Adobe InDesign, Cloud Computing, Distributed Systems, Monitoring of Systems, Reliability Engineering, Prometheus, Software Engineering, Statistical Process Control (SPC), Datadog, Data Logging, Grafana, Kubernetes, Performance Monitor, Splunk, New Relic (SaaS), Servicenow, Microservices - **Published:** September 27, 2026 - **Apply:** https://dejobs.org/x/x/6D44E04C16244999B34338F63E49D426/job/ ## About the Role * Bachelor's degree, or equivalent work experience * Five to seven years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development Preferred Skills/Experience * Expertise in Site Reliability Engineering (SRE) , or Reliability Engineering . * Strong knowledge of SLIs, SLOs, Error Budgets, and Customer Journey Monitoring . * Demonstrated ability to understand stakeholder needs and guide the development of reliability requirements for large, complex multi-system products. * Hands-on experience with APM, RUM, synthetics, monitoring, logging, tracing, and telemetry frameworks . * Proficiency with Datadog, , Splunk, Grafana, Prometheus, New Relic, Elastic, or OpenTelemetry . * Experience building, standardizing, and tuning operational dashboards and actionable alerts that communicate service health, customer impact, dependency health, performance trends, failure conditions, severity, ownership, routing, and runbook linkage. * Strong understanding of distributed systems, microservices, cloud platforms, and Kubernetes . * Ability to leverage incident analysis, RCA, and performance data to drive reliability improvements. * Excellent stakeholder management, communication, and technical leadership skills. * Hands on experience with ServiceNow. ## Description As a Reliability Engineer, your role will be a combination of supporting production applications and proactively looking for ways to automate your discoveries, eliminate incidents from recurring and/or reduce the time it takes to get our customers back up and running. In addition, you'll focus on improving the following for our applications: availability, latency, performance, efficiency, and effective proactive monitoring. The reliability engineer interfaces with business users, development teams and system administrators to ensure systems perform to meet their business needs and specifications. Responsibilities include: Developing, coordinating, and conducting technical reliability studies on engineering designs to assess the likelihood that a product/process performs its intended function over the intended lifecycle. Measuring and analyzing the reliability of the design, materials, processes, cost, and final products of production. Recommending design or test methods and statistical process control procedures for achieving required levels of product reliability. Completing risk analysis studies of new designs and processes. Undertaking testing and analysis on failures, proposing changes in design or formulation to improve system and/or process reliability. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Monitoring as Code - Managing your dashboards at scale](https://www.wearedevelopers.com/videos/753-monitoring-as-code-managing-your-dashboards-at-scale) - [The Memory Leak That Ate Our Cluster: A Postmortem](https://www.wearedevelopers.com/videos/2057-the-memory-leak-that-ate-our-cluster-a-postmortem) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Keycloak case study: Making users happy with service level indicators and observability](https://www.wearedevelopers.com/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london) - [Best Companies to work for in London: Top 25 Companies in 2023](https://www.wearedevelopers.com/magazine/187-best-companies-to-work-for-in-london-top-25-companies-in-2023)