> Markdown version of [/jobs/ext/1243317-sr-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1243317-sr-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Site Reliability Engineer - **Company:** O.C. Tanner - **Location:** Salt Lake City, UT, United States - **Experience:** Expert - **Salary:** $98,800.0 - $148,200.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Apache ActiveMQ, Application Performance Management, Software as a Service, DevOps, Distributed Data Store, Python (Programming Language), PostgreSQL, Redis, Reliability Engineering, Amazon Simple Notification Service (SNS), Software Engineering, Data Streaming, Datadog, Data Logging, Pulumi, Spring Cloud, Amazon ElastiCache, System Availability, Containerization, Kubernetes, Apache Kafka, Amazon Simple Queue Service (SQS), Terraform, Dynatrace, Golang, Programming Languages - **Published:** July 11, 2026 - **Apply:** https://www.careerjet.com/job/us185dc9973250ab0ba6ccb520486fb1d3/eaa ## About the Role * 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, or related roles, with a strong background in production triage, incident response, and operational excellence. * Experience operating large-scale, customer-facing SaaS platforms with high availability and uptime requirements. * Proficiency in Go, Python, Java, or similar programming languages, with demonstrated experience building automation, production tooling, and reliability-focused engineering solutions. * Deep experience with modern Infrastructure-as-Code and GitOps technologies such as Terraform, OpenTofu, CDKTF, Pulumi, ArgoCD, Helm, and Kubernetes. * Hands-on experience with OpenTelemetry, Datadog, Coralogix, or similar observability platforms. * Strong knowledge of AWS services and Kubernetes in production environments. * Deep understanding of monitoring, logging, and distributed tracing for complex systems. * Ability to partner effectively with software engineering and testing teams to design reliable systems, improve application performance, and strengthen quality practices across the software development lifecycle. * Comfortable with participating in on-call rotations and handling high-pressure environments. Bonus Qualifications: * Experience with multiple cloud or cloud-agnostic environments. * Familiarity with security, compliance, and governance frameworks * Experience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora. * Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS, or similar event-driven technologies. ## Description As a Senior Site Reliability Engineer, you will help define the future of reliability for our world-class employee recognition platform. You'll leverage software engineering, automation, and cloud-native technologies to build and operate highly available, scalable systems that serve millions of users. We're looking for someone who is passionate about reliability engineering, continuous improvement, and building self-healing platforms that enable development teams to move faster while delivering exceptional customer experiences., * Improve the availability, scalability, and performance of cloud-native applications through automation, monitoring, and engineering best practices. * Build and evolve observability platforms using OpenTelemetry, Datadog, Coralogix, or similar tools. Establish standards for metrics, logs, traces, and service-level objectives (SLOs) that enable proactive issue detection and resolution. * Lead production triage efforts, rapidly diagnosing and resolving service disruptions. Drive incident management, root cause analysis, and blameless post incident reviews to improve system resilience and reduce recurring issues. * Partner with Engineering, Support and Product teams to embed reliability, observability, and operational excellence throughout the software development lifecycle. * Champion a reliability-first engineering culture by establishing automation standards, monitoring best practices, shift-left quality approaches and shared ownership models that proactive improve resilience, reduce operational risk, and protect the availability of business-critical services. * Collaborate with global engineering teams in a follow-the-sun support model, ensuring seamless 24x7 coverage, effective handoffs, and shared ownership of production services. * Participate in an on-call rotation focused on maintaining service health, reducing operational toil, improving alert quality, and automating repetitive operational tasks., Description: Weyerhaeuser Company is seeking a Process Control/Automation Engineer for our Columbia Falls MDF facility. The successful candidate will be responsible for mentoring… + 15 days ago, About the Role The Site Reliability Engineering team at iCapital is fundamental to ensuring our platform delivers consistent, reliable service to our client base. As a Site Relia… + 2 days ago ## Related Videos - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)