Sr. Site Reliability Engineer

O.C. Tanner
Salt Lake City, UT, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$98,800.0 - $148,200.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Apache ActiveMQ Application Performance Management Software as a Service DevOps Distributed Data Store Python (Programming Language) PostgreSQL Redis Reliability Engineering Amazon Simple Notification Service (SNS) Software Engineering
+15 more
Data Streaming Datadog Data Logging Pulumi Spring Cloud Amazon ElastiCache System Availability Containerization Kubernetes Apache Kafka Amazon Simple Queue Service (SQS) Terraform Dynatrace Golang Programming Languages

Job description

As a Senior Site Reliability Engineer, you will help define the future of reliability for our world-class employee recognition platform. You’ll leverage software engineering, automation, and cloud-native technologies to build and operate highly available, scalable systems that serve millions of users. We’re looking for someone who is passionate about reliability engineering, continuous improvement, and building self-healing platforms that enable development teams to move faster while delivering exceptional customer experiences., * Improve the availability, scalability, and performance of cloud-native applications through automation, monitoring, and engineering best practices.

  • Build and evolve observability platforms using OpenTelemetry, Datadog, Coralogix, or similar tools. Establish standards for metrics, logs, traces, and service-level objectives (SLOs) that enable proactive issue detection and resolution.
  • Lead production triage efforts, rapidly diagnosing and resolving service disruptions. Drive incident management, root cause analysis, and blameless post incident reviews to improve system resilience and reduce recurring issues.
  • Partner with Engineering, Support and Product teams to embed reliability, observability, and operational excellence throughout the software development lifecycle.
  • Champion a reliability-first engineering culture by establishing automation standards, monitoring best practices, shift-left quality approaches and shared ownership models that proactive improve resilience, reduce operational risk, and protect the availability of business-critical services.
  • Collaborate with global engineering teams in a follow-the-sun support model, ensuring seamless 24x7 coverage, effective handoffs, and shared ownership of production services.
  • Participate in an on-call rotation focused on maintaining service health, reducing operational toil, improving alert quality, and automating repetitive operational tasks., Description: Weyerhaeuser Company is seeking a Process Control/Automation Engineer for our Columbia Falls MDF facility. The successful candidate will be responsible for mentoring…
  • 15 days ago, About the Role The Site Reliability Engineering team at iCapital is fundamental to ensuring our platform delivers consistent, reliable service to our client base. As a Site Relia…
  • 2 days ago

Requirements

  • 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, or related roles, with a strong background in production triage, incident response, and operational excellence.

  • Experience operating large-scale, customer-facing SaaS platforms with high availability and uptime requirements.
  • Proficiency in Go, Python, Java, or similar programming languages, with demonstrated experience building automation, production tooling, and reliability-focused engineering solutions.
  • Deep experience with modern Infrastructure-as-Code and GitOps technologies such as Terraform, OpenTofu, CDKTF, Pulumi, ArgoCD, Helm, and Kubernetes.
  • Hands-on experience with OpenTelemetry, Datadog, Coralogix, or similar observability platforms.
  • Strong knowledge of AWS services and Kubernetes in production environments.
  • Deep understanding of monitoring, logging, and distributed tracing for complex systems.
  • Ability to partner effectively with software engineering and testing teams to design reliable systems, improve application performance, and strengthen quality practices across the software development lifecycle.
  • Comfortable with participating in on-call rotations and handling high-pressure environments.

Bonus Qualifications:

  • Experience with multiple cloud or cloud-agnostic environments.
  • Familiarity with security, compliance, and governance frameworks
  • Experience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora.
  • Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS, or similar event-driven technologies.

About the company

O.C. Tanner is the global leader in software and services that improve workplace culture through meaningful employee experiences. Our Culture Cloud is a suite of apps designed to enhance the employee experience with strategic recognition, service awards, wellbeing, leadership, and events that help people thrive at work. Our Culture by Design approach provides expert services to organizations looking to create great workplaces. Our global team of 1,500 people hail from 58 countries and speak 62 languages. As programmers, researchers, designers, client professionals and craftspeople we create the tech, tools and awards that connect employees to purpose at thousands of companies. Join us as we help people all over the world thrive at work. Location: Salt Lake City, UT

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · WWC Europe 2026

1:33 min

Case study on adopting Kubernetes and Golang effectively

Andrew Holway · LIVE

Videos

See all

Related articles

See all