Senior Site Reliability Engineer

Abbott Laboratories
Sunnyvale, CA, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Systems Engineering Microsoft Azure Bash Shell Software Debugging DevOps Distributed Systems Fault Tolerance Python (Programming Language) Windows PowerShell Reliability Engineering Cloud Services Prometheus
+13 more
Software Engineering Datadog Data Logging Scripting Load Balancing Cloud Monitoring System Availability Grafana Kubernetes Information Technology Deployment Automation Docker Golang

Job description

Experteer Overview As a Senior Site Reliability Engineer in Abbott’s Cardiac Rhythm Management DevOps team, you ensure Merlin.net operates with high availability, scalability, and safety for patient data. You will bridge development and operations to build a resilient, real-time platform used by doctors and care teams. You’ll tackle performance bottlenecks, implement monitoring and automation, and drive scalable cloud and distributed systems. This role offers the chance to contribute to mission-critical healthcare infrastructure at scale and influence reliability practices across the organization. Compensation / Benefits * Design and maintain highly available, fault-tolerant systems meeting uptime and safety constraints * Identify performance bottlenecks and ensure low-latency, high-throughput real-time responsiveness * Define and monitor SLIs/SLOs and error budgets * Develop and implement monitoring, logging, tracing, and alerting for system health at scale * Automate provisioning, deployment, testing, and recovery * Plan and implement scaling strategies for cloud deployments (Azure) and distributed systems * Collaborate with software, security, quality, and compliance teams to embed SRE best practices * Create runbooks and playbooks; lead blameless postmortems and track action items * Participate in architecture reviews, capacity planning, and reliability roadmap planning Tasks * Bachelor’s in Computer Science, Software Engineering, Systems Engineering, or related technical discipline (or equivalent experience) * Strong communication skills for cross-functional collaboration * Analytical, problem-solving, and debugging skills for distributed systems under pressure * Proficiency in at least one scripting/automation language (Python, Go, Bash, PowerShell) * Expertise with Microsoft Azure (AKS, Azure Monitor, Azure DevOps, Azure Policy) and managed services * Container orchestration experience with Kubernetes and Docker at scale * Observability experience with Prometheus, Grafana, ELK/EFK, Datadog, or similar * Experience designing and operating CI/CD pipelines with safe deployment strategies * Strong understanding of distributed systems, load balancing, service meshes, and fault tolerance * Incident management experience including on-call, RCA, and recurrence prevention Key requirements *

Requirements

deployment, testing, and recovery * Plan and implement scaling strategies for cloud deployments (Azure) and distributed systems * Collaborate with software, security, quality, and compliance teams to embed SRE best practices * Create runbooks and playbooks; lead blameless postmortems and track action items * Participate in architecture reviews, capacity planning, and reliability roadmap planning Tasks * Bachelor’s in Computer Science, Software Engineering, Systems Engineering, or related technical discipline (or equivalent experience) * Strong communication skills for cross-functional collaboration * Analytical, problem-solving, and debugging skills for distributed systems under pressure * Proficiency in at least one scripting/automation language (Python, Go, Bash, PowerShell) * Expertise with Microsoft Azure (AKS, Azure Monitor, Azure DevOps, Azure Policy) and managed services * Container orchestration experience with Kubernetes and Docker at scale * Observability experience with aaaaaaa aay_ Grafana, ELK/EFK, Datadog, or similar * Experience designing and operating CI/CD pipelines with safe deployment strategies * Strong understanding of distributed systems, load balancing, service meshes, and fault tolerance * Incident management experience including on-call, RCA, and recurrence prevention Key requirements *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · WWC Europe 2026

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

Videos

See all

Related articles

See all