Engineer III, Site Reliability
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+12 more
Job description
As a Site Reliability Engineer, you will own the reliability, scalability, and operational health of a defined set of cloud services that support mission-critical pharmacy automation systems used by healthcare providers worldwide., * Own reliability outcomes for assigned services, ensuring strong instrumentation, actionable alerts, meaningful dashboards, and up-to-date runbooks.
- Define and implement SLIs and SLOs in partnership with product and engineering teams, and surface reliability performance in regular Cloud Operations reviews.
- Identify operational toil and design automation to eliminate repetitive manual work.
- Drive continuous improvement initiatives that increase observability, automation coverage, and system resilience.
Incident Response & Operational Excellence
- Participate in the SRE on-call rotation, progressing from secondary to primary ownership as readiness increases.
- Command Sev-2 and Sev-3 incidents independently over time, with pairing and coaching from a Senior SRE; act as technical lead during Sev-1 incidents.
- Lead blameless post-incident reviews and own follow-up actions through completion.
- Partner closely with managed services providers (IBM, HCL) to ensure clean escalation paths from L1/L2 monitoring into SRE ownership.
Platform, CI/CD & Observability
- Design, build, and operate CI/CD pipelines supporting cloud-native application delivery using tools such as GitHub Actions, CodeFresh, TeamCity, and Octopus Deploy.
- Automate infrastructure and platform services using Infrastructure as Code (Terraform preferred).
- Contribute to the evolution of Omnicell’s observability platform, including intelligent alerting, ML-based anomaly detection, and automated diagnostics.
- Participate in architecture and launch readiness reviews, bringing a reliability lens to system design.
- Help establish reference implementations and “golden paths” that enable product teams to launch services with reliability built in from day one., * Collaborate: Partner closely with product engineering, security, and operations teams to build shared ownership of reliability.
- Inspire: Influence reliability best practices across teams by modeling calm, structured incident leadership.
- Develop: Continuously build your technical depth while learning directly from a senior SRE mentor.
- Execute: Take ownership of services, incidents, and follow-through-turning lessons learned into measurable improvements.
- Impact: Help shape foundational SRE practices and introduce modern reliability and AIOps capabilities that scale with the business.
Growth & Career Path
This role is intentionally designed as a growth role . With strong performance and increasing ownership, the natural progression is into a Senior Site Reliability Engineer position as the practice scales. Omnicell also supports lateral growth into platform engineering, security engineering, or product engineering for SREs who discover adjacent passions.
Work Conditions
- Remote or hybrid work environment supported.
- Up to 10% travel as needed.
- Participation in an SRE on-call rotation is required.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical field.
- 5+ years of experience in software or platform engineering, including 3+ years in an SRE, DevOps, or reliability-focused role.
- Strong hands-on experience with at least one major public cloud platform (AWS, Azure, or GCP).
- Proficiency in Python or another object-oriented programming language for automation and tooling.
- Production experience with Kubernetes, Docker, and Helm.
- Experience implementing Infrastructure as Code using Terraform or similar frameworks.
- Working knowledge of modern observability tools across metrics, logs, and tracing.
- Real-world incident response experience, including on-call participation and post-incident write-ups.
- Solid Linux system administration skills.
- Collaborative, coachable mindset with a desire to grow under senior mentorship., * Experience working in regulated environments such as healthcare, financial services, or government (HIPAA, SOC 2, or similar).
- Familiarity with managed service provider models for L1/L2 operations.
- Exposure to AIOps, ML-based anomaly detection, or LLM-assisted incident triage.
- Understanding of GitOps principles and tools such as ArgoCD or Flux.
- Experience operating secure, compliant Kubernetes platforms.
- Familiarity with chaos engineering, messaging systems (Kafka, RabbitMQ), or stateful services in Kubernetes.
About the company
At Omnicell, success isn’t just about what you deliver-it’s about how you deliver it. Our Elevate Behaviors guide how we work together and create impact
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Is Software Engineering Over-Saturated?
Fully Remote Software Engineer Jobs
Why Upskilling And Reskilling is Important For Developers