> Markdown version of [/jobs/ext/2309525-engineer-iii-site-reliability](https://www.wearedevelopers.com/jobs/ext/2309525-engineer-iii-site-reliability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Engineer III, Site Reliability - **Company:** Omnicell, Inc. - **Location:** Cranberry Township, PA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Cloud Computing, Continuous Integration, DevOps, Github, Python (Programming Language), Linux System Administration, Enterprise Messaging Systems, Octopus Deploy, Object-Oriented Software Development, RabbitMQ, Reliability Engineering, Site Reliability Engineering Practices, Cloud Services, Software Deployment, Large Language Models, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Apache Kafka, Teamcity, Terraform, Docker - **Published:** August 30, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3369820383&tx=JL11003LFU&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * Bachelor's degree in Computer Science, Engineering, or a related technical field. * 5+ years of experience in software or platform engineering, including 3+ years in an SRE, DevOps, or reliability-focused role. * Strong hands-on experience with at least one major public cloud platform (AWS, Azure, or GCP). * Proficiency in Python or another object-oriented programming language for automation and tooling. * Production experience with Kubernetes, Docker, and Helm. * Experience implementing Infrastructure as Code using Terraform or similar frameworks. * Working knowledge of modern observability tools across metrics, logs, and tracing. * Real-world incident response experience, including on-call participation and post-incident write-ups. * Solid Linux system administration skills. * Collaborative, coachable mindset with a desire to grow under senior mentorship., * Experience working in regulated environments such as healthcare, financial services, or government (HIPAA, SOC 2, or similar). * Familiarity with managed service provider models for L1/L2 operations. * Exposure to AIOps, ML-based anomaly detection, or LLM-assisted incident triage. * Understanding of GitOps principles and tools such as ArgoCD or Flux. * Experience operating secure, compliant Kubernetes platforms. * Familiarity with chaos engineering, messaging systems (Kafka, RabbitMQ), or stateful services in Kubernetes. ## Description As a Site Reliability Engineer, you will own the reliability, scalability, and operational health of a defined set of cloud services that support mission-critical pharmacy automation systems used by healthcare providers worldwide., * Own reliability outcomes for assigned services, ensuring strong instrumentation, actionable alerts, meaningful dashboards, and up-to-date runbooks. * Define and implement SLIs and SLOs in partnership with product and engineering teams, and surface reliability performance in regular Cloud Operations reviews. * Identify operational toil and design automation to eliminate repetitive manual work. * Drive continuous improvement initiatives that increase observability, automation coverage, and system resilience. Incident Response & Operational Excellence * Participate in the SRE on-call rotation, progressing from secondary to primary ownership as readiness increases. * Command Sev-2 and Sev-3 incidents independently over time, with pairing and coaching from a Senior SRE; act as technical lead during Sev-1 incidents. * Lead blameless post-incident reviews and own follow-up actions through completion. * Partner closely with managed services providers (IBM, HCL) to ensure clean escalation paths from L1/L2 monitoring into SRE ownership. Platform, CI/CD & Observability * Design, build, and operate CI/CD pipelines supporting cloud-native application delivery using tools such as GitHub Actions, CodeFresh, TeamCity, and Octopus Deploy. * Automate infrastructure and platform services using Infrastructure as Code (Terraform preferred). * Contribute to the evolution of Omnicell's observability platform, including intelligent alerting, ML-based anomaly detection, and automated diagnostics. * Participate in architecture and launch readiness reviews, bringing a reliability lens to system design. * Help establish reference implementations and "golden paths" that enable product teams to launch services with reliability built in from day one., * Collaborate: Partner closely with product engineering, security, and operations teams to build shared ownership of reliability. * Inspire: Influence reliability best practices across teams by modeling calm, structured incident leadership. * Develop: Continuously build your technical depth while learning directly from a senior SRE mentor. * Execute: Take ownership of services, incidents, and follow-through-turning lessons learned into measurable improvements. * Impact: Help shape foundational SRE practices and introduce modern reliability and AIOps capabilities that scale with the business. Growth & Career Path This role is intentionally designed as a growth role . With strong performance and increasing ownership, the natural progression is into a Senior Site Reliability Engineer position as the practice scales. Omnicell also supports lateral growth into platform engineering, security engineering, or product engineering for SREs who discover adjacent passions. Work Conditions * Remote or hybrid work environment supported. * Up to 10% travel as needed. * Participation in an SRE on-call rotation is required. ## Related Videos - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read)