Senior Site Reliability Engineer, CloudOps

ICU Medical
Columbus, OH, United States
1 day ago
Apply on www.jofdav.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Amazon Web Services Amazon Cloudfront Amazon Elastic Compute Cloud Amazon S3 ARM Architecture Bash Shell Ubuntu (Operating System) Cloud Computing Cloud Computing Security Databases Continuous Delivery
+27 more
Data Stores Linux DevOps Identity and Access Management Python (Programming Language) Linux System Administration Network Segmentation Reliability Engineering Software Deployment Zabbix Datadog Delivery Pipeline Spring-boot Cloudformation Fastapi Containerization Git Flow Kubernetes Infrastructure Automation Frameworks AWS Fargate Cloud Migration Functional Programming Cloudwatch Api Gateway Serverless Computing Docker Jenkins

Job description

We are seeking a Senior Site Reliability Engineer, CloudOps to support, scale, and optimize a multi-account AWS environment hosting healthcare-oriented applications and analytics platforms. In this role, you will bridge infrastructure engineering, operational reliability, production support, and cloud modernization initiatives across complex microservices architectures. The ideal candidate brings strong AWS expertise, solid Linux administration skills, and a proven track record of managing production systems in HIPAA/HiTrust regulated environments. You will participate in incident response, on-call rotations, and continuous deployment workflows while helping drive our transition toward containerized and Kubernetes-based platforms. This collaborative position is built for an analytical engineer who excels at resolving production incidents, partnering with developers, and continuously elevating operational excellence., * Manage, maintain, and troubleshoot a multi-account AWS Organization environment (35+ accounts) and core services, including EC2, ECS/Fargate, Lambda, S3, CloudFront, API Gateway, and Aurora/RDS databases.

  • Support production deployments, CI/CD pipelines (Jenkins, AWS CodePipeline), and infrastructure automation using Python, Bash, and AWS CloudFormation.
  • Monitor system health and performance using Datadog, CloudWatch, and Zabbix; investigate alerts, execute root-cause analysis, and refine monitoring coverage to reduce operational noise.
  • Participate in a shared on-call rotation, managing incident response and performing failover/recovery validation for production applications and data stores.
  • Maintain HIPAA/HiTrust compliance and security posture by managing tools like Prisma/Cortex Cloud, Security Hub, and GuardDuty, while enforcing proper IAM policies and network segmentation.
  • Support Java (Spring Boot) and Python applications running in containers, assisting developers during investigations and preparing for future Kubernetes (EKS) modernization initiatives.

Requirements

  • Deep hands-on expertise with AWS core services (networking, compute, serverless, and database technologies) and CloudFormation IaC automation.
  • Strong Linux administration skills (primarily Ubuntu) along with proficiency in Python and Bash scripting for operational automation.
  • Experience with containerization technologies (Docker, ECS/Fargate) and familiarity with modern Kubernetes ecosystems (EKS, Helm, ArgoCD).
  • Solid understanding of observability tools (Datadog, CloudWatch, Zabbix) and CI/CD pipelines (Jenkins, CodePipeline, Git workflows).
  • Knowledge of cloud security best practices, access management (IAM), and compliance frameworks within regulated sectors (HIPAA/HiTrust).
  • Proven diagnostic, incident-management, and analytical troubleshooting skills for complex microservices architectures., * Must be at least 18 years of age.
  • High School Diploma required.
  • Bachelor’s degree from an accredited college or university is required.
  • 7+ years of hands-on experience in AWS Cloud Engineering, DevOps, Site Reliability Engineering (SRE), or Infrastructure Engineering.
  • Practical background supporting production workloads in Linux/AWS environments, reading application logs, and making minor code fixes.
  • Direct experience participating in on-call rotations and incident response protocols.
  • Prior experience in the healthcare industry maintaining HIPAA/HiTrust-compliant infrastructure.

About the company

ICU Medical has consistently provided you with clinical innovations that help solve real-world challenges.

With the acquisition of Hospira Infusion Systems in 2017 and Smiths Medical in 2022, we are now a global market leader with a complete line of clinically-essential IV therapy and high-value critical care products for hospital, alternate site, and home care settings.

We’re ready to bring you consistent quality, innovation, and value in more areas than ever. Our focus allows us to bring you:

  • Dedicated and non-dedicated IV sets and needlefree connectors clinically proven to provide an effective barrier against bacterial transfer and colonization.
  • The industry’s broadest IV smart pump offering covering large volume, pain management, and ambulatory needs.
  • IV medication safety software providing full IV-EHR interoperability with the highest customer satisfaction and compatibility with more EHR systems than any other company.
  • Significant US IV solutions manufacturing and supply capabilities.

This role is based remotely; the incumbent may be remote in any state other than Colorado; California; Connecticut; Montana, Maine or New York.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jofdav.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:37 min

Why differing legacy workflows complicate monitoring tool migrations

Mathias Palmersheim Mathias Palmersheim · Europe 2026 Virtual

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all