Site Reliability Engineer

Enzo Health, Inc.
Lehi, UT, United States
5 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Kubernetes Security Amazon Web Services Build Automation Computer Programming Databases Continuous Integration Github Key Management PostgreSQL Reliability Engineering Runbook Software Vulnerability Management
+7 more
Backup and Restore Policy as Code Scripting Grafana Database Migration Kubernetes Terraform

Job description

Enzo builds software for home health care teams. Our systems support important clinical and business workflows, so they must be available, secure, and easy to operate., We are hiring a Senior Site Reliability Engineer to join our Security and Site Reliability team. You will focus on a stable, scalable AWS and Kubernetes platform, reliable Postgres operations, and safer delivery. You will also help mature observability, on-call, and incident response practices across engineering.

This is a hands-on role. You will diagnose production problems, improve infrastructure, write automation, and help product engineers operate their services with confidence. You will make our systems safer without slowing product delivery.

What you’ll do

Strengthen the production platform

  • Operate and improve our AWS and Kubernetes environments.
  • Own infrastructure changes through Terraform, including modules, state, review standards, and drift control.
  • Improve environment management, cluster practices, and deployment reliability.
  • Make CI/CD and releases safer through clear checks, reliable rollbacks, environment consistency, and release observability.
  • -mprove capacity planning, resource controls, resilience, and cost visibility.
  • Build automation that removes repetitive operational work and reduces avoidable failures.

Improve Postgres reliability

  • Own the operational health of Postgres in production.
  • Verify backups and regularly test restore procedures against agreed recovery targets.
  • Improve monitoring for connections, storage, slow queries, locks, and other important failure signals.
  • Guide safe schema migrations, database access, and production change procedures.
  • Find performance and capacity risks before they affect customers.
  • Maintain clear database runbooks for common failures and emergency work.

Mature production operations

  • Improve dashboards, monitors, logs, traces, and alert routing.
  • Validate existing service-level indicators and objectives, then close important coverage gaps.
  • Participate in and improve the company-wide on-call rotation.
  • Write practical runbooks and make escalation paths clear for engineers across the company.
  • Lead or support incident response, recovery, and blameless post-incident reviews.
  • Use incident and reliability data to set priorities and measure improvement.

Work across Security and engineering

  • Work with Security to keep infrastructure changes consistent with SOC 2 and health care security requirements.
  • Apply least-privilege access, secure defaults, secrets management, encryption, and auditable change practices.
  • Support vulnerability remediation, disaster-recovery exercises, and secure production access.
  • Help product teams include reliability and operational risk in technical decisions.

What success looks like in the first six months

  • Kubernetes, AWS, and Terraform have clear operating standards and fewer manual failure points.
  • Deployments are observable, repeatable, and easy to roll back.
  • Postgres backups and restores are tested, and database health and capacity risks are visible.
  • Database migrations and emergency access follow safe, documented procedures.
  • Production alerts are useful and actionable, with clear ownership and less noise.
  • On-call responders have the runbooks, access, and escalation paths that they need.
  • Production incidents result in tracked corrective work and measurable reliability gains., This is an in-office role in Lehi, Utah. You will be part of the Security and Site Reliability team and work closely with engineering leads and product engineers. The role includes participation in the engineering on-call rotation.

Requirements

  • At least 5 years of experience in site reliability, platform, infrastructure, or production engineering.
  • Strong hands-on experience with AWS and production Kubernetes.
  • Strong experience with Terraform and infrastructure as code.
  • Experience operating Postgres in production, including backup and restore, performance, and safe migrations.
  • Experience building and operating CI/CD and deployment systems.
  • Experience with modern observability tools across metrics, logs, and traces.
  • Strong programming or scripting skills for automation and operational tools.
  • A strong record of diagnosing production failures and leading incidents through recovery.
  • Ability to write clear automation, runbooks, and technical standards.
  • Ability to work independently and partner well with application engineers.

Nice to have

  • Experience with Kubernetes security controls, policy as code, and software supply-chain security.
  • Experience in health care or another regulated environment.
  • Experience supporting SOC 2 controls or audit evidence.
  • Experience improving cloud cost and capacity efficiency.

Our stack

  • AWS, Kubernetes, and Terraform
  • Postgres
  • GitHub-based CI/CD

Benefits & conditions

  • Competitive salary and meaningful equity
  • 401k & insurance (medical, dental, vision, HSA, and more)
  • High ownership and the ability to shape the security function from day one
  • Direct collaboration with founders and engineering leadership
  • A fast-paced, product-driven engineering culture
  • The opportunity to defend technology that meaningfully improves healthcare operations

About the company

Enzo Health is a healthcare technology company transforming home health operations through purpose-built artificial intelligence. We deliver a secure, HIPAA-compliant AI platform that unifies intake, clinical documentation, coding, and quality assurance-enabling agencies to reclaim time and revenue while elevating patient care.

Enzo addresses the critical challenges facing home health agencies today: rising operational costs, clinician burnout, shrinking reimbursement margins, and increasing compliance demands. Our integrated AI solution automates documentation workflows from referral to final QA, allowing clinical staff to focus on delivering exceptional patient care.

Our Solutions

  • Enzo Intake: Delivers intake decisions in seconds, automatically extracting key data from referrals to increase admissions and reduce processing delays.
  • Enzo Scribe: Auto-generates OASIS documentation, clinical narratives, and care plans, reducing documentation time by up to 75%.
  • Enzo QA: Ensures documentation meets the highest clinical standards with approximately 95% coding accuracy, reducing compliance risk while increasing reimbursement by an average of $185 per episode.

Our Impact

Trusted by top-performing home health agencies nationwide, Enzo delivers measurable results: documentation time under 25 minutes per visit, referral intake under 5 minutes, and 30-50% savings per episode of care. These efficiencies effectively double staff capacity while maintaining exceptional quality and compliance.

As reimbursement pressures intensify, Enzo Health empowers agencies to navigate cost-cutting measures without compromising care quality, positioning AI as the essential strategy for sustainable growth in home health.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

2:50 min

Introduction and the value of runbooks

Hila Fish · World Congress 2023

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

1:32 min

Structuring automated incident workflows between runbooks and raw models

Aram Hakobyan Aram Hakobyan +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all