Senior Site Reliability Engineer, AI Infrastructure

POINTCLICKCARE
Salt Lake City, UT, United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Artificial Intelligence Continuous Integration Disaster Recovery Key Management Network Segmentation Reliability Engineering Azure Machine Learning Runbook AI Infrastructure Automatic Programming Data Processing Cloud Platform System
+6 more
AI Platforms Git Flow Kubernetes Machine Learning Operations Terraform Databricks

Job description

Experteer Overview In this AI SRE role, you will ensure PointClickCare’s AI platforms run safely, reliably, and efficiently while protecting patient data. You’ll own reliability and security across data processing, ML workspaces, labeling systems, and model serving, using automation, SLOs, and guardrails. You’ll enable data scientists and ML engineers to move fast without compromising stability. You’ll collaborate with research, platform, data, and security teams to improve observability, incident response, and cost efficiency. A meaningful hook is shaping hospital-grade AI infrastructure at scale. Compensation / Benefits * Own service level objectives, error budgets, and reliability targets for AI/ML infra with full observability (metrics, logs, traces) and telemetry * Design, build, and maintain infrastructure-as-code and automation to reduce toil and ensure repeatability * Implement platform security controls (network segmentation, secrets, encryption) aligned to compliance * Lead incident response and blameless postmortems; validate backups and disaster recovery; conduct resiliency testing * Mentor engineers; influence design reviews; improve platform resiliency, cost efficiency, and capacity planning * Collaborate with cross-functional teams across AI engineering, security, platform, and data teams Tasks * 5+ years in SRE, platform engineering, or infrastructure roles * Strong observability and incident management experience * Proficiency with Terraform, GitOps, and CI/CD for infra and platforms * Cloud platform administration experience with AI/ML platforms (Databricks, Azure ML, Kubernetes) * Platform security experience (network segmentation, secrets management, encryption, KMS) * Strong automation programming skills * Clear communication and ability to write runbooks and lead postmortems Key requirements * Benefits starting from Day 1 * Retirement Plan Matching * Flexible Paid Time Off * Wellness Support Programs * Parental & Caregiver Leaves * Continuous Development Support Program

Requirements

secrets response and blameless postmortems; validate backups and disaster recovery; conduct resiliency testing * Mentor engineers; influence design reviews; improve platform resiliency, cost efficiency, and capacity planning * Collaborate with cross-functional teams across AI engineering, security, platform, and data teams Tasks * 5+ years in SRE, platform engineering, or infrastructure roles * Strong observability and incident management experience * Proficiency with Terraform, GitOps, and CI/CD for infra and platforms * Cloud platform administration experience with AI/ML platforms (Databricks, Azure ML, Kubernetes) * Platform security experience (network segmentation, secrets management, encryption, KMS) * Strong automation programming skills * Clear communication and ability to write runbooks and lead postmortems Key requirements * Benefits starting from Day 1 * Retirement Plan Matching * Flexible Paid Time Off * Wellness Support Programs * Parental & Caregiver Leaves * aaaaaaaar _ Development Support Program

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

5:02 min

Mapping Git flow branches to application tester segments

Majid Hajian · LIVE

2:50 min

Introduction and the value of runbooks

Hila Fish · WWC 2023

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

1:06 min

Empowering site reliability engineers with integrated AI agents

Osmar Matos Osmar Matos · WWC Europe 2026

3:53 min

Introduction to git flow and clean feature branches

Johannes Haux · WWC 2022

Videos

See all

Related articles

See all