> Markdown version of [/jobs/ext/2001640-senior-site-reliability-engineer-ai-infrastructure](https://www.wearedevelopers.com/jobs/ext/2001640-senior-site-reliability-engineer-ai-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer, AI Infrastructure - **Company:** POINTCLICKCARE - **Location:** Salt Lake City, UT, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Continuous Integration, Disaster Recovery, Key Management, Network Segmentation, Reliability Engineering, Azure Machine Learning, Runbook, AI Infrastructure, Automatic Programming, Data Processing, Cloud Platform System, AI Platforms, Git Flow, Kubernetes, Machine Learning Operations, Terraform, Databricks - **Published:** August 9, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/senior-site-reliability-engineer-ai-infrastructure-salt-lake-city-ut-usa-58870965 ## About the Role secrets response and blameless postmortems; validate backups and disaster recovery; conduct resiliency testing * Mentor engineers; influence design reviews; improve platform resiliency, cost efficiency, and capacity planning * Collaborate with cross-functional teams across AI engineering, security, platform, and data teams Tasks * 5+ years in SRE, platform engineering, or infrastructure roles * Strong observability and incident management experience * Proficiency with Terraform, GitOps, and CI/CD for infra and platforms * Cloud platform administration experience with AI/ML platforms (Databricks, Azure ML, Kubernetes) * Platform security experience (network segmentation, secrets management, encryption, KMS) * Strong automation programming skills * Clear communication and ability to write runbooks and lead postmortems Key requirements * Benefits starting from Day 1 * Retirement Plan Matching * Flexible Paid Time Off * Wellness Support Programs * Parental & Caregiver Leaves * aaaaaaaar _ Development Support Program ## Description Experteer Overview In this AI SRE role, you will ensure PointClickCare's AI platforms run safely, reliably, and efficiently while protecting patient data. You'll own reliability and security across data processing, ML workspaces, labeling systems, and model serving, using automation, SLOs, and guardrails. You'll enable data scientists and ML engineers to move fast without compromising stability. You'll collaborate with research, platform, data, and security teams to improve observability, incident response, and cost efficiency. A meaningful hook is shaping hospital-grade AI infrastructure at scale. Compensation / Benefits * Own service level objectives, error budgets, and reliability targets for AI/ML infra with full observability (metrics, logs, traces) and telemetry * Design, build, and maintain infrastructure-as-code and automation to reduce toil and ensure repeatability * Implement platform security controls (network segmentation, secrets, encryption) aligned to compliance * Lead incident response and blameless postmortems; validate backups and disaster recovery; conduct resiliency testing * Mentor engineers; influence design reviews; improve platform resiliency, cost efficiency, and capacity planning * Collaborate with cross-functional teams across AI engineering, security, platform, and data teams Tasks * 5+ years in SRE, platform engineering, or infrastructure roles * Strong observability and incident management experience * Proficiency with Terraform, GitOps, and CI/CD for infra and platforms * Cloud platform administration experience with AI/ML platforms (Databricks, Azure ML, Kubernetes) * Platform security experience (network segmentation, secrets management, encryption, KMS) * Strong automation programming skills * Clear communication and ability to write runbooks and lead postmortems Key requirements * Benefits starting from Day 1 * Retirement Plan Matching * Flexible Paid Time Off * Wellness Support Programs * Parental & Caregiver Leaves * Continuous Development Support Program ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Why Git Still Matters](https://www.wearedevelopers.com/videos/100288-why-git-still-matters) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix)