Senior Site Reliability Engineer, AI Infrastructure
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+6 more
Job description
Experteer Overview In this AI SRE role, you will ensure PointClickCare’s AI platforms run safely, reliably, and efficiently while protecting patient data. You’ll own reliability and security across data processing, ML workspaces, labeling systems, and model serving, using automation, SLOs, and guardrails. You’ll enable data scientists and ML engineers to move fast without compromising stability. You’ll collaborate with research, platform, data, and security teams to improve observability, incident response, and cost efficiency. A meaningful hook is shaping hospital-grade AI infrastructure at scale. Compensation / Benefits * Own service level objectives, error budgets, and reliability targets for AI/ML infra with full observability (metrics, logs, traces) and telemetry * Design, build, and maintain infrastructure-as-code and automation to reduce toil and ensure repeatability * Implement platform security controls (network segmentation, secrets, encryption) aligned to compliance * Lead incident response and blameless postmortems; validate backups and disaster recovery; conduct resiliency testing * Mentor engineers; influence design reviews; improve platform resiliency, cost efficiency, and capacity planning * Collaborate with cross-functional teams across AI engineering, security, platform, and data teams Tasks * 5+ years in SRE, platform engineering, or infrastructure roles * Strong observability and incident management experience * Proficiency with Terraform, GitOps, and CI/CD for infra and platforms * Cloud platform administration experience with AI/ML platforms (Databricks, Azure ML, Kubernetes) * Platform security experience (network segmentation, secrets management, encryption, KMS) * Strong automation programming skills * Clear communication and ability to write runbooks and lead postmortems Key requirements * Benefits starting from Day 1 * Retirement Plan Matching * Flexible Paid Time Off * Wellness Support Programs * Parental & Caregiver Leaves * Continuous Development Support Program
Requirements
secrets response and blameless postmortems; validate backups and disaster recovery; conduct resiliency testing * Mentor engineers; influence design reviews; improve platform resiliency, cost efficiency, and capacity planning * Collaborate with cross-functional teams across AI engineering, security, platform, and data teams Tasks * 5+ years in SRE, platform engineering, or infrastructure roles * Strong observability and incident management experience * Proficiency with Terraform, GitOps, and CI/CD for infra and platforms * Cloud platform administration experience with AI/ML platforms (Databricks, Azure ML, Kubernetes) * Platform security experience (network segmentation, secrets management, encryption, KMS) * Strong automation programming skills * Clear communication and ability to write runbooks and lead postmortems Key requirements * Benefits starting from Day 1 * Retirement Plan Matching * Flexible Paid Time Off * Wellness Support Programs * Parental & Caregiver Leaves * aaaaaaaar _ Development Support Program
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
Dev Digest 137 - AI'm not sure about this
Dev Digest 121 - AI goes offline
Navigating the AI Shift