> Markdown version of [/jobs/ext/2594056-sr-iam-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2594056-sr-iam-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. IAM Site Reliability Engineer - **Company:** EVERPURE LLC - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $186,000.0 - $279,000.0 - **Contract:** Permanent contract - **Skills:** Bash Shell, Configuration Management, Software Documentation, Identity and Access Management, Information Security Management, Python (Programming Language), Windows PowerShell, Reliability Engineering, Ansible, Prometheus, Datadog, Scripting, System Availability, Grafana, Terraform, Splunk - **Published:** August 10, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=05860dd9b1ab2753 ## About the Role * Production Automation & Infrastructure Skills: Hands-on proficiency operating production-grade enterprise infrastructure, using Infrastructure as Code (Terraform), configuration management (Ansible), and scripting (Python, PowerShell, or Bash) to automate workflows and identity lifecycle management. * Observability & Reliability Engineering: Demonstrated experience implementing telemetry using enterprise monitoring tools (such as Datadog, Prometheus, Grafana, or Splunk) and applying SLAs, SLOs, and operational KPIs to measure and elevate service health. * Incident Management & Root Cause Analysis: Expertise in incident response and structured problem management, with the ability to lead resolution efforts during service disruptions and implement effective preventive controls. * Technical Communication & Operational Documentation: Exceptional written and verbal communication skills to translate complex operations into clear updates for diverse audiences, paired with disciplined habits for creating maintainable runbooks, SOPs, and system documentation. We are primarily an in-office environment and therefore, you will be expected to work from the Santa Clara office in compliance with Pure's policies, unless you are on PTO, or work travel, or other approved leave. ## Description Do you think in playbooks and dashboards? Do you sleep better knowing an IAM platform is healthy, patched, and instrumented end to end? At Everpure, identity is the front door to everything we build, and we're looking for an IAM Site Reliability Engineer to keep that door running smoothly - and to make it better every day. The Global Information Security Office (GISO) at Everpure is seeking an IAM Site Reliability Engineer to operate, automate, and continuously improve our enterprise Identity and Access Management services. This role is centered on keeping IAM platforms reliable and observable, automating operations, and making sure the rest of the organization always knows the state of the systems you run. WHAT YOU'll DO * Enterprise IAM Platform Ownership: Own the end-to-end operation, high availability, and resilient performance of enterprise IAM platforms-including identity providers and lifecycle services-guaranteeing seamless, secure access for global users. * Infrastructure as Code & Workflow Automation: Build and maintain configuration management and automation frameworks using Terraform, Ansible, and Tines to streamline system patching, provisioning, and routine operations, eliminating manual overhead and deployment risk. * Observability, SLIs/SLOs & Incident Response: Establish, track, and tune observability tooling (metrics, logs, alerting) alongside SLIs, SLOs, and operational KPIs, driving proactive issue detection and rapid incident resolution during scheduled on-call rotations. * Change Governance & Cross-Functional Collaboration: Partner across technical and business teams to lead change management risk assessments, deploy new identity capabilities, and maintain actionable runbooks and architecture documentation to ensure transparent, repeatable operations. ## Related Videos - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Eclipse Che for Infrastructure Automation](https://www.wearedevelopers.com/videos/1611-eclipse-che-for-infrastructure-automation) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026)