> Markdown version of [/jobs/ext/2722662-lead-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2722662-lead-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Site Reliability Engineer - **Company:** KONTAKT LLC - **Location:** New York, NY, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Health Informatics, Software as a Service, Cloud Computing, Cloud Engineering, Continuous Integration, Disaster Recovery, Distributed Systems, Fault Tolerance, Identity and Access Management, Interoperability, Network Security, Reliability Engineering, Prometheus, Datadog, Data Logging, Fast Healthcare Interoperability Resources, Grafana, Mttr, Event Driven Architecture, Kubernetes, Deployment Automation, Performance Monitor, Health Level Seven International, Terraform, Stream Processing, Data Pipelines, Docker - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/lead-site-reliability-engineer-kontakt-io-8566149 ## About the Role * 10+ years of experience in Site Reliability Engineering or Cloud Infrastructure. * Proven success scaling high-traffic, mission-critical platforms in SaaS, IoT, or healthcare. * Deep expertise in cloud platforms (AWS), Kubernetes, and distributed systems. * Strong background in monitoring, logging, and observability with Prometheus, OpenTelemetry, or similar tools. * Hands-on experience with incident management, postmortems, and building resilient systems. * Deep knowledge of CI/CD automation, GitOps, and infrastructure as code (Terraform, etc.). * A mature leadership approach, with the ability to drive technical strategy while growing and mentoring a high-performance SRE team. * Strong understanding of network security, access management, and compliance frameworks (HIPAA, SOC 2). Bonus Points If You Have: * Experience with healthcare IT, including EHR data, FHIR, and HL7 interoperability. * Expertise in real-time distributed systems, event-driven architectures, or large-scale data pipelines. * Prior experience leading on-call rotations and major incident management processes. ## Description * Ensure 99.99% uptime across our cloud platform, meeting strict SLAs for healthcare customers. * Design and implement self-healing, fault-tolerant systems to prevent failures before they happen. * Define SLIs, SLOs, and SLAs, ensuring proactive performance monitoring and incident resolution. * Architect and manage scalable cloud infrastructure (AWS) for massive real-time data processing. * Optimize containerized environments (Kubernetes, Docker) to support multi-region deployments. * Lead the adoption of infrastructure as code (Terraform) to fully automate infrastructure management. * Build and refine a world-class monitoring, alerting, and logging system using Prometheus, Grafana, OpenTelemetry, and Datadog. * Lead incident response and on-call operations, reducing mean time to detection (MTTD) and mean time to resolution (MTTR). * Conduct blameless postmortems and continuously improve system resilience. * Reduce manual intervention through automated deployment, scaling, and failover mechanisms. * Partner with Security & Compliance teams to ensure infrastructure meets HIPAA and SOC 2 standards. * Lead disaster recovery and business continuity planning to ensure critical healthcare services are always available. * Drive technical strategy and roadmap for scalability, monitoring, and reliability engineering. * Collaborate with Product, Engineering, and Infrastructure teams to align SRE initiatives with business priorities. ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025)