> Markdown version of [/jobs/ext/624467-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/624467-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** HavocAI Inc - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $150,000.0 - $185,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Systems Engineering, Build Automation, Cloud Computing, Computer Programming, Continuous Integration, Information Engineering, Linux, DevOps, Disaster Recovery, Distributed Systems, Python (Programming Language), Key Management, Reliability Engineering, Cloud Services, Prometheus, Runbook, Datadog, Data Logging, Pulumi, Scripting, Cloud Platform System, Real Time Systems, Delivery Pipeline, Grafana, Reliability of Systems, Containerization, Kubernetes, Infrastructure Automation Frameworks, Build Tools, Terraform, Data Pipelines - **Published:** June 12, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=3cc9cd3e5e9e2a26 ## About the Role Do you have experience in Team leadership?, The ideal candidate is deeply technical, calm under pressure, and experienced in owning reliability outcomes end to end., * 7+ years of experience in SRE, infrastructure engineering, systems engineering, or related roles * Strong experience operating large-scale distributed production systems * Deep understanding of Linux systems, networking, cloud infrastructure, and distributed systems fundamentals * Hands-on experience with Kubernetes and container orchestration * Programming or scripting experience in Go, Python, or similar languages * Experience designing and operating observability systems for production environments * Proven ability to lead incident response and drive reliability improvements * Strong communication skills and ability to collaborate across engineering teams * Ability to operate calmly and effectively under pressure * Must be a U.S. Citizen and eligible to obtain a U.S. Government security clearance if required, * Experience supporting autonomy, robotics, simulation, real-time systems, or data-intensive platforms * Familiarity with AWS and large-scale cloud infrastructure * Experience with chaos engineering, fault injection, or resilience testing * Knowledge of CI/CD systems and progressive delivery practices * Experience working in high-reliability, safety-critical, defense, or mission-critical environments * Experience with Infrastructure as Code tools such as Terraform or Pulumi * Experience with Prometheus, Grafana, OpenTelemetry, Datadog, ELK/OpenSearch, or similar observability tools ## Description HavocAI is seeking a Senior Site Reliability Engineer with 7+ years of experience designing, operating, and scaling highly reliable distributed systems. In this role, you will serve as a key technical leader within the Cloud Platform team, responsible for ensuring the availability, performance, and resilience of mission-critical services supporting autonomy, simulation, and data-intensive workloads., You will work closely with Cloud Platform, DevOps, Data Engineering, and Autonomy teams to establish reliability standards, improve operational maturity, and build systems that scale safely under real-world conditions., * Design and evolve reliability architecture for distributed and cloud-hosted systems * Define and implement SRE best practices, including SLIs, SLOs, error budgets, and capacity planning * Partner with platform and application teams to design systems for reliability, scalability, and operability * Identify and mitigate systemic reliability risks across infrastructure, applications, services, and data pipelines * Establish reliability patterns that support autonomy, simulation, and mission-critical cloud workloads, * Lead incident response processes, including on-call rotations, escalation paths, and post-incident reviews * Conduct root cause analysis for complex production incidents and drive long-term corrective actions * Improve operational readiness through runbooks, automation, resilience testing, and production-readiness reviews * Reduce operational toil through tooling, automation, and process improvements * Help build a culture of ownership, accountability, and continuous improvement across production systems, * Design, implement, and maintain observability systems for metrics, logging, tracing, alerting, and service health * Ensure services and data pipelines are observable, debuggable, and performant in production * Drive performance analysis and tuning across infrastructure, application, and service layers * Improve alert quality, reduce noise, and ensure operational signals are actionable * Partner with engineering teams to define meaningful reliability and performance metrics, * Build automation to improve system reliability, deployment safety, and recovery processes * Partner with DevOps and Cloud Platform teams on CI/CD reliability, rollout strategies, and safe deployment patterns * Support and improve Kubernetes-based environments and containerized workloads * Contribute to infrastructure-as-code practices and platform automation * Help define operational standards for cloud infrastructure, deployment workflows, and production services, * Collaborate with security teams to ensure secure and resilient system design * Participate in disaster recovery planning, backup strategy, and resilience testing * Maintain strong operational practices around access control, secrets management, change management, and production access * Support secure operations for systems that may serve defense, autonomy, or mission-sensitive use cases, A successful Senior Site Reliability Engineer at HavocAI will raise the reliability, performance, and operational maturity of the systems that support our autonomy and cloud platform work. You will help ensure that mission-critical services are observable, resilient, scalable, and recoverable. You will bring structure to incident response, reduce operational toil, improve deployment safety, and partner with engineering teams to design systems that can handle real-world operational demands. This role is ideal for someone who combines deep systems expertise with strong ownership, practical judgment, and a bias toward building durable solutions. ## Related Videos - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)