> Markdown version of [/jobs/ext/543577-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/543577-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Harbor Compliance - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $110,761.0 - $158,230.0 - **Contract:** Permanent contract - **Skills:** Computer Programming, Continuous Integration, Data Loss, Linux, Disaster Recovery, Identity and Access Management, Virtual Private Networks (VPN), Python (Programming Language), MySQL, Software Architecture, Reliability Engineering, Ansible, Prometheus, Software Engineering, Datadog, Scripting, Delivery Pipeline, Reliability of Systems, Firewalls (Computer Science), Kubernetes, Terraform, New Relic (SaaS), Dynatrace - **Published:** June 10, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a6cdd1a3ccaef8a4 ## About the Role Do you have experience in Tooling?, * 4-7 years of professional experience building and managing resilient, modern infrastructure within a fast-paced environment. * Expert-level proficiency in managing and troubleshooting Linux-based servers across multiple distributions. * Advanced capability in developing modular, reusable infrastructure templates using tools such as Terraform and Ansible. * Proven success in managing containerized workloads at scale using Kubernetes and Helm. * Extensive experience configuring and optimizing high-performance database environments, specifically MySQL. * Demonstrated ability to build robust, secure CI/CD deployment pipelines that include automated rollback and quality gates. * Strong technical documentation skills, including the creation of architectural diagrams, detailed specifications, and operational playbooks. * Ability to lead cross-functional projects independently while mentoring junior engineers and driving team-wide initiatives. Skills and Knowledge: * Deep understanding of observability platforms such as New Relic, Datadog, or Prometheus to measure and improve system reliability. * Expertise in designing secure cloud networking strategies including firewalls, VPNs, and identity management best practices. * Advanced scripting and programming proficiency in Python or similar languages to automate complex operational workflows. * Strategic insight into infrastructure ROI and the ability to align technical roadmaps with broad business priorities. * Practical knowledge of disaster recovery planning and the execution of failure-resilient system designs. ## Description The Site Reliability Engineer is a senior-level technical leader responsible for the proactive design, implementation, and predictable management of our business-critical Linux infrastructure. You will collaborate cross-functionally with Software Development and technical stakeholders to execute resilient infrastructure strategies that support high-growth business goals. Success in this role is defined by the successful delivery of scalable technical solutions and the consistent maintenance of exceptional system performance and reliability., * Design and execute a comprehensive infrastructure strategy that proactively supports evolving business requirements and operational excellence. * Own the predictable delivery of high-complexity technical solutions through deep automation using Kubernetes and sophisticated CI/CD pipelines. * Maintain superior portal availability and system health by implementing advanced observability and distributed tracing strategies. * Lead high-severity incident response efforts and drive systemic improvements through insightful, blameless postmortem analysis. * Architect failure-resilient and self-healing infrastructure systems to ensure continuous operational stability and zero data loss. * Serve as the internal subject matter expert to influence software architecture decisions toward maximum scalability and performance. * Facilitate regular knowledge-sharing and training sessions to elevate technical standards and process predictability across the entire technology department. * Direct security initiatives and design secure networking strategies to maintain a high-standard protection framework for all client data and assets. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [The OpenTelemetry mistakes I keep seeing (and how to stop making them)](https://www.wearedevelopers.com/videos/100158-the-opentelemetry-mistakes-i-keep-seeing-and-how-to-stop-making-them) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)