> Markdown version of [/jobs/ext/3053114-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3053114-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Space X - **Location:** Hawthorne, CA, United States - **Experience:** Experienced - **Salary:** $125,000.0 - $160,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Computational Fluid Dynamics, Configuration Management, Nvidia CUDA, Continuous Integration, Software Debugging, Linux, DevOps, Firmware, Python (Programming Language), Machine Learning, Nagios, Reliability Engineering, Ansible, Tensorflow, Prometheus, Scientific Computating, Scripting, Pytorch, Grafana, Kubernetes, Information Technology, Slurm, Puppet, Terraform, Network Server, Software Version Control, Docker - **Published:** September 24, 2026 - **Apply:** https://www.thejobnetwork.com/job/56c863e9-7c08-4ea0-ba19-f44254292635/site-reliability-engineer-high-performance-computing ## About the Role - Bachelor's degree in computer science, engineering, math, or a scientific discipline; OR 2+ years of professional experience operating production infrastructure in lieu of a degree - 2+ years of experience with Linux operating systems in production - 2+ years of experience operating production infrastructure (servers, services, or networks), including monitoring, debugging, and repairing what you own **PREFERRED SKILLS AND EXPERIENCE:** - 2+ years of professional experience in SRE, DevOps, or production infrastructure engineering - Experience with monitoring and alerting (Prometheus, Grafana, Nagios, or similar) - Experience deploying and maintaining configuration management or infrastructure as code (Ansible, Puppet, Terraform, or similar) - Experience writing scripts/code (eg. Python or similar languages) to automate common tasks - Experience with containers (Docker, Podman, Singularity/Apptainer) - Experience with Kubernetes administration for on-premise deployment - Experience with distributed or high-performance storage (VAST or similar), including capacity, performance, and lifecycle management - Familiarity with HPC clusters, schedulers (Slurm, PBS, LSF), or GPU compute - not required; we will teach this - Familiarity with scientific computing (CFD, FEA) and/or ML training workloads (PyTorch, TensorFlow, CUDA) and/or AI inference workloads - Good understanding of version control, testing, continuous integration, build, deployment and monitoring - Ability to communicate clearly with users, peers, and vendors in both incident and design settings - Comfortable working with mission-critical and sensitive systems, with a sense of urgency appropriate to the responsibilities - Eligibility for access to classified material up to TS/SCI with polygraph ## Description We are looking for a Site Reliability Engineer who wants to own everything from Linux machines and our Infrastructure as Code, storage, and user facing applications - the whole ecosystem as a product, not as a ticket queue. You do not need a prior HPC title. You do need production instincts - you have operated real infrastructure, you write code to delete toil, and you care about whether users can actually get work done, not just whether nodes ping. You'll work alongside HPC systems engineers who design and commission clusters to help make them more reliable and provide world class services for world class engineers. Aerospace experience is not required. We value engineers who treat teammates with fairness and respect, who are self-critical, and who will take ownership of hard production problems. **RESPONSIBILITIES:** - Participate in the team's on-call rotation; practice sustainable incident response and blameless postmortems - Manage node lifecycle with infrastructure as code: OS images, firmware, configuration management, kernel and driver stack - Build observability for both HPC administrators and end users - cluster, node, and storage health for operators, and job/workflow-level signal for the people running work on the platform - Reduce toil with automation; split time between operating production systems and writing the software that makes that work smaller - Sustainably manage resources, including compute and storage - Lead capacity planning with users across the company: understand what they will need next, and turn that into a concrete picture of tomorrow's compute and storage - Collaborate with HPC systems engineers and with engineers across all disciplines across the company on operable, maintainable infrastructure ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering)