> Markdown version of [/jobs/ext/2697995-site-reliability-engineer-hpc-automation-silicon-engineering](https://www.wearedevelopers.com/jobs/ext/2697995-site-reliability-engineer-hpc-automation-silicon-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SITE RELIABILITY ENGINEER - HPC & AUTOMATION (SILICON ENGINEERING) - **Company:** Space Exploration Technologies Corp. - **Location:** Redmond, WA, United States - **Salary:** $125,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Information Systems, Databases, Continuous Integration, Linux, Python (Programming Language), PostgreSQL, MySQL, NetApp Applications, Data ONTAP (Server Appliance), Release Management, Reliability Engineering, Ansible, Prometheus, SQLite, TCP/IP, High Performance Computing, Computer Network Technologies, Application Specific Integrated Circuits, Large Language Models, Grafana, Concurrency, Containerization, Kubernetes, Information Technology, Build Tools, Slurm, Puppet, Restful APIs, Ansys, Terraform, Software Version Control, Atlassian Bamboo, Docker, Jenkins, Programming Languages - **Published:** September 3, 2026 - **Apply:** https://www.themuse.com/jobs/spacex/site-reliability-engineer-hpc-automation-silicon-engineering ## About the Role * Bachelor's degree in computer science, information systems, or an engineering discipline; OR 2+ years of professional experience in system administration, high performance computing, or site reliability engineering * 1+ years of development experience with Bash, Python, and/or other programming languages * 1+ years of experience with Linux operating systems, * Familiarity with containerization technologies (i.e. Docker, Kubernetes) * Knowledge in computer system concepts (computer architecture, computer organization, operating systems and concurrency) * Experience with databases and data modeling (e.g., MySQL, PostgreSQL, SQLite) * Networking knowledge of TCP/IP * Experience with high performance computing and workload managers (e.g., Slurm, LSF) * Experience with Terraform, Ansible, Puppet, or similar automation frameworks * Experience building monitoring and alerting as code (e.g., Grafana, Prometheus, custom exporters) * Experience with CI/CD automation at scale (e.g., Jenkins, Bamboo, build systems) * Experience with infrastructure as code (IaC) tools for managing fleets of servers * Experience with using & building REST API clients/servers * Experience with enterprise/networked storage automation (e.g., NetApp ONTAP REST API/CLI, NFS) * Experience with ASIC design flows and tools (e.g., Cadence, Synopsys, Ansys, Keysight, Siemens) * Strong desire to find performance bottlenecks and performance improvement techniques * Excellent communication skills with the ability to communicate with customers, peers, management, etc. in both formal and informal situations * Ability to quickly learn new tools and frameworks * Interest in or experience with AI/LLM-assisted tooling (e.g., Grok, Claude Code) ADDITIONAL REQUIREMENTS: * Ability to work extended hours and weekends as needed to meet critical milestones, * To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here. ## Description * Deploy, upgrade, operate, maintain, and scale our suite of clusters and services * Collaborate with engineers to develop automated, full turnkey solutions for silicon simulation workflows to speed up project timelines * Manage our underlying infrastructure as code and use modern observability tools to provide a complete picture of cluster and infrastructure health * Operate the continuous integration pipeline, build and release systems, and version control across the environment * Identify and eliminate performance bottlenecks using measurement and creative engineering ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Cloud Vendor Lock-In - Is it just a new version of the Database Abstraction Layers?](https://www.wearedevelopers.com/videos/1185-cloud-vendor-lock-in-is-it-just-a-new-version-of-the-database-abstraction-layers) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Your next 10x engineer isn't in your city. Refactor accordingly.](https://www.wearedevelopers.com/videos/100103-your-next-10x-engineer-isn-t-in-your-city-refactor-accordingly) - [Colorful quantum randomness](https://www.wearedevelopers.com/videos/100257-colorful-quantum-randomness) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)