> Markdown version of [/jobs/ext/3077943-site-reliability-engineer-ai-infrastructure-starshield](https://www.wearedevelopers.com/jobs/ext/3077943-site-reliability-engineer-ai-infrastructure-starshield). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer, AI Infrastructure (Starshield - **Company:** Space Exploration Technologies Corp. - **Location:** Washington, DC, United States - **Salary:** $125,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Booting (BIOS), C++ (Programming Language), Cloud Computing, Information Systems, Databases, System Configuration, Continuous Integration, Linux, DevOps, Distributed Data Store, Make (Software), Python (Programming Language), Reliability Engineering, Ansible, TCP/IP, AI Infrastructure, Computer Networking Systems, System Availability, Containerization, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Bare Metal, Terraform, Oracle Cloud Infrastructure, Network Server - **Published:** September 25, 2026 - **Apply:** https://www.themuse.com/jobs/spacex/site-reliability-engineer-ai-infrastructure-starshield ## About the Role * Bachelor's degree in computer science, information systems/IT, or an engineering discipline and 1+ years of professional experience in site reliability engineering or DevOps; OR 3+ years of professional experience in site reliability engineering or DevOps in lieu of a degree * 1+ years of professional experience with Linux operating systems * Experience with Terraform, Ansible, or other infrastructure tools * Experience with containerization technologies (i.e. OCI containers, Kubernetes) * Experience scripting in Bash, Python, or other similar languages * Development experience in Python, C++, or Go PREFERRED SKILLS AND EXPERIENCE: * 1+ years of experience with Python and Python-based development frameworks * Experience managing Kubernetes clusters, not just using them * Knowledge of Linux boot process and systems configuration * Deep understanding of testing, continuous integration, build, deployment & continuous monitoring * Understanding of relevant build technologies, such as Bazel and Makefiles * Focus on performance bottlenecks and performance improvement techniques * Understanding of distributed databases and data modeling * Experience with automatically managing thousands of servers (eg: Terraform or Ansible) * Strong networking knowledge of TCP/IP * Cloud virtualization experience (as the cloud provider) * Experience using NVIDIA GPU deployment stacks (Blackwell/Rubin) * Excellent communications skills with the ability to communicate with customers, peers, management etc. in both formal and informal situations * Active Top Secret, Top Secret SCI, or DOE Level Q clearance ADDITIONAL REQUIREMENTS: * Must be willing to work extended hours and weekends as needed * Must be willing to travel domestically and globally in the future when needed * This position requires successfully obtaining and maintaining a Top Secret Security Clearance as a condition of employment. While the clearance may not be immediately necessary upon hire, we encourage you to initiate the application process promptly upon accepting this offer. Your ability to secure the necessary clearance is essential for fulfilling key responsibilities of the role. Should you be unable to obtain it, SpaceX reserves the right to modify or terminate your employment to align with operational needs., * To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here. ## Description * Manage GPU/CPU infrastructure deployments to Top Secret datacenters * Manage and provide support for GPU as a service for external customers on bare metal hardware and virtualized platforms * Design, validate, and productize solutions for AI clusters (100k+ GPU scale) * Develop automation to deploy and manage on-premise Kubernetes\AI clusters, and operating systems * Deploy and manage core infrastructure such as databases, monitoring and distributed storage * Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products * Engage in and improve the whole lifecycle of services -- from inception and design, through deployment, operation and refinement * Monitoring and alerting supporting systems to have high availability * Identify areas for improvement and create innovative solutions that enable high system availability ## Related Videos - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)