> Markdown version of [/jobs/ext/623219-site-reliability-engineer-kubernetes-platform-starshield](https://www.wearedevelopers.com/jobs/ext/623219-site-reliability-engineer-kubernetes-platform-starshield). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer, Kubernetes Platform (Starshield) - **Company:** Space Exploration Technologies Corp. - **Location:** Hawthorne, CA, United States - **Salary:** $125,000.0 - $175,000.0 - **Contract:** Permanent contract - **Skills:** Bash Shell, Booting (BIOS), C++ (Programming Language), Information Systems, Databases, System Configuration, Continuous Integration, Linux, DevOps, Distributed Data Store, Make (Software), Python (Programming Language), Reliability Engineering, Ansible, TCP/IP, Computer Networking Systems, System Availability, Containerization, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Terraform, Oracle Cloud Infrastructure - **Published:** June 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=bf86045a3a7c4154 ## About the Role Do you have experience in Tooling?, * Bachelor's degree in computer science, information systems/IT, or an engineering discipline and 1+ years of professional experience in site reliability engineering or DevOps; OR 3+ years of professional experience in site reliability engineering or DevOps in lieu of a degree * 1+ years of professional experience with Linux operating systems * Experience with Terraform, Ansible, or other infrastructure tools * Experience with containerization technologies (i.e. OCI containers, Kubernetes) * Experience scripting in Bash, Python, or other similar languages * Development experience in Python, C++, or Go PREFERRED SKILLS AND EXPERIENCE: * 1+ years of experience with Python and Python-based development frameworks * Experience managing Kubernetes clusters, not just using them * Knowledge of Linux boot process and systems configuration * Deep understanding of testing, continuous integration, build, deployment & continuous monitoring * Understanding of relevant build technologies, such as Bazel and Makefiles * Focus on performance bottlenecks and performance improvement techniques * Understanding of distributed databases and data modeling * Experience with automatically managing dozens, hundreds, or thousands of servers (eg: Terraform or Ansible) * Strong networking knowledge of TCP/IP * Excellent communications skills with the ability to communicate with customers, peers, management etc. in both formal and informal situations * Active Top Secret, Top Secret SCI, or DOE Level Q clearance ## Description * Deploy and manage core infrastructure such as databases, monitoring and distributed storage * Closely collaborate with software engineers to create highly scalable, operable, and maintainable products Engage in and improve the whole lifecycle of services * - from inception and design, through deployment, operation and refinement * Monitoring and alerting supporting systems to have high availability * Hands-on integration and troubleshooting across the entire Starshield stack * Identify areas for improvement and create innovative solutions that enable high system availability, * Must be willing to work extended hours and weekends as needed * This position requires successfully obtaining and maintaining a Top Secret Security Clearance as a condition of employment. While the clearance may not be immediately necessary upon hire, we encourage you to initiate the application process promptly upon accepting this offer. Your ability to secure the necessary clearance is essential for fulfilling key responsibilities of the role. Should you be unable to obtain it, SpaceX reserves the right to modify or terminate your employment to align with operational needs. ## Related Videos - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [From Space to Software: Reliability Lessons 40 Years After Challenger](https://www.wearedevelopers.com/videos/100283-from-space-to-software-reliability-lessons-40-years-after-challenger) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)