> Markdown version of [/jobs/ext/2155549-software-engineer-inference-ai-data-engineering](https://www.wearedevelopers.com/jobs/ext/2155549-software-engineer-inference-ai-data-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Inference (AI Data Engineering) - **Company:** Space Exploration Technologies Corp. - **Location:** Palo Alto, CA, United States - **Experience:** Experienced - **Salary:** $135,000.0 - $175,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Software Applications, Application Performance Management, C++ (Programming Language), Code Generation, Computer Programming, Databases, Continuous Delivery, Continuous Integration, Information Engineering, Distributed Systems, Internet Services, Python (Programming Language), PostgreSQL, MongoDB, Data Streaming, System Programming, AI Infrastructure, Load Balancing, Large Language Models, Multi-Agent Systems, Caching, Parallel Computation, Backend, Rate Limiting, Containerization, Kubernetes, Information Technology, Low Latency, Build Tools, TensorRT, Vertica, Software Version Control, Docker - **Published:** August 20, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/88089235/1 ## About the Role * Bachelor's degree in computer science, engineering, math, or scientific discipline; OR 2+ years of professional experience building software in lieu of a degree * Experience in designing, implementing, and maintaining reliable and horizontally scalable distributed systems * 1+ years of experience in full stack development or backend development with production systems * 1+ years of experience with Rust or C++, * Experience with LLM inference engines and serving frameworks (e.g., SGLang, vLLM, Triton, TensorRT-LLM) * Deep low-level systems programming and optimizations: GPU kernels, code generation, batching, caching, parallelism, quantization, and speculative decoding * Experience with large-scale, high-concurrency production serving systems * Knowledge of service observability and reliability best practices * Experience operating commonly used databases such as PostgreSQL, ClickHouse, or MongoDB * Experience designing or building with agent SDKs and agent orchestration frameworks * Experience with Docker, Kubernetes, and containerized applications * Expert knowledge of gRPC (unary, response streaming, bi-directional streaming, REST mapping) * Programming experience in Python, Go, or similar languages * Experience with version control, continuous integration, continuous delivery, build systems, and monitoring * Expertise in profiling and improving application performance ## Description The application software team is the central nervous system of SpaceX - we create mission critical applications that are used throughout SpaceX to accelerate launch vehicle production and flight as well as systems that allow Starlink to grow into a worldwide fast, reliable Internet service. We are looking for engineers who treat fellow teammates with fairness, respect, and support. Our team maintains a high-performance AI inference platform that serves the best models internally at SpaceX to accelerate our most ambitious engineering goals. As part of this effort in Palo Alto, you will design and optimize large-scale model serving systems end-to-end, owning everything from distributed infrastructure to deep low-level optimizations. You will work on systems that deliver reliable, high-throughput inference to power SpaceX's mission-critical applications while maintaining the highest standards of performance and availability. Aerospace experience is not required to be successful here - rather we look for smart, motivated, respectful, collaborative engineers who love solving problems and want to make an impact on a super inspiring mission. You will have full ownership of challenging problems, working with a team of enthusiastic engineers with diverse perspectives to design and produce solutions that enable SpaceX to achieve its loftiest engineering goals at a rapid pace. The success of the missions at SpaceX depends on the software that you and your team produce. This role will report through SpaceX Internal AI Infrastructure, as we'll also be providing support for training workloads., * Develop highly reliable, high-throughput inference systems that serve the best AI models internally across SpaceX * Architect and implement scalable distributed infrastructure for model serving, including load balancing, auto-scaling, batch scheduling, global KV cache, and continuous batching * Optimize latency and throughput of model inference under real production workloads, including low-level GPU kernel work, quantization, speculative decoding, and other acceleration techniques * Build reliable, high-concurrency serving systems with 100% uptime, low tail latency, and excellent observability * Own end-to-end components such as request routing, SDK development, rate limiting, and efficient scaling for internal SpaceX AI inference platforms * Benchmark, fine-tune, and accelerate inference engines (e.g., SGLang, vLLM, TensorRT-LLM) * Develop custom tools for tracing, replaying, and resolving issues across the full stack - from orchestration down to GPU kernels * Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates * Collaborate across SpaceXAIteams to integrate inference capabilities into broader systems and workflows ## Related Videos - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [You are not an AI developer](https://www.wearedevelopers.com/videos/1148-you-are-not-an-ai-developer) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Hacking AI at the Edge of the Indian Ocean](https://www.wearedevelopers.com/videos/100177-hacking-ai-at-the-edge-of-the-indian-ocean) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)