> Markdown version of [/jobs/ext/1279143-software-engineer-ml-infrastructure](https://www.wearedevelopers.com/jobs/ext/1279143-software-engineer-ml-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - ML Infrastructure - **Company:** WATNEY CORPORATION - **Location:** San Francisco, CA, United States - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Software Debugging, Distributed Computing Environment, Memory Management, Network Topologies, Python (Programming Language), Machine Learning, Pytorch, Machine Learning Operations, Decoding - **Published:** July 15, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=253907ae056fd5be ## About the Role * Bring a experience building machine learning platforms and large-scale distributed training * Possess deep professional experience with distributed training backbones (FSDP, DeepSpeed, Megatron, Ray Train) or large-scale inference serving layers (vLLM, Triton, Ray Serve). * Exhibit fluency in Python alongside Rust or C/C++, with a strong mathematical background and practical knowledge of GPU kernel optimization or network topologies. * Have experience navigating structural edge-case hardware bottlenecks, specifically regarding video decoding, multimodal alignment, or high-throughput real-time playback. We're committed to building a diverse, inclusive team. At Watney Robotics, we welcome people of all backgrounds and identities, and we make hiring decisions based on skills, experience, and potential. If you're passionate about robotics but don't meet every requirement, we still encourage you to apply! ## Description At Watney, ML Infrastructure software engineers build the high-performance foundations that allow our perception and intelligence models to scale. You will architect the high-performance computing foundation that powers our physical intelligence models., * Own Training & Inference Infrastructure: Design and maintain multi-tenant scheduling systems that automatically place training and inference jobs based on hardware topology, cost, and priority, while enforcing fair resource sharing and preemption policies. * Scale Distributed Training: Partner with researchers to scale JAX and PyTorch-based training loops across heterogeneous GPU/TPU clusters with minimal friction, ensuring rock-solid checkpointing and metrics collection. * Optimize Performance & Hardware Bounds: Profile and improve memory usage, device utilization, throughput, and distributed synchronization, specifically navigating edge hardware bottlenecks like on-chip video decoders and memory bandwidth. * Enable Rapid Iteration: Build clean abstractions for launching, monitoring, debugging, and reproducing experiments so researchers can submit massive jobs without needing to manage underlying cluster state. * Contribute to Core Training Code: Evolve our core JAX model code and training pipelines to natively support new architectures, multimodal video/telemetry data streams, and robust evaluation metrics. * Manage Compute Resources: Ensure highly efficient allocation and utilization of massive cloud-based compute clusters while aggressively monitoring and controlling resource costs. ## Related Videos - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [How Machine Learning is turning the Automotive Industry upside down](https://www.wearedevelopers.com/videos/61-how-machine-learning-is-turning-the-automotive-industry-upside-down) - [What non-automotive Machine Learning projects can learn from automotive Machine Learning projects](https://www.wearedevelopers.com/videos/397-what-non-automotive-machine-learning-projects-can-learn-from-automotive-machine-learning-projects) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)