> Markdown version of [/jobs/ext/2163915-software-engineer](https://www.wearedevelopers.com/jobs/ext/2163915-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - **Company:** Nimble Robotics - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $300,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, C++ (Programming Language), Computer Clusters, Program Optimization, Code Review, Nvidia CUDA, Computer Programming, Information Engineering, Data Files, Software Debugging, File Systems, Distributed Computing Environment, Distributed Systems, Memory Management, Python (Programming Language), Linux Kernel, Machine Learning, Network File Systems, Performance Tuning, Tensorflow, Software Engineering, System Programming, Extensible Markup Language (XML), Parquet, Scripting, Graphics Processing Unit (GPU), Data Ingestion, Pytorch, Delivery Pipeline, Kubernetes, Information Technology, Machine Learning Operations, Golang, Programming Languages - **Published:** August 21, 2026 - **Apply:** https://www.careerbuilder.com/job-details/senior-staff-software-engineer-infrastructure-ml-san-francisco-ca--72932266-420f-40e9-a6ec-f13fd86b81da ## About the Role * Bachelor's, Master's, or PhD in Computer Science or a related field, or equivalent practical experience. * 4+ years of industry experience in infrastructure, distributed systems, ML systems, robotics, or a related area. * Experience with programming languages such as Rust, Go, Python, or C++. * Experience with ML frameworks such as PyTorch or JAX. * Strong understanding of distributed systems, systems programming fundamentals, memory management, and performance optimization. * Experience with Kubernetes orchestration, resource scheduling for large distributed jobs, and containerized deployment pipelines. * Ability to debug and optimize bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operations. * Ability to reason from first principles and optimize systems for both memory-bound and compute-bound workloads. * Strong cross-functional communication skills, ownership, and a growth mindset. Nice to Have * Hands-on experience with distributed training frameworks and techniques such as PyTorch DDP/FSDP, DeepSpeed, Megatron, or NCCL. * Hands-on experience with GPU kernel development. * Experience with data engineering technologies such as Parquet, Arrow, or similar systems., Alliance/Partner Management, Artificial Intelligence (AI), Best Practices, C++ Programming Language, CUDA (Compute Unified Device Architecture), Code Reviews, College Level Faculty, Communication Skills, Computer Science, Cross-Functional, Data Partitioning, Data Sets, Debugging Skills, Distributed Computing, Environmental Impact, Funding, GPU (Graphics Processing Unit), Go Programming Language (Golang), High Throughput, JAX (Java API for XML), Kernel Programming, Memory Hardware, Memory Management, Mentoring, NFS (Network File System), Performance Tuning/Optimization, Product Development, Programming Languages, Python Programming/Scripting Language, Research Skills, Retirement Planning, Robotics, Software Engineering, Supply Chain, Systems Scalability, Systems/Internals Programming, Technical Analysis, Testability, Trade-Off Analysis, Warehousing ## Description We're looking for a Software Engineer to join our ML Infrastructure team. In this role, you'll help build the training and inference systems that power our general-purpose warehouse robots. You'll own training infrastructure end to end: keeping GPUs highly utilized, making runs reproducible, and ensuring every researcher can launch the next experiment with a single command. You'll work closely with ML and Robotics teams to design, build, and scale the systems that turn our GPU clusters into a reliable, high-throughput platform for model development. Responsibilities * Design, develop, and maintain ML training infrastructure that enables the AI team to run training jobs efficiently, manage and iterate experiments quickly. * Build low-latency inference pipelines for production robotics workloads. * Develop, tune, and optimize low-level CUDA kernels. * Design training-platform systems for scalable model training, including high-throughput data ingestion, dataset sharding and sampling for distributed training. * Participate in and lead design reviews with peers and stakeholders to evaluate technical tradeoffs and select appropriate technologies. * Review code and provide feedback to uphold best practices around style, correctness, testability, performance, and maintainability. * Contribute to documentation and educational materials, adapting content as systems and workflows evolve. * Mentor junior engineers and help raise the technical bar across the team. ## Related Videos - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Parquet, Delta, Iceberg & Ducklake - An introduction for developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)