> Markdown version of [/jobs/ext/1242175-senior-ai-infrastructure-engineer-model-training](https://www.wearedevelopers.com/jobs/ext/1242175-senior-ai-infrastructure-engineer-model-training). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior AI Infrastructure Engineer - Model Training - **Company:** Kodiak, LLC - **Location:** Mountain View, CA, United States - **Experience:** Expert - **Salary:** $190,000.0 - $260,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Computer Clusters, Program Optimization, Profiling, Nvidia CUDA, Extract Transform Load (ETL), Shard (Database Architecture), Distributed Computing Environment, InfiniBand, Python (Programming Language), Data Streaming, AI Infrastructure, Graphics Processing Unit (GPU), Pytorch, Information Technology, Machine Learning Operations, Stream Processing, Lidar, Data Pipelines - **Published:** July 11, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=3097777cb2855cdb ## About the Role * BS, MS, or PhD in Computer Science or a related field, and at least 2-3 years of industry experience in ML systems or infrastructure * Hands-on experience with distributed training frameworks and techniques (PyTorch DDP/FSDP, DeepSpeed, Megatron, NCCL) and a strong grasp of parallelism trade-offs * Experience building high-performance data pipelines for large-scale training, including streaming dataset formats (WebDataset, MosaicML Streaming/MDS, or similar), sharding, and storage/network-aware loading * Deep understanding of GPU performance: mixed precision, memory hierarchy, kernel fusion, profiling tools (Nsight, PyTorch Profiler), and interconnects (NVLink, InfiniBand) * Strong Python skills and proficiency in PyTorch internals; systems-level experience (C++/CUDA/Triton) a plus * Passion for building the infrastructure that lets AI for the physical world train faster, scale further, and improve continuously ## Description * Design high-throughput data loading and streaming systems for multimodal sensor data (camera, LiDAR, radar), including dataset formats, sharding strategies, and prefetching pipelines that keep GPUs saturated * Build and optimize distributed training infrastructure across multi-node GPU clusters, applying data, tensor, pipeline, and fully sharded (FSDP/ZeRO) parallelism to models that don't fit on a single device * Maximize utilization of modern accelerators such as NVIDIA B200s through mixed-precision training (BF16/FP8), fused kernels, memory optimization, and communication/computation overlap * Profile end-to-end training pipelines to find and eliminate bottlenecks across storage, network, CPU preprocessing, and GPU compute * Develop scalable dataset construction pipelines that convert petabytes of raw driving logs into training-ready, streamable formats * Partner with ML teams to scale new architectures from prototype to full-cluster training runs efficiently and reliably ## Related Videos - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [How to develop an autonomous car end-to-end: Robotic Drive and the mobility revolution](https://www.wearedevelopers.com/videos/22-how-to-develop-an-autonomous-car-end-to-end-robotic-drive-and-the-mobility-revolution) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Enhancing Workload Security in Kubernetes](https://www.wearedevelopers.com/videos/356-enhancing-workload-security-in-kubernetes) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)