> Markdown version of [/jobs/ext/2407425-ml-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2407425-ml-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Infrastructure Engineer - **Company:** Finoit Inc - **Location:** Redwood City, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Computer Clusters, Distributed Computing Environment, Distributed Systems, Memory Management, Job Scheduling, Machine Learning, Performance Tuning, Google Cloud, Pytorch, Containerization, Kubernetes, Slurm, Machine Learning Operations, TensorRT, Data Pipelines - **Published:** August 6, 2026 - **Apply:** https://www.dice.com/job-detail/36a10000-c41f-4113-933d-a1e87110a363 ## About the Role * 7+ years of experience in ML Infrastructure, HPC, or Distributed Systems. * Strong experience with PyTorch, DeepSpeed, FSDP, ZeRO, or similar distributed training frameworks. * Hands-on experience with Kubernetes, cloud platforms (AWS/Google Cloud Platform), and containerized environments. * Strong understanding of distributed systems, GPU optimization, NCCL, memory management, and performance tuning. * Experience building scalable ML infrastructure from development through production., * Experience with multimodal AI, robotics data pipelines, Triton, TensorRT, custom ML kernels, or ML compiler/runtime optimization. ## Description * Design and scale distributed ML training infrastructure for large GPU clusters. * Build and optimize training pipelines using PyTorch, DeepSpeed, and distributed training frameworks. * Develop and maintain job scheduling systems using Kubernetes and/or SLURM. * Create high-throughput data pipelines for large-scale multimodal datasets. * Optimize GPU utilization, memory efficiency, and overall system performance. * Build low-latency inference pipelines for production ML deployments. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)