> Markdown version of [/jobs/ext/2702705-ml-infrastructure-tech-lead](https://www.wearedevelopers.com/jobs/ext/2702705-ml-infrastructure-tech-lead). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Infrastructure Tech Lead - **Company:** Reducto, Inc. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Nvidia CUDA, Software Debugging, Distributed Systems, Python (Programming Language), Open Source Technology, Software Architecture, Software Deployment, Graphics Processing Unit (GPU), Pytorch, Large Language Models, Kubernetes, Low Latency, Machine Learning Operations, TensorRT - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/machine-learning-infrastructure-tech-lead-reducto-8763969 ## About the Role * Have 5+ years of experience building production infrastructure, including significant ML systems experience. * Have led complex technical projects from an ambiguous problem through production deployment. * Are equally comfortable setting direction and personally implementing the hardest parts. * Have strong Python and systems-engineering skills. * Understand the performance characteristics of modern GPU training or inference workloads. * Are comfortable with Kubernetes and distributed training or serving frameworks. * Can reason across low-level model performance and higher-level platform architecture. * Hold yourself to a high bar for quality, precision, and operational reliability. * Operate well in a fast-changing, high-growth environment. * Take full ownership from strategy through execution. Bonus Points If You * Have optimized or implemented CUDA, Triton, or custom model-serving kernels. * Have contributed meaningfully to frameworks such as vLLM, SGLang, PyTorch, TensorRT-LLM, Ray, or related open-source systems. * Have operated distributed inference or training across hundreds or thousands of GPUs. * Have built observability, scheduling, or capacity-management systems for GPU workloads. * Have experience at an early-stage or high-growth startup. * Care deeply about connecting technical excellence to measurable business impact. ## Description As our ML Infrastructure Tech Lead, you'll own the systems that make high-performance model training and inference possible at Reducto. This is a deeply hands-on role: roughly 80% of your time will be spent building, debugging, and optimizing our infrastructure. The remaining 20% will focus on setting technical direction - identifying bottlenecks, planning our infrastructure roadmap, and helping the ML and Platform teams make strong architectural decisions. You'll work across the stack, from model-serving kernels and GPU utilization to distributed systems and Kubernetes. We're looking for someone with the experience and judgment to lead ambiguous, high-impact infrastructure projects while remaining close to the code. This is a fully in-person role at our San Francisco office. What You'll Do * Own the technical direction and roadmap for Reducto's ML infrastructure. * Build and maintain our training and inference stack, balancing fast experimentation with high-performance production serving. * Optimize model serving at every layer, including kernels, runtimes, batching, scheduling, and distributed inference. * Design systems for reliable multi-node, multi-GPU training and inference. * Improve GPU utilization, latency, throughput, reliability, observability, and cost efficiency. * Develop benchmarks that identify bottlenecks and guide infrastructure investments. * Evaluate state-of-the-art advances in training and inference and apply the ones that matter. * Build the tooling and abstractions that help ML engineers move quickly from experiments to production. * Partner with ML and Platform engineers on architecture, capacity planning, and technical prioritization. * Raise the engineering bar through design reviews, mentorship, and hands-on technical leadership. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)