> Markdown version of [/jobs/ext/585677-founding-machine-learning-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/585677-founding-machine-learning-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Founding Machine Learning Infrastructure Engineer - **Company:** AI MODELS LLC - **Location:** Palo Alto, CA, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Nvidia CUDA, Software Debugging, Distributed Computing Environment, Distributed Systems, Memory Management, Machine Learning, Open Source Technology, Graphics Processing Unit (GPU), High Performance Computing, Pytorch, Large Language Models, Low Latency, Machine Learning Operations, TensorRT - **Published:** June 18, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=dadfa553d7432e3d ## About the Role Do you have experience in Customer communication?, * Strong experience in ML systems, distributed systems, or high-performance computing. * Experience optimizing inference or training workloads for large models. * Familiarity with TPUs, GPUs, or other accelerators. * Experience with one or more of CUDA, Triton, NCCL, JAX/XLA, PyTorch internals, vLLM, SGLang, TensorRT-LLM, distributed inference, or distributed training. * Strong systems debugging skills. * Comfort working across model code, runtime, infrastructure, and product requirements. * High ownership and the ability to operate effectively in an early-stage startup environment. Cultural Fit * Hands-on technical excellence and strong engineering judgment. * End-to-end ownership, from design to implementation to production outcomes. * Bias for action: ship quickly, learn from failures, and iterate. * High intensity during critical milestones, with a focus on real customer impact. * Ability to do deep, focused work and sustain execution. * Clear communication with teammates, customers, and stakeholders. * Comfort with ambiguity, rapid change, and wearing multiple hats. * Low ego, high integrity, high accountability, and strong collaboration. * Continuous learning and a belief that judgment, intelligence, and capability compound over time. If you are excited to build the infrastructure and agent systems behind the next generation of AI applications, push open-source models to production-grade performance, and turn ambitious research ideas into real-world impact, Model AI is the place for you. ## Description You will work on model serving performance, accelerator utilization, long-context inference, batching, scheduling, KV cache management, runtime efficiency, and cost reduction. This is a deeply technical role at the intersection of ML systems, infrastructure, and product. Direct TPU experience is a strong plus, but not required. We care most about strong ML systems fundamentals, performance intuition, and the ability to ship reliable systems quickly. What You'll Do * Optimize large-scale LLM inference and serving systems. * Improve total tokens per second, decode tokens per second, latency, throughput, and cost efficiency. * Work on serving infrastructure for open-source models across different types of accelerators. * Improve batching, scheduling, KV cache management, memory usage, and accelerator utilization. * Support long-context inference, including workloads targeting up to 1M context. * Debug performance bottlenecks across model execution, runtime, networking, and infrastructure. * Work with frameworks such as JAX/XLA, PyTorch, vLLM, SGLang, TensorRT-LLM, or related systems. * Collaborate closely with the application team to ensure infrastructure is optimized for agentic workloads, not just generic chatbot inference. * Help turn research prototypes into reliable, high-performance production systems. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Localized Open Models in Production: What Builders Need to Know](https://www.wearedevelopers.com/videos/100270-localized-open-models-in-production-what-builders-need-to-know) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)