> Markdown version of [/jobs/ext/2681867-ai-engineer-llm-infra](https://www.wearedevelopers.com/jobs/ext/2681867-ai-engineer-llm-infra). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Engineer - LLM Infra - **Company:** Yutori, Inc. - **Location:** San Francisco, CA, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computer Clusters, Nvidia CUDA, Large Language Models, Low Latency, Machine Learning Operations - **Published:** August 31, 2026 - **Apply:** https://www.careerboard.com/us/en/find-jobs-in-United-States/-B5A2466B29EC93853E/ ## About the Role * Experience with ML infrastructure (GPU clusters) and supporting networking (NCCL) * Experience optimizing post-training and inference performance of multimodal LLMs (data/tensor/pipeline/context/expert parallelism, optimizing MFU, throughput, latency) * Low level systems experience (Triton, CUDA) * High IQ, high EQ, high agency, high craftsmanship, low ego. Proactive, clear communication. ## Description Yutori is reimagining how people interact with the web by building AI agents that can reliably do everyday digital tasks. We are building the entire stack to be agent-first, from training our own models to generative product interfaces. Towards this goal, we are looking for a member of the AI technical staff to join the founding team. Someone technically strong, and excited about building superhuman AI agents that take actions on the web. Our founders - Devi Parikh, Abhishek Das, Dhruv Batra - have decades of experience in AI research and product spanning generative, multimodal and embodied AI at Meta. Our team combines AI experience with design-minded product thinking to build and deliver on Yutori's mission. Yutori is backed by a stellar set of visionary investors - Elad Gil, Sarah Guo, Jeff Dean, Fei-Fei Li, Amjad Masad, Guillermo Rauch, Akshay Kothari, Soleio, Oliver Cameron, Julien Chaumond, Logan Kilpatrick, Bryan McCann, Vladlen Koltun, Jamie Cuffe, Michele Catasta, etc. Responsibilities: * Scale infra for post-training of multimodal LLMs (CPT, SFT, RL, search, reward models) * Scale infra for agentic inference (throughput and latency of perception-planning-action loops) * Build the foundations of a superhuman generalist web-agent * Work closely with product engineers to translate cutting-edge AI capabilities into reliable product experiences. ## Related Videos - [How to develop an autonomous car end-to-end: Robotic Drive and the mobility revolution](https://www.wearedevelopers.com/videos/22-how-to-develop-an-autonomous-car-end-to-end-robotic-drive-and-the-mobility-revolution) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)