> Markdown version of [/jobs/ext/2722393-ml-performance-engineer](https://www.wearedevelopers.com/jobs/ext/2722393-ml-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Performance Engineer - **Company:** AI INFRASTRUCTURE LLC - **Location:** United States - **Contract:** Permanent contract - **Skills:** Adobe Flash, Algorithm Design, Nvidia CUDA, Memory Management, InfiniBand, Open Source Technology, AI Infrastructure, Pytorch, Large Language Models, Low Latency, TensorRT - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/founding-engineer-ml-performance-urun-8355421 ## About the Role * Deep, hands-on CUDA expertise: you have written custom kernels in production, not just called into cuBLAS * Strong background in model inference and post-training optimization at scale * Fluency in GPU memory hierarchy, warp scheduling, kernel fusion, and hardware-aware algorithm design * Experience profiling and benchmarking complex inference pipelines: you know where the time goes and how to get it back * Able to operate at the frontier with minimal guidance - you identify the problem, design the approach, and ship the fix Things that will give you an edge * Public work in GPU optimization or inference efficiency - open source contributions, a published paper, or a side project that shows your depth (vLLM, Flash-Attention, TensorRT-LLM, PyTorch, or equivalent) * Experience with hardware-aware optimization frameworks: CuTe, Triton, TileLang, or similar * Familiarity with distributed memory and communication primitives: NCCL, InfiniBand, NVLink, RoCE * Contributions to or deep familiarity with PyTorch Distributed, Ray core, or similar systems * Experience optimizing for video generation or other high-throughput, latency-sensitive generative workloads * Prior work at an inference-focused company or research lab pushing the boundary of what GPU hardware can do ## Description Performance is uRun's core differentiator. We're not chasing incremental gains - we're building infrastructure that runs 10-100x faster than the status quo. As our ML Performance Engineer, you will be the person who makes that true. This is a founding technical hire. You will write custom CUDA kernels, push GPU utilization to its limits, and own inference latency end-to-end across the stack. You will work directly with the founding team on the hardest performance problems in production AI infrastructure - and your fingerprints will be on everything we ship. What you'll actually be doing day-to-day * Write custom CUDA kernels that unlock performance headroom unavailable through off-the-shelf frameworks * Optimize model inference end-to-end, targeting sub-50ms latency across our inference platform * Drive 10x performance improvements across the stack: memory bandwidth, kernel fusion, operator scheduling, and beyond * Implement zero-copy distributed memory optimizations across multi-GPU and multi-node environments * Own GPU utilization and memory management, squeezing every available FLOP out of the hardware we run * Profile, benchmark, and instrument the full inference pipeline to find and eliminate bottlenecks systematically * Set the performance engineering bar for the team: define what fast looks like and build the tooling to measure it ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)