> Markdown version of [/jobs/ext/146867-founding-ml-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/146867-founding-ml-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Founding ML infrastructure Engineer - **Company:** Urun LLC - **Location:** United States (Remote available) - **Salary:** $200,000.0 - $350,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Computing Platforms, Microsoft Azure, Computer Clusters, Distributed Systems, InfiniBand, Performance Tuning, Dynamic Routing, Google Cloud, System Availability, Large Language Models, Multi-Cloud, Kubernetes, Bare Metal, Slurm, Machine Learning Operations, TensorRT - **Published:** May 20, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=6e8472a8deecd73d ## About the Role Do you have experience in Team management?, * Proven experience designing and operating large-scale distributed infrastructure at 1,000+ nodes or equivalent complexity, in any domain * Deep expertise in distributed systems, cluster orchestration (Kubernetes, Slurm, or custom schedulers), and large-scale resource scheduling * Strong production reliability instincts: observability, incident response, capacity planning, and SLA ownership across complex systems * Experience building infrastructure that other engineers build on top of, not just operating it * Ability to operate as a technical lead: set direction, make tradeoffs under uncertainty, and raise the bar for the team around you * Startup orientation. You are energised by ambiguity, move fast, and build for scale from day one Things that will give you an edge * Exposure to ML infrastructure concepts: GPU networking (NCCL, InfiniBand, RoCE), model serving frameworks (vLLM, SGLang, TensorRT-LLM), or hardware-aware performance tuning (CuTe, Triton, TileLang) * Experience with multi-cloud GPU procurement and capacity management across AWS, GCP, Azure, and bare metal providers * Familiarity with inference marketplace architectures, dynamic routing, or spot/preemptible workload management * Prior experience at a Series A or earlier stage company scaling from early infrastructure to production ## Description We are building the next generation of AI inference infrastructure. As our ML Infrastructure and Platform Engineer, you will own the architecture and scaling of our GPU compute platform from the ground up. This is a founding technical hire with end-to-end ownership across the full infrastructure stack, from bare metal to model serving. You will work directly with the founding team and define how we build. What you'll actually be doing day-to-day * Design and scale our GPU compute platform to support 1,000+ GPU clusters, ensuring high availability and low-latency inference across the fleet * Build and maintain the infrastructure layer for our compute marketplace, including multi-tenant scheduling, isolation, and billing-aware resource allocation * Own production reliability for ML systems end-to-end: observability, incident response, and SLA achievement across model serving and infrastructure * Architect feature stores and model registry systems that support rapid iteration and reproducibility at scale * Design an experiment tracking infrastructure capable of handling thousands of concurrent runs with full auditability * Build resource orchestration and scheduling systems that optimise for throughput, cost, and latency across heterogeneous hardware * Set engineering standards for infrastructure reliability, capacity planning, and operational excellence as an early technical leader ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Single Server, Global Reach: Running a Worldwide Marketplace on Bare Metal in a Cloud-Dominated World](https://www.wearedevelopers.com/videos/1206-single-server-global-reach-running-a-worldwide-marketplace-on-bare-metal-in-a-cloud-dominated-world) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [From Model to Metal: An Open Source Stack for Accelerating Intelligence](https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)