> Markdown version of [/jobs/ext/1310104-software-engineer-ml-infrastructure](https://www.wearedevelopers.com/jobs/ext/1310104-software-engineer-ml-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, ML Infrastructure - **Company:** The Eeo - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $200,000.0 - $275,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Code Review, Data Structures, Distributed Systems, Data Flow Control, Python (Programming Language), Open Source Technology, Large Language Models, Concurrency, Information Technology, Low Latency, Machine Learning Operations, TensorRT - **Published:** July 17, 2026 - **Apply:** https://www.dice.com/job-detail/3d22b29d-e414-4ec3-a3b3-ed38bdcfc01c ## About the Role * Bachelor's degree in Computer Science, Electrical Engineering, or related field * 5+ years of industry experience building and operating large-scale distributed systems, ideally in ML serving * Strong software engineering fundamentals: algorithms, data structures, concurrency, and systems design * Experience designing and maintaining production services with strict latency, throughput, and availability requirements * Working knowledge of modern LLM inference techniques and familiarity with open-source serving stacks such as vLLM, TensorRT-LLM, or SGLang * Proficiency in Python * Experience collaborating across teams to deliver complex, system-level engineering solutions ## Description The Senior Software Engineer, ML Infrastructure will be responsible for designing, building, and operating the production-grade inference infrastructure that powers SambaNova's serving stack on our Reconfigurable Dataflow Unit (RDU) architecture. SambaNova is an inference-first company, and this role sits at the heart of that mission: turning state-of-the-art inference techniques into reliable, high-throughput, low-latency services exposed to customers through SambaStack and SambaCloud. The engineer will own end-to-end systems spanning request scheduling, advanced decoding algorithms, caching layers, API surfaces, and the accuracy infrastructure that keeps the stack trustworthy. This role partners closely with ML, compiler, runtime, and product teams to ship inference features from prototype to production., * Design and productionize advanced inference techniques on RDU to optimize for performance and cost. Key areas include speculative decoding, constrained decoding, function/tool calling, prompt caching, and long-context inference. * Own SambaNova's integration with vLLM and adjacent serving frameworks, adapting them to RDU's architecture. * Own the public inference API surface exposed through SambaStack and SambaCloud. * Build and maintain the accuracy verification and regression infrastructure that gates every inference feature shipped to customers. * Partner with ML, compiler, runtime, and product teams to take inference features from prototype to production. * Contribute to technical design discussions, code reviews, and architectural decisions as a senior individual contributor. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Concurrency with Go](https://www.wearedevelopers.com/videos/191-concurrency-with-go) - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [How E.On productionizes its AI model & Implementation of Secure Generative AI.](https://www.wearedevelopers.com/videos/623-how-e-on-productionizes-its-ai-model-implementation-of-secure-generative-ai) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)