> Markdown version of [/jobs/ext/2774897-ml-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2774897-ml-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Infrastructure Engineer - **Company:** Greenhouse Software, Inc. - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $200,000.0 - $275,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Code Review, Data Structures, Distributed Systems, Data Flow Control, Python (Programming Language), Open Source Technology, Large Language Models, Concurrency, Information Technology, Low Latency, Machine Learning Operations, TensorRT - **Published:** September 7, 2026 - **Apply:** https://boards.greenhouse.io/embed/job_app?for=sambanovasystems&token=6007942004 ## About the Role * B.S. in Computer Science, Electrical Engineering, or related field * 5+ years of industry experience building and operating large-scale distributed systems, ideally in ML serving * Strong software engineering fundamentals: algorithms, data structures, concurrency, and systems design * Experience designing and maintaining production services with strict latency, throughput, and availability requirements * Working knowledge of modern LLM inference techniques and familiarity with open-source serving stacks such as vLLM, TensorRT-LLM, or SGLang * Proficiency in Python * Experience collaborating across teams to deliver complex, system-level engineering solutions ## Description The Senior Software Engineer, ML Infrastructure will be responsible for designing, building, and operating the production-grade inference infrastructure that powers SambaNova's serving stack on our Reconfigurable Dataflow Unit (RDU) architecture. SambaNova is an inference-first company, and this role sits at the heart of that mission: turning state-of-the-art inference techniques into reliable, high-throughput, low-latency services exposed to customers through SambaStack and SambaCloud. The engineer will own end-to-end systems spanning request scheduling, advanced decoding algorithms, caching layers, API surfaces, and the accuracy infrastructure that keeps the stack trustworthy. This role partners closely with ML, compiler, runtime, and product teams to ship inference features from prototype to production., * Design and productionize advanced inference techniques on RDU to optimize for performance and cost. Key areas include speculative decoding, constrained decoding, function/tool calling, prompt caching, and long-context inference. * Own SambaNova's integration with vLLM and adjacent serving frameworks, adapting them to RDU's architecture. * Own the public inference API surface exposed through SambaStack and SambaCloud. * Build and maintain the accuracy verification and regression infrastructure that gates every inference feature shipped to customers. * Partner with ML, compiler, runtime, and product teams to take inference features from prototype to production. * Contribute to technical design discussions, code reviews, and architectural decisions as a senior individual contributor. ## Related Videos - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Concurrency with Go](https://www.wearedevelopers.com/videos/191-concurrency-with-go) - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) - [Structured Concurrency in Practice: CoroutineScope vs StructuredTaskScope](https://www.wearedevelopers.com/videos/100349-structured-concurrency-in-practice-coroutinescope-vs-structuredtaskscope) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)