ML Infrastructure Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+1 more
Job description
The Senior Software Engineer, ML Infrastructure will be responsible for designing, building, and operating the production-grade inference infrastructure that powers SambaNova’s serving stack on our Reconfigurable Dataflow Unit (RDU) architecture. SambaNova is an inference-first company, and this role sits at the heart of that mission: turning state-of-the-art inference techniques into reliable, high-throughput, low-latency services exposed to customers through SambaStack and SambaCloud. The engineer will own end-to-end systems spanning request scheduling, advanced decoding algorithms, caching layers, API surfaces, and the accuracy infrastructure that keeps the stack trustworthy. This role partners closely with ML, compiler, runtime, and product teams to ship inference features from prototype to production., * Design and productionize advanced inference techniques on RDU to optimize for performance and cost. Key areas include speculative decoding, constrained decoding, function/tool calling, prompt caching, and long-context inference.
- Own SambaNova’s integration with vLLM and adjacent serving frameworks, adapting them to RDU’s architecture.
- Own the public inference API surface exposed through SambaStack and SambaCloud.
- Build and maintain the accuracy verification and regression infrastructure that gates every inference feature shipped to customers.
- Partner with ML, compiler, runtime, and product teams to take inference features from prototype to production.
- Contribute to technical design discussions, code reviews, and architectural decisions as a senior individual contributor.
Requirements
- B.S. in Computer Science, Electrical Engineering, or related field
- 5+ years of industry experience building and operating large-scale distributed systems, ideally in ML serving
- Strong software engineering fundamentals: algorithms, data structures, concurrency, and systems design
- Experience designing and maintaining production services with strict latency, throughput, and availability requirements
- Working knowledge of modern LLM inference techniques and familiarity with open-source serving stacks such as vLLM, TensorRT-LLM, or SGLang
- Proficiency in Python
- Experience collaborating across teams to deliver complex, system-level engineering solutions
Benefits & conditions
SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits. We cover 95% premium coverage for employee medical insurance, and 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution. We also offer Dental, Vision, Short/Long term Disability, Basic Life, Voluntary Life, and AD&D insurance plans in addition to Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care. Our library of well-being benefits available to you and your dependents includes a full subscription to Headspace, Gympass+ membership with access to physical gyms, One Medical membership, counseling services with an Employee Assistance Program, and much more.
About the company
The era of pervasive AI has arrived. In this era, organizations will use generative AI to unlock hidden value in their data, accelerate processes, reduce costs, drive efficiency and innovation to fundamentally transform their businesses and operations at scale.
SambaNova Suite is the first full-stack, generative AI platform, from chip to model, optimized for enterprise and government organizations. Powered by the intelligent SN40L chip, the SambaNova Suite is a fully integrated platform, delivered on-premises or in the cloud, combined with state-of-the-art open-source models that can be easily and securely fine-tuned using customer data for greater accuracy. Once adapted with customer data, customers retain model ownership in perpetuity, so they can turn generative AI into one of their most valuable assets., The ML Infrastructure team builds and operates the inference stack that serves SambaNova’s models on RDU accelerators, from request scheduling and caching through the public APIs in SambaStack and SambaCloud. We take inference techniques like speculative decoding, constrained decoding, and long-context serving from prototype to production, and own the accuracy infrastructure that gates every feature we ship. We work alongside the ML, compiler, runtime, and product teams, since most of what we build touches all four.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
How to Become an AI Engineer
MLOps – What’s the deal behind it?