> Markdown version of [/jobs/ext/1429010-senior-software-engineer-quantized-inference](https://www.wearedevelopers.com/jobs/ext/1429010-senior-software-engineer-quantized-inference). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer, Quantized Inference - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $152,000.0 - $287,500.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Data Analysis, C++ (Programming Language), Code Review, Software Debugging, Python (Programming Language), Linux Kernel, Open Source Technology, Performance Tuning, Software Engineering, Pytorch, Large Language Models, Information Technology, Code Testing, HuggingFace, Build Tools - **Published:** July 24, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=e458011d2d422220 ## About the Role * Proficient in Python; familiarity with C++ * Strong software engineering fundamentals: concise, well-tested code; fluent with AI-assisted tooling * Experience with ML accelerators with a basic understanding of how certain ML layers affect execution time * Familiarity with PyTorch internals (custom ops, autograd, export) or equivalent framework * Experience reading, modifying, or contributing to a large open-source codebase * MS/PhD in Computer Science or related field, or equivalent experience. * 4+ years in a relevant software engineering role * Demonstrated ability to move fast with ambiguous requirements, with strong written and verbal communication Ways to stand out from the crowd: * Experience contributing to inference serving frameworks (vLLM, TRT-LLM, SGLang) or Triton kernel development * Track record of debugging numerical issues across mixed-precision boundaries * Deep experience with model compression techniques: PTQ, QAT, structured/unstructured sparsity ## Description We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery and deployment of efficient inference recipes for LLMs. A recipe defines which operators are transformed into low-precision or sparsified variants - unlocking throughput and latency gains without regressing accuracy or verbosity. Recipes may incorporate techniques such as rotations, block scaling to attenuate outlier impact, or improved calibration data drawn from SFT/RL pipelines. Each new recipe demands corresponding kernel and model-level implementations in inference engines (vLLM, TRT-LLM, SGLang). The candidate will translate recipe specifications into functionally correct, performant code, e.g., writing Triton kernels, inserting quantize/dequantize nodes into prefill and decode paths, and ensuring per-expert scaling in MoE layers is handled correctly. From there, the candidate will collaborate with partner inference teams to further optimize throughput and interactivity on target workloads. This work is a core component of our productization effort across Megatron-LM, ModelOpt, and vLLM. What you'll be doing: * Implement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang) * Own model export pipelines (ModelOpt, Megatron-LM <-> HuggingFace), ensuring quantized checkpoints serialize correctly for downstream serving * Build prototypes and benchmarking harnesses to evaluate recipe throughput/interactivity before full optimization * Develop data analysis tooling and visualizations for numerics debugging * Improve developer productivity across the team: CI, build systems, training infrastructure, pipeline friction * Participate in code reviews and incorporate feedback ## Related Videos - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Build a CI/CD pipeline to automate code reviews and ensure code quality](https://www.wearedevelopers.com/videos/349-build-a-ci-cd-pipeline-to-automate-code-reviews-and-ensure-code-quality) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [The Fastest-Growing Tech Sectors to Look Out for in 2025](https://www.wearedevelopers.com/magazine/373-the-fastest-growing-tech-sectors-to-look-out-for-in-2025)