> Markdown version of [/jobs/ext/2719131-end-inference-engineers](https://www.wearedevelopers.com/jobs/ext/2719131-end-inference-engineers). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # end inference engineers - **Company:** d Matrix - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, C++ (Programming Language), Profiling, Nvidia CUDA, Computer Programming, Microprocessors, Firmware, Python (Programming Language), Open Source Technology, Graphics Processing Unit (GPU), Application Specific Integrated Circuits, Large Language Models, Information Technology, ONNX (Open Neural Network Exchange) Format, Machine Learning Operations, TensorRT - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-staff-llm-inference-engineer-d-matrix-8631582 ## About the Role * Bachelor's degree in Computer Science, Electrical Engineering, or a related field, and 10+ years of relevant engineering experience; or equivalent demonstrated experience. * Master's or PhD in Computer Science, Electrical Engineering, or a related field preferred, with 6+ years of relevant industry experience. * Strong proficiency in Python and C/C++. * Hands-on experience optimizing LLM inference - attention kernels, KV cache, batching strategies, quantization (INT8/FP8/INT4). * Experience with at least one major inference framework (vLLM, SGLang, TensorRT-LLM, ONNX Runtime, or similar) at a contributor level. * Familiarity with GPU kernel programming (CUDA/Triton) and performance profiling tools. Preferred Qualifications * Experience with heterogeneous compute deployments - scheduling inference workloads across dissimilar hardware (accelerators, CPUs, GPUs). * Familiarity with custom silicon or ASIC-based inference (beyond GPU-only environments). * Experience with distributed inference: tensor parallelism, pipeline parallelism, disaggregated serving. * Contributions to open-source inference or ML systems projects. * Experience with production inference serving at scale (latency SLOs, continuous batching, multi-model serving). * Familiarity with speculative decoding, mixture-of-experts routing, or long-context serving techniques. * Working familiarity with the material in the JAX Scaling Book or equivalent systems-level understanding of modern LLM training and inference. ## Description We are hiring end-to-end inference engineers who are comfortable going from a novel research idea to a deployed, optimized system. You will work at every layer of the inference stack - from kernel-level optimization to distributed orchestration to high-level serving APIs. This role could be a great match for you if you: * Have deep intuition for modern generative AI architectures and how to squeeze performance out of them at inference time. * Are familiar with the internals of open-source inference frameworks (vLLM, SGLang, TensorRT-LLM, etc.) and can extend or replace them when needed. * Enjoy pathfinding new use cases - exploring heterogeneous deployment topologies and building early-stage POCs that prove out new ideas. * Are results-oriented with a strong bias toward action; you own problems end-to-end from prototype to optimization to handoff. * Are energized by working at the intersection of novel hardware and frontier models, and want your work to directly influence how next-generation AI silicon is used. * Value clear communication and thrive in a small, high-ownership team environment. Responsibilities * Identify and prototype emerging LLM inference use cases suited to heterogeneous hardware deployments. * Build compelling proof-of-concept systems that demonstrate D-Matrix capabilities to customers, partners, and internal stakeholders. * Develop and tune custom kernels and operator-level optimizations to maximize throughput and minimize latency. * Drive quantization, sparsity, and batching strategies tailored to D-Matrix computational model. * Build and maintain inference runtimes, serving frameworks, and evaluation tooling. * Contribute to distributed inference systems: tensor/pipeline parallelism, disaggregated prefill/decode, KV-cache management. * Work closely with hardware architects to provide firmware and compiler teams with actionable inference workload insights. * Partner with product and business development to translate POCs into customer-facing demonstrations. * Contribute to technical publications, whitepapers, and open-source projects that advance D-Matrix visibility. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Enhancing Workload Security in Kubernetes](https://www.wearedevelopers.com/videos/356-enhancing-workload-security-in-kubernetes) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)