> Markdown version of [/jobs/ext/2653267-large-language-model-inference-system-engineer-graduate-applied-machine-learning-2027-start](https://www.wearedevelopers.com/jobs/ext/2653267-large-language-model-inference-system-engineer-graduate-applied-machine-learning-2027-start). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Large Language Model Inference System Engineer Graduate (Applied Machine Learning) - 2027 Start - **Company:** BYTEDANCE INC. - **Location:** San Jose, CA, United States - **Experience:** Starter - **Salary:** $128,000.0 - $256,000.0 - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Program Optimization, Nvidia CUDA, Data Structures, Software Design Patterns, Distributed Systems, Python (Programming Language), Machine Learning, Network Service, Remote Direct Memory Access, High Performance Computing, Kubernetes, Information Technology - **Published:** August 6, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=dc85e0f0ffa77242 ## About the Role * Individuals who are completing or have recently completed a Bachelor's or Master's degree in Computer Science or a related discipline. * Strong command of algorithms, design patterns, and data structures, with solid knowledge of operating systems and computer architecture. * Proficient in one or more programming languages such as C++ or Python, with good coding style. * Understands GPU hardware architecture, is familiar with high-performance computing software stacks such as CUDA, and has experience in GPU performance analysis. * Strong interest in distributed systems and large-scale heterogeneous inference; enjoys studying low-level principles and performance bottlenecks, and actively follows progress in related fields., * Experience in large model inference system optimization, with practical understanding of PD disaggregation, KV Cache systems, and multi-node inference. * Experience in distributed network communication optimization, with deep understanding of distributed communication operator implementation and RDMA principles. * Familiarity with resource orchestration and scheduling frameworks such as Kubernetes and Ray., Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment ## Description * Participate in the engineering development of the Volcano Ark MaaS inference system, optimizing large model inference performance, cost, and stability in ultra-large-scale heterogeneous inference clusters. * Reduce large model inference costs through system-level approaches such as disaggregated multi-role inference, distributed KV Cache systems, heterogeneous inference, elastic computing, and multi-tenant co-located inference. ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Phel, a native Lisp for PHP](https://www.wearedevelopers.com/videos/791-phel-a-native-lisp-for-php) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Distributed search under the hood](https://www.wearedevelopers.com/videos/256-distributed-search-under-the-hood) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix)