> Markdown version of [/jobs/ext/2643115-ai-infrastructure-engineer-recommendation-llm](https://www.wearedevelopers.com/jobs/ext/2643115-ai-infrastructure-engineer-recommendation-llm). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Infrastructure Engineer - Recommendation & LLM - **Company:** Tiktok Accommodation - **Location:** San Jose, CA, United States - **Experience:** Experienced - **Salary:** $128,000.0 - $316,800.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Nvidia CUDA, Computer Programming, Computer Engineering, Data Structures, Distributed Computing Environment, Distributed Systems, General-Purpose Computing on Graphics Processing Units, Python (Programming Language), Open Source Technology, Recommender Systems, Tensorflow, Software Engineering, AI Infrastructure, Graphics Processing Unit (GPU), High Performance Computing, Pytorch, Large Language Models, Gpu Programming, Information Technology, Low Latency, Machine Learning Operations - **Published:** August 24, 2026 - **Apply:** https://www.dice.com/job-detail/73269a6b-2da8-4b15-8c4f-67c4144fe693 ## About the Role Minimum Qualifications - Bachelor's, Master's, or Ph.D. degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience. - 3+ years of software engineering, ML systems, distributed systems, or related industry experience. - Strong programming skills in C++ or Python and experience building production-quality software. - Strong understanding of data structures, algorithms, operating systems, distributed systems, or computer architecture. - Hands-on experience with PyTorch, TensorFlow, or other machine learning frameworks. - Experience designing, developing, or optimizing large-scale production systems. - Strong problem-solving skills and the ability to independently drive complex technical projects. Preferred Qualifications - Experience building infrastructure for large-scale recommendation systems, LLMs, or foundation models. - Strong understanding of distributed training and model parallelism, including DP, TP, PP, FSDP, ZeRO, or related technologies. - Hands-on experience with GPU programming and optimization, including CUDA, Triton, GPU kernels, or similar technologies. - Experience with LLM inference and serving, including KV Cache, Continuous Batching, FlashAttention, CUDA Graph, speculative decoding, or related techniques. - Experience optimizing GPU utilization, memory efficiency, communication performance, latency, and throughput at scale. - Experience with distributed systems, high-performance computing, networking, or ML infrastructure. - Experience designing and owning large-scale production systems from architecture through deployment and operation. - Contributions to open-source projects, research publications, or technical projects in machine learning systems, distributed systems, GPU computing, or LLM infrastructure. - Experience mentoring engineers or leading technical projects is a plus. ## Description About the Team We are looking for experienced Software Engineers / ML Systems Engineers to join our Model Infrastructure team and build the next generation of AI infrastructure powering TikTok's For You recommendation system and Large Language Models (LLMs). Our team develops the core training and serving infrastructure behind one of the world's largest recommendation systems, enabling billions of personalized recommendations every day. We are also building next-generation infrastructure for foundation models and LLMs, covering large-scale model training, online inference, GPU optimization, distributed systems, and AI serving. As a member of the team, you will take ownership of challenging infrastructure problems at massive scale, working across model, framework, runtime, GPU, distributed systems, and production serving. You will collaborate closely with researchers, algorithm engineers, and infrastructure teams to turn state-of-the-art AI technologies into highly scalable, reliable, and efficient production systems. This role is ideal for engineers who are passionate about AI systems, distributed computing, GPU optimization, Recommendation, LLMs, and building high-performance infrastructure at massive scale. Responsibilities - Design, develop, and optimize large-scale AI training and online inference infrastructure for recommendation models and LLMs. - Drive the architecture and implementation of distributed training and serving systems with high scalability, reliability, and efficiency. - Optimize end-to-end model performance across GPU computation, communication, memory, networking, and runtime systems. - Develop and optimize LLM training and serving infrastructure, including model parallelism, KV Cache, Continuous Batching, and efficient inference. - Work closely with researchers and algorithm engineers to productionize new model architectures and algorithms. - Identify and resolve performance bottlenecks across the full AI stack, from model and framework to GPU kernels and distributed runtime. - Improve system latency, throughput, GPU utilization, scalability, and infrastructure cost efficiency. - Drive technical design, implementation, performance benchmarking, and production rollout of critical infrastructure components. - Mentor junior engineers and contribute to the team's technical direction and engineering standards. ## Related Videos - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [How We Built a Machine Learning-Based Recommendation System (And Survived to Tell the Tale)](https://www.wearedevelopers.com/videos/752-how-we-built-a-machine-learning-based-recommendation-system-and-survived-to-tell-the-tale) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)