> Markdown version of [/jobs/ext/1465915-machine-learning-engineer-inference](https://www.wearedevelopers.com/jobs/ext/1465915-machine-learning-engineer-inference). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer - Inference - **Company:** Together Ai - **Location:** San Francisco, CA, United States - **Experience:** Experienced - **Salary:** $160,000.0 - $230,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Compilers, Nvidia CUDA, Memory Management, Fault Tolerance, Python (Programming Language), Machine Learning, Multithreading, Data Ingestion, Pytorch, Large Language Models, Build Management, Production Code, TensorRT, Decoding - **Published:** July 28, 2026 - **Apply:** https://www.dice.com/job-detail/c09bda10-616a-4323-be86-b3e6426992bf ## About the Role * 3+ years of experience writing high-performance, well-tested, production-quality code. * Proficiency with Python and PyTorch. * Demonstrated experience in building high performance libraries and tooling. * Excellent understanding of low-level operating systems concepts including multi-threading, memory management, networking, storage, performance, and scale. * Preferred: Knowledge of existing AI inference systems such as TGI, vLLM, TensorRT-LLM, Optimum * Preferred: Knowledge of AI inference techniques such as speculative decoding. * Preferred: Knowledge of CUDA/Triton programming. * Nice to have: Knowledge of Rust, Cython and compilers. ## Description Together AI is seeking a Machine Learning Engineer to join ourInference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role involves working with state-of-the-art large language models models and ensuring they run efficiently and effectively at scale. If you are passionate about AI inference, PyTorch, and developing high-performance systems, we want to hear from you. This position offers the chance to collaborate closely with AI researchers and engineers to create cutting-edge AI solutions. Join us in shaping the future at Together AI! Responsibilities * Design and build the production systems that power the Together AI inference engine, enabling reliability and performance at scale. * Develop and optimize runtime inference services for large-scale AI applications. * Collaborate with researchers, engineers, product managers, and designers to bring new features and research capabilities to the world. * Conduct design and code reviews to ensure high standards of quality. * Create services, tools, and developer documentation to support the inference engine. * Implement robust and fault-tolerant systems for data ingestion and processing. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Challenges and Solutions for Efficient, Large-Scale Video Analysis](https://www.wearedevelopers.com/videos/2022-challenges-and-solutions-for-efficient-large-scale-video-analysis) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)