> Markdown version of [/jobs/ext/1469887-backend-engineer-inference-platform](https://www.wearedevelopers.com/jobs/ext/1469887-backend-engineer-inference-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Backend Engineer, Inference Platform - **Company:** Together Ai - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $160,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Nvidia CUDA, Computer Engineering, Data Centers, Distributed Systems, Memory Management, Fault Tolerance, InfiniBand, Python (Programming Language), Open Source Technology, TypeScript, Multithreading, Load Balancing, Large Language Models, Generative AI, Kubernetes, Information Technology, Microservices - **Published:** July 28, 2026 - **Apply:** https://www.dice.com/job-detail/0d9c44bd-02c5-4b06-9fb2-27fe1ed1a946 ## About the Role * 5+ years of demonstrated experience building large-scale, fault-tolerant, distributed systems and API microservices. * Strong background in designing, analyzing, and improving efficiency, scalability, and stability of complex systems. * Excellent understanding of low-level OS concepts: multi-threading, memory management, networking, and storage performance. * Expert-level programming in one or more of: Rust, Go, Python, or TypeScript. * Knowledge of modern LLMs and generative models and how they are served in production is a plus. * Experience working with the open source ecosystem around inference is highly valuable; familiarity with SGLang, vLLM, or NVIDIA Dynamo will be especially handy. * Experience with Kubernetes or container orchestration is a strong plus. * Familiarity with GPU software stacks (CUDA, Triton, NCCL) and HPC technologies (InfiniBand, NVLink, MPI) is a plus. * Bachelor's or Master's degree in Computer Science, Computer Engineering, or related field, or equivalent practical experience. ## Description * Build and optimize global and local request routing, ensuring low-latency load balancing across data centers and model engine pods. * Develop auto-scaling systems to dynamically allocate resources and meet strict SLOs across dozens of data centers. * Design systems for multi-tenant traffic shaping, tuning both resource allocation and request handling - including smart rate limiting and regulation - to ensure fairness and consistent experience across all users. * Engineer trade-offs between latency and throughput to serve diverse workloads efficiently. * Optimize prefix caching to reduce model compute and speed up responses. * Collaborate with ML researchers to bring new model architectures into production at scale. * Continuously profile and analyze system-level performance to identify bottlenecks and implement optimizations. ## Related Videos - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Do TypeScript without TypeScript](https://www.wearedevelopers.com/videos/327-do-typescript-without-typescript) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [Vuejs and TypeScript- Working Together like Peanut Butter and Jelly](https://www.wearedevelopers.com/videos/127-vuejs-and-typescript-working-together-like-peanut-butter-and-jelly) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)