> Markdown version of [/jobs/ext/1716034-engineer-inference-data-plane](https://www.wearedevelopers.com/jobs/ext/1716034-engineer-inference-data-plane). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Engineer, Inference Data Plane - **Company:** DigitalOcean, LLC - **Location:** Denver, CO, United States (Remote available) - **Experience:** Expert - **Salary:** $139,000.0 - $174,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Confluence, Cloud Computing, Python (Programming Language), Open Source Technology, Performance Tuning, Digitalocean, Software Engineering, Systems Integration, User-Centered Design, Load Balancing, Large Language Models, Kubernetes, Optimization Algorithms, Free and Open-Source Software, TensorRT, Golang - **Published:** July 24, 2026 - **Apply:** https://www.digitalocean.com/careers/position/apply/?gh_jid=7749096&gh_src=dc409e81us ## About the Role * AI/ML Domain Knowledge: Hands-on experience hosting large language or multimodal models using inference engines like vLLM, SGLang, or TensorRT. * Inference Frameworks: Familiarity with distributed inference serving frameworks such as llm-d, NVIDIA Dynamo, or Ray Serve. * Inference Engine Depth: Hands-on experience with vLLM or alternatives (SGLang, TensorRT-LLM, TGI, Modular MAX), including internals like continuous batching, paged attention, and prefix caching. * Distributed Inference Fluency: Understanding of why cluster-scale serving is hard: KV-cache locality is partitioned across workers, naive round-robin routing destroys cache hit rates and tail latency, and disaggregated prefill/decode requires fast cross-pod KV transfer (e.g., NIXL). * Upstream Track Record: Merged contributions to vLLM, llm-d, SGLang, or similar projects strongly preferred. * Architecture Proficiency: Knowledge of common LLM architectures and optimization techniques (e.g., continuous batching, quantization). * Software Engineering: Expert-level proficiency in GoLang or Python and familiarity with gRPC. * Cloud Operations: Proven experience shipping customer-facing software products and running critical services in a high-scale environment similar to DigitalOcean. * Open Source Mindset: Experience integrating and building with open-source software., * We innovate with purpose. You'll be a part of a cutting-edge technology company with an upward trajectory, who are proud to simplify cloud and AI so builders can spend more time creating software that changes the world. As a member of the team, you will be a Shark who thinks big, bold, and scrappy, like an owner with a bias for action and a powerful sense of responsibility for customers, products, employees, and decisions. ## Description DigitalOcean is expanding its AI Infrastructure layer to support the next generation of AI-driven applications. We are seeking a Senior Engineer 2 to join our AI Inference Data Plane team. In this role, you will be a key technical leader responsible for designing, developing, and delivering high-scale, resilient data plane services that power our "Inference as a Service" offering. You will work at the intersection of distributed systems and specialized AI hardware to ensure our customers can deploy and scale their models with industry-leading performance and reliability. This is a hands-on role, requiring you to be able to develop high quality software while availing of all the productivity boosts granted by the latest AI coding agents. What You'll Do: * Technical Leadership: Act as a technical leader on the team, driving the end-to-end design, development, and delivery of critical data plane components hosting large generative AI models. * System Design: Architect and refine system design proposals for our high-scale, multi-tenant AI inference cloud ecosystem, ensuring they meet rigorous availability and resiliency standards. * Performance Optimization: Implement and optimize distributed inference hosting using techniques like tensor/data parallelism, KV cache optimizations, and smart routing. * Collaboration: Work cross-functionally with Product Managers, customer-facing teams, and other engineering teams to align technical roadmaps with customer needs. * Distributed Serving at Scale: Build on Kubernetes-native distributed inference frameworks like llm-d (or alternatives such as NVIDIA Dynamo, Ray Serve, KServe) to deliver prefill/decode disaggregation, KV-cache-aware routing, tiered prefix caching, and wide expert parallelism for MoE models. * Flow Control & Load Balancing: Solve the distributed-systems problems unique to LLM serving - inference-aware load balancing on queue depth, cache locality, and predicted latency; flow control and fairness across tenants; autoscaling inference pools; and moving gigabytes of KV-cache between prefill and decode instances with negligible overhead. * Open Source Contributions: Contribute upstream to llm-d, vLLM, and the inference gateway ecosystem, and represent DigitalOcean in these communities. * Mentorship: Coach and mentor junior engineers, fostering a culture of technical excellence and continuous improvement. * Operational Excellence: Maintain and operate critical, high-scale services, utilizing observability tools and defining SLOs to ensure superior platform health., Lead architecture and low-level optimizations for high-performance AI inference. Benchmark and tune GPU kernels, memory and precision management, parallelize across multi-node GPU clusters, implement quantization (FP8/INT8/FP4), advise on hardware/software stacks (CUDA/ROCm/Triton/TensorRT), mentor engineers, and collaborate to turn hardware limits into shippable inference products. Top Skills: Amd AiterAsk (Assembly)Bf16Ck (Composable Kernel)CudaFlashattentionFp4Fp8Int8Openai TritonRocmTensorrtTriton Compiler DigitalOcean ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [42 x 2 Canvases Later: Two Years, Two Minds, Many Lessons](https://www.wearedevelopers.com/videos/1458-42-x-2-canvases-later-two-years-two-minds-many-lessons) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)