> Markdown version of [/jobs/ext/1014588-ai-inference-engineer](https://www.wearedevelopers.com/jobs/ext/1014588-ai-inference-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Inference Engineer - **Company:** Triune Infomatics Inc - **Location:** San Jose, CA, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, C++ (Programming Language), Computer Clusters, Profiling, Nvidia CUDA, Software Debugging, Distributed Systems, Fault Tolerance, Python (Programming Language), Node.Js, OpenShift, Rust (Programming Language), Datadog, Network Routers, Graphics Processing Unit (GPU), Load Balancing, System Availability, Large Language Models, Kubernetes, Hardware Acceleration, Front End Software Development, TensorRT, Hardware Infrastructure, Api Design, Microservices - **Published:** June 10, 2026 - **Apply:** https://www.dice.com/job-detail/e4872a74-a68b-49a0-9c67-0e9d12cc5126 ## About the Role The ideal candidate has hands-on experience building or operating production-grade inference serving systems and is comfortable working close to the hardware, from CUDA/ROCm kernels to distributed multi-node, multi-GPU clusters serving large language models at scale., · Proven experience with production model-serving frameworks (vLLM, SGLang, Triton Inference Server, TensorRT-LLM, TorchServe, KServe, or custom runtimes) · Strong proficiency in C++, Python, and Rust for building high-performance, memory-efficient systems · Hands-on experience writing GPU kernels using CUDA and/or ROCm · Solid understanding of LLM inference internals, including attention mechanisms, KV cache management, continuous batching, and quantization · Experience with distributed, multi-node, multi-GPU serving environments · Experience deploying and managing services on Kubernetes, OpenShift, or similar orchestration platforms · Strong background in performance profiling, benchmarking, and debugging latency or throughput issues, · Direct experience working with NVIDIA Dynamo or similar distributed serving architectures (router, worker discovery, multi-model routing) · Experience supporting diverse model types in production, including MoE, multimodal, and hybrid attention/SSM architectures · Familiarity with OpenAI-compatible API design and implementation · Experience with telemetry and observability tooling for large-scale GPU infrastructure ## Description We are seeking a highly skilled AI Inference Engineer to join our team and drive the performance, scalability, and reliability of our large-scale model serving infrastructure. This role sits at the intersection of systems engineering, GPU optimization, and distributed infrastructure, and is ideal for someone who thrives on squeezing maximum performance out of production AI workloads., Inference Serving & Optimization · Build, operate, and optimize production model-serving stacks using frameworks such as vLLM, SGLang, Triton Inference Server, TensorRT-LLM, TorchServe, or KServe · Develop and maintain custom high-throughput microservices for model inference using C++, Python, and Rust GPU & Hardware Acceleration · Write and optimize custom GPU kernels using CUDA, ROCm, or Triton · Apply deep understanding of GPU architecture, including memory hierarchies and tensor cores, to improve compute efficiency LLM Inference Internals · Optimize prefill and decode stages, attention mechanisms, and continuous batching · Implement and tune quantization, speculative decoding, tensor parallelism, pipeline parallelism, and Mixture of Experts (MoE) serving strategies Memory & KV Cache Management · Design and implement KV cache optimization strategies, including PagedAttention, chunked prefill, prefix caching, and quantized KV · Develop cache transfer and offload strategies to manage memory pressure under high-volume, irregular workloads Distributed Systems & Infrastructure · Build and operate fault-tolerant, high-concurrency serving systems deployed on Kubernetes, OpenShift, Helm, or similar orchestration platforms · Implement tensor parallelism, pipeline parallelism, and distributed computing across multi-node, multi-GPU clusters Distributed Serving Platform (Dynamo) · Contribute to distributed serving architecture components including frontend, router, worker discovery, multi-model routing, and health checks · Build and maintain OpenAI-compatible endpoints across multiple backends, including SGLang, TensorRT-LLM, and vLLM Performance & Reliability · Conduct deep profiling and benchmarking to identify and resolve latency and throughput regressions · Build telemetry-driven observability platforms ensuring high availability, load balancing, and dynamic request scheduling Model Support · Bring up and support a broad range of model classes in production, including decoder-only LLMs, MoE models, hybrid attention/SSM models, multimodal models, embedding models, reward models, and classification models ## Related Videos - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Stop Using Node.js Like It’s 2020! - Alfonso Graziano](https://www.wearedevelopers.com/videos/1863-stop-using-node-js-like-it-s-2020-alfonso-graziano) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)