> Markdown version of [/jobs/ext/1212851-ai-inference-engineer](https://www.wearedevelopers.com/jobs/ext/1212851-ai-inference-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Inference Engineer - **Company:** EPICSOFT CORPORATION - **Location:** San Jose, CA, United States - **Salary:** $135,200.0 - $156,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, C++ (Programming Language), Program Optimization, Nvidia CUDA, Computer Programming, Software Debugging, Distributed Systems, Fault Tolerance, Python (Programming Language), Node.Js, OpenShift, Performance Tuning, Software Engineering, System Programming, AI Infrastructure, Rust (Programming Language), Graphics Processing Unit (GPU), Large Language Models, Grafana, Parallel Computation, Reliability of Systems, Gpu Programming, Containerization, Kubernetes, Information Technology, Optimization Algorithms, Hardware Acceleration, TensorRT, Hardware Infrastructure, Decoding, Microservices - **Published:** July 9, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=05015933760340ca ## About the Role * Bachelor's degree in Computer Science, Software Engineering, or a related field (or equivalent experience). * 5+ years of software engineering experience. * Hands-on experience building production AI inference systems. * Strong programming skills in: * C++ * Python * Rust * Experience with one or more of the following: * vLLM * TensorRT-LLM * Triton Inference Server * SGLang * TorchServe * KServe * Experience with CUDA and/or ROCm GPU programming. * Knowledge of LLM inference optimization techniques, including: * KV Cache * Continuous Batching * Quantization * Attention Optimization * Experience deploying workloads on Kubernetes or OpenShift. * Experience working with distributed GPU infrastructure. * Strong debugging, benchmarking, and performance tuning skills. Preferred Qualifications * Experience with NVIDIA Dynamo or similar distributed inference platforms. * Experience serving: * Large Language Models (LLMs) * Multimodal Models * Mixture of Experts (MoE) * Embedding Models * Familiarity with OpenAI-compatible APIs. * Experience with telemetry, monitoring, and observability tools. Pay: $65.00 - $75.00 per hour Experience: * CUDA/ROCm: 1 year (Required) * production LLM inference: 3 years (Required) * vLLM/TensorRT-LLM/Triton production: 1 year (Required) ## Description We are seeking an experienced AI Inference Engineer to help build and optimize high-performance AI model serving infrastructure for production-scale Large Language Models (LLMs). You will work on GPU acceleration, distributed inference systems, Kubernetes deployments, and model optimization to deliver low-latency, scalable AI services. This position is ideal for engineers with strong systems programming experience who enjoy solving performance challenges across GPUs, distributed computing, and AI infrastructure. Responsibilities * AI Model Serving * Build, deploy, and optimize production AI inference services. * Work with frameworks such as vLLM, TensorRT-LLM, Triton Inference Server, SGLang, TorchServe, or KServe. * Develop scalable inference microservices using C++, Python, and Rust. * GPU Performance Optimization * Develop and optimize GPU kernels using CUDA or ROCm. * Improve GPU utilization, memory efficiency, and inference throughput. * Optimize tensor operations and hardware acceleration. * Large Language Model Optimization * Improve inference performance through: * Continuous batching * KV Cache optimization * Quantization * Tensor Parallelism * Pipeline Parallelism * Speculative Decoding * Mixture of Experts (MoE) * Distributed Infrastructure * Deploy and manage inference services on Kubernetes, OpenShift, or similar container platforms. * Support distributed multi-node, multi-GPU environments. * Build reliable, fault-tolerant AI serving infrastructure. * Performance & Reliability * Profile and benchmark production workloads. * Troubleshoot latency and throughput bottlenecks. * Implement monitoring, telemetry, and observability solutions. * Improve system reliability and scalability. ## Related Videos - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Stop Using Node.js Like It’s 2020! - Alfonso Graziano](https://www.wearedevelopers.com/videos/1863-stop-using-node-js-like-it-s-2020-alfonso-graziano) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)