> Markdown version of [/jobs/ext/2965307-senior-software-engineer-inference](https://www.wearedevelopers.com/jobs/ext/2965307-senior-software-engineer-inference). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer, Inference - **Company:** Hewlett-Packard Enterprise - **Location:** Spring, TX, United States (Remote available) - **Experience:** Expert - **Salary:** $144,000.0 - $273,000.0 - **Contract:** Permanent contract - **Skills:** Multitier Architecture, Artificial Intelligence, C++ (Programming Language), Nvidia CUDA, Computer Programming, Software Debugging, InfiniBand, Python (Programming Language), Remote Direct Memory Access, Software Engineering, Software Technical Review, Private Cloud Environment, Enterprise Software Applications, Large Language Models, Kubernetes, Information Technology, TensorRT, Nim (Programming Language), Decoding - **Published:** September 17, 2026 - **Apply:** https://hpe.wd5.myworkdayjobs.com/Jobsathpe/job/Spring-Texas-United-States-of-America/Senior-Software-Engineer--Inference_1211356-2 ## About the Role · Familiar with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals · Strong understanding of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding · Working knowledge of tensor and pipeline parallelism, NCCL collective operations, and the GPU memory hierarchy and interconnect characteristics that govern them · Advanced proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling · Strong programming proficiency in Go and Python, with the ability to read, debug, and profile C++/CUDA using tools such as Nsight · Familiar with debugging/profiling multi-tier application workloads such as RAG, Agents · Excellent analytical, debugging, and problem-solving abilities Preferred · Upstream contribution to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe · Disaggregated prefill/decode serving, or KV cache offload and reuse at scale · RDMA, GPUDirect Storage, InfiniBand, or RoCE · MIG, fractional GPU allocation, and multi-tenant GPU isolation · On-premises, air-gapped, or regulated enterprise software delivery, · Minimum of 8 years of experience in Software Engineering, including 1-2+ years working directly on LLM inference runtimes or production model serving · Degree in Computer Science or related field ## Description This role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office., HPE's Private Cloud AI organization is seeking a Senior Software Engineer to build and evolve the model runtime within HPE AI Essentials, the inference platform used by enterprises to operate large language models on infrastructure they own, including air-gapped and sovereign environments. The core engineering challenge in this domain is not model deployment but sustained execution efficiency: achieving low tail latency and high GPU utilization on customer-owned hardware of varying generation and configuration. In this role you will design and implement key components of that runtime - engine integration, batching, KV cache management, and distributed execution - together with the Kubernetes orchestration layer that supports it. The primary work location is as listed, but could be any other HPE site location in the US; however, remote work options will be considered. Responsibilities · Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution · Partner with inference engineering teams and contribute to improving time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency · Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage · Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and make well-supported recommendations on adoption · Contribute to the orchestration layer supporting the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling · Triage and resolve customer issues end-to-end, identifying root causes and improving systems and processes to prevent recurrence · Provide insightful code and design reviews, mentor team members, and lead by example on engineering practices within the team ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [Challenges and Solutions for Efficient, Large-Scale Video Analysis](https://www.wearedevelopers.com/videos/2022-challenges-and-solutions-for-efficient-large-scale-video-analysis) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)