Principal Software Engineer

Hewlett-Packard Enterprise
Durham, NC, United States
1 day ago
Apply on www.careerbuilder.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Compensation
$160,000.0 - $303,000.0
Working hours
Regular working hours

Tech stack

Multitier Architecture Artificial Intelligence C++ (Programming Language) Cloud Computing Profiling Nvidia CUDA Computer Programming Continuous Integration Software Debugging InfiniBand Python (Programming Language) Language Modeling
+13 more
Remote Direct Memory Access Software Engineering Private Cloud Environment Scripting Graphics Processing Unit (GPU) Enterprise Software Applications Autoscaling Large Language Models Caching Kubernetes Information Technology TensorRT Nim (Programming Language)

Job description

This role has been designed as ‘Hybrid’ with a requirement that you will work on average 2 days per week from an HPE office., HPE’s Private Cloud AI organization is seeking a Principal Software Engineer to lead the model runtime within HPE AI Essentials, the inference platform used by enterprises to operate large language models on infrastructure they own, including air-gapped and sovereign environments. The principal engineering challenge in this domain is not model deployment but sustained execution efficiency: achieving low tail latency and high GPU utilization on customer-owned hardware of varying generation and configuration. In this role you will define the architecture of that runtime - engine integration, batching, KV cache management, and distributed execution - together with the Kubernetes orchestration layer that supports it. The primary work location is as listed, but could be any other HPE site location in the US; however, remote work options will be considered. Responsibilities

· Define and own the technical direction of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution

· Partner with inference performance engineering teams, with accountability for time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency

· Define distributed inferencing strategy, including disaggregated prefill/decode, tensor and pipeline parallelism, KV cache offload across GPU memory, host memory, and RDMA-attached storage

· Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and determine whether each runtime is adopted, developed in-house, or declined

· Define the orchestration layer supporting the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling

· Mentor engineers, lead design and architecture reviews, and present technical direction to business unit and executive audiences Knowledge and Skills, All legitimate job opportunities will come through official company channels, and candidates are responsible for verifying the credentials of any third party claiming to represent the company. Any reliance on fraudulent communication is at the individual’s own risk, and HPE disclaims legal liability for any resulting damages. If you suspect recruitment fraud, do not share personal information or make any payments and report the incident to your local authorities immediately.

Requirements

· Production experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals

· Comprehensive understanding of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding

· Tensor and pipeline parallelism, NCCL collective operations, and the GPU memory hierarchy and interconnect characteristics that govern them

· Expert level proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling

· Strong programming proficiency in Go and Python, with the ability to read, debug, and profile C++/CUDA using tools such as Nsight

· Experience with debugging/profiling multi-tier application workloads such as RAG, Agents, etc

· Excellent analytical, debugging, and problem-solving abilities

Preferred

· Upstream contribution to vLLM, SGLang, TensorRT-LLM, LLM-D, LMCache, or KServe

· Disaggregated prefill/decode serving, or KV cache offload and reuse at scale

· RDMA, GPUDirect Storage, InfiniBand, or RoCE

· MIG, fractional GPU allocation, and multi-tenant GPU isolation

· On-premises, air-gapped, or regulated enterprise software delivery Experience and Education

· Minimum of 12 years of experience in Software Engineering, including +1 years working directly on LLM inference runtimes or production model serving

· Degree in Computer Science or related field, Analysis Skills, Artificial Intelligence (AI), Autoscaling, C++ Programming Language, CUDA (Compute Unified Device Architecture), Caching, Cloud Computing, Computer Programming, Computer Science, Continuous Integration, Debugging Skills, Employee Benefits, Enterprise Applications, GPU (Graphics Processing Unit), Hewlett-Packard Product Family, Inference Engine, Legal, Memory Hardware, Mentoring, Modeling Languages, Multi-tier Architecture, Performance Engineering, Private Cloud, Problem Solving Skills, Python Programming/Scripting Language, Recruiting/Staffing Agency, Return on Capital Employed (ROCE), Risk, Social Media, Software Engineering, Technical Presentation

Benefits & conditions

“The expected salary/wage range for this position is provided below. Actual offer may vary from this range based upon geographic location, work experience, education/training, and/or skill level.

  • United States of America: Annual Salary USD 160,000 - 303,000 in Colorado // 152,000 - 349,000 in North Carolina & Texas The listed salary range reflects base salary. Variable incentives may also be offered.”

About the company

Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · World Congress 2026 Europe

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · World Congress 2024

2:17 min

Comparing code profiling with surface level monitoring

Jérôme Vieilledent · LIVE

Videos

See all

Related articles

See all