> Markdown version of [/jobs/ext/3530341-hpc-ai-performance-engineer](https://www.wearedevelopers.com/jobs/ext/3530341-hpc-ai-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # HPC & AI Performance Engineer - **Company:** Hewlett-Packard Enterprise - **Location:** Houston, TX, United States - **Experience:** Expert - **Salary:** $105,500.0 - $243,000.0 - **Contract:** Internship / Graduate position - **Skills:** C (Programming Language), Artificial Intelligence, C++ (Programming Language), Compilers, Profiling, Nvidia CUDA, Software Debugging, Linux, Microprocessors, Fortran (Programming Language), Python (Programming Language), OpenMP, Performance Tuning, Shell Script, System Software, Graphics Processing Unit (GPU), Large Language Models, Distributed Programming, Parallel Computation, Information Technology, Deepseek (AI Talent Sourcing Platform), Engineering Base - **Published:** October 2, 2026 - **Apply:** https://hpe.wd5.myworkdayjobs.com/Jobsathpe/job/Houston-Texas-United-States-of-America/HPC---AI-Performance-Engineer_1215573-2 ## About the Role * 6+ years of Experience with HPC and AI workloads, including scientific and engineering applications, MLPerf, large language models such as DeepSeek, Kimi K2.6 or K3, and gpt-oss-120b, AI training and inference, storage benchmarks, and performance-profiling tools. Relevant coursework and internship experience will be considered. * Knowledge of HPC and AI system architecture, including CPUs, GPUs and other accelerators, memory, networking, storage, and software stacks, with the ability to explain their impact on application and benchmark performance. * Understanding of parallel and distributed programming techniques, including MPI, OpenMP, OpenSHMEM, algorithms, and performance considerations for HPC and AI workloads. * Ability to lead complex HPC and AI performance projects, work effectively across technical teams, and translate findings into customer-focused recommendations. * Demonstrated ability to analyze and optimize computational applications and benchmarks on Linux-based HPC and AI systems. * Experience with HPC and AI software environments, including C, C++, Fortran, Python, Linux scripting, compilers, MPI, MPI-IO, OpenMP, and relevant AI frameworks and libraries. * Experience with NVIDIA GPUs and AMD Instinct MI-series accelerators, including offloading computational routines using CUDA, HIP, OpenMP target offload, OpenACC, or comparable programming models for HPC and AI workloads. * Experience using performance-profiling, tracing, and debugging tools on Linux-based HPC and AI systems. * Ability to interpret benchmark results, identify performance bottlenecks, and recommend improvements for HPC and AI applications and systems. * Excellent analytical and problem-solving skills. * Excellent written and verbal communication skills, with the ability to present complex technical findings clearly to engineering teams, customers, and business stakeholders; professional proficiency in English required. * Master's degree in computer science, engineering, mathematics, physics, chemistry, environmental science, or a related technical field; PhD preferred. ## Description This role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office., * Lead HPC and AI benchmarking projects across CPU, GPU, network, memory, and storage platforms. * Evaluate HPE and competitive HPC and AI architectures using performance models and benchmark data. * Run and analyze scientific, engineering, large language model (LLM), AI training and inference, storage, and I/O workloads. * Identify performance bottlenecks and optimize applications, AI frameworks, libraries, compilers, runtimes, and system software. * Apply parallel computing, GPU acceleration, profiling, and memory and I/O optimization to deliver credible, repeatable results. ## Related Videos - [Just-in-time Compilation in JVM](https://www.wearedevelopers.com/videos/240-just-in-time-compilation-in-jvm) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Building a Compiler with C#](https://www.wearedevelopers.com/videos/116-building-a-compiler-with-c) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) ## Related Articles - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)