> Markdown version of [/jobs/ext/2714610-hpc-performance-engineer](https://www.wearedevelopers.com/jobs/ext/2714610-hpc-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # HPC Performance Engineer - **Company:** Coreweave, Inc. - **Location:** Bellevue, United States - **Experience:** Expert - **Salary:** $165,000.0 - $242,000.0 - **Contract:** Permanent contract - **Skills:** Systems Engineering, Automation of Tests, Ubuntu (Operating System), Computer Engineering, Software Debugging, Linux, Microprocessors, Distributed Systems, Ethernet, Network Interface Controllers, InfiniBand, Python (Programming Language), Linux Kernel, Machine Learning, Regression Analysis, Open Source Technology, Quick EMUlator (QEMU), Regression Testing, Prometheus, Software Engineering, Systems Architecture, Virtualization Technology, Graphics Processing Unit (GPU), Enterprise Software Applications, Cloud Platform System, High Performance Computing, Grafana, Containerization, Kubernetes, Information Technology, Software Version Control, Docker, Golang, Programming Languages - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/hpc-performance-engineer-coreweave-2-7242302 ## About the Role * 5+ years of professional experience in Systems/HPC Performance Engineering, Benchmarking, and/or Validation. * Bachelor's degree in Computer Engineering, Electrical Engineering, Computer Science, or a related field. * Strong experience with MPI workloads and distributed system performance analysis * Familiarity with RoCE, InfiniBand, and GPUDirect/Data Direct I/O, NUMA, etc in HPC workloads * Hands-on use of public HPC benchmarks (HPCC, HPL, OSU, MLPerf-HPC, STREAM, IO500) * Extensive, deep experience in Linux internals * Fluency with a programming language geared toward automation (Python preferred, but others possible) * Experience writing robust, testable code * Experience diagnosing and fixing systems performance issues * Experiencing with implementing automation testing * Ability to effectively prioritize and communicate proposed features and fixes in a remote-employee environment * Strong passion for automation, with a commitment to automating processes comprehensively * Excellent documentation skills and attention to detail * Strong analytical and problem-solving abilities Preferred: * Familiarity with QA/QE best practices * Familiarity with Golang * Opinions about software version control and team collaboration * Experience working in Cloud environments * Experience as a software engineer writing large-scale applications * Experience in open-source community software development * Experience with machine learning is a huge bonus Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. ## Description CoreWeave is seeking a highly skilled and motivated HPC Performance Engineer to join our HAVOCK Team, reporting into the Manager of Systems Engineering. In this role, you will play a crucial part in the design, development, and optimization of our bare-metal systems from POST through joining a Kubernetes cluster. The team's primary responsibilities include maintaining a custom Linux kernel, various OS images (Ubuntu-based), the virtualization stack (kubevirt/qemu/vfio), and the container/pod runtime stack (containerd/nydus/kubelet). You will collaborate closely with cross-functional teams, up stack engineering teams, and stakeholders to ensure our low-level software stack is performant in the context of hardware updates; and providing data, metrics, dashboards, and analysis to substantiate performance assertions. Kernel Hardware - Acceleration - Virtualization - Operating Systems - Containerization - Kubelet Our Team's Stack: * Python, Go, bash/sh, C * Prometheus, Victoria Metrics, Grafana * Linux Kernel (custom build), Ubuntu * Intel/AMD/ARM CPUs, Nvidia GPUs, DPUs, Infiniband and Ethernet NICs * Docker, kubernetes (k8s), KubeVirt, containerd, kubelet, * Develop and maintain tools for establishing systems performance baselines * Develop and maintain performance regression analysis testing automation * Design and maintain performance regression test pipelines for HPC workloads * Debug and Tune fabric-level performance to ensure low-latency high throughput configurations * Development of telemetry for performance analysis across distributed clusters of servers * Triage and fix performance issues in Linux * Collect data, produce metrics and visualizations that communicate performance information compared to benchmarks; this data should lead to appropriate business decisions and toward greater automation that improves customer experience in relation to performance * Define Linux and OS requirements, specifications, and system architecture in relation to systems performance, in collaboration with cross-functional teams. Along with these responsibilities there will also be cross team collaboration to triage and resolve bottlenecks ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Should senior developers refuse interview coding challenges?](https://www.wearedevelopers.com/magazine/29-should-senior-developers-refuse-interview-coding-challenges) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [How to Answer the Interview Question: “Why Do You Want to Be a Software Engineer?”](https://www.wearedevelopers.com/magazine/392-how-to-answer-the-interview-question-why-do-you-want-to-be-a-software-engineer) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers)