> Markdown version of [/jobs/ext/2726778-hpc-engineer](https://www.wearedevelopers.com/jobs/ext/2726778-hpc-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # HPC Engineer - **Company:** Coreweave Uk Ltd. - **Location:** London, UK - **Salary:** £79,000.0 - £131,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Intelligent Platform Management Interface, Bash Shell, Big Data, Network Operating System (NOS), Command-Line Interface, Program Optimization, Data Centers, Software Debugging, Linux, Firmware, IBM Hardware Management Console, InfiniBand, Networking Hardware, Python (Programming Language), Linux System Administration, Networking Basics, Remote Direct Memory Access, Ansible, Prometheus, Software Engineering, Diagnostic Tools, Scripting, Graphics Processing Unit (GPU), High Performance Computing, Grafana, Kubernetes, Low Latency, Golang - **Published:** September 5, 2026 - **Apply:** https://startup.jobs/hpc-engineer-metal-net-coreweave-europe-8136056 ## About the Role * Strong Linux system administration and engineering troubleshooting skills. * Solid grasp of networking fundamentals and common diagnostic/troubleshooting tools. * Hands-on production debugging experience using logs, metrics, and command-line interfaces. * Technical experience troubleshooting server, network, GPU, or data centre hardware. * Practical scripting or automation experience using Python, Go, Bash, or similar languages. * Clear written and verbal communication, documentation skills, and readiness to participate in an on-call rotation. * High curiosity to deeply learn specialized GPU interconnect technologies such as NVLink, NVSwitch, and InfiniBand. Preferred: * Experience with Ansible or other infrastructure-as-code and configuration automation tooling. * Kubernetes application development or live platform operations experience. * Familiarity with modern observability systems, including Grafana, Prometheus, PromQL, or similar stack components. * Experience managing large fleet operations across Linux systems, network devices, GPUs, or infrastructure components. * Deep understanding of InfiniBand, RDMA, HPC networking, or low-latency/high-bandwidth fabrics. * Experience with BMC, Redfish, IPMI, firmware lifecycle management, or hardware management APIs. * Exposure to NVLink, NVSwitch, NVIDIA GPU platforms, NVUE, SONiC, or specialized network operating systems. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk. * You love to dive headfirst into production troubleshooting, hardware-adjacent systems work, and bringing robust automation to infrastructure at scale. * You're curious about specialized GPU interconnect technologies, high-bandwidth platforms, and continuous system optimization. * You're an expert in driving assigned work to completion with clear communication, thoughtful prioritisation, and keeping operations running smoothly. ## Description We are looking for an HPC Engineer to join our team to deploy, operate, and support NVLink/NVSwitch platforms across large data centre environments. This role is a strong fit for engineers who enjoy production troubleshooting, hardware-adjacent systems work, automation, observability, and learning specialized infrastructure deeply. You will be responsible for troubleshooting Linux, networking, hardware, firmware, performance, and stability issues in production, while building automation to improve runbooks, dashboards, alerts, and lifecycle workflows. Additionally, you will participate in rotating on-call shifts, lead incident responses, conduct root cause analyses, and collaborate cross-functionally across CoreWeave to ensure reliable workflows scale effectively as our global fleet grows. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)