HPC Engineer

Coreweave Uk Ltd.
London, UK
2 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
£79,000.0 - £131,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Intelligent Platform Management Interface Bash Shell Big Data Network Operating System (NOS) Command-Line Interface Program Optimization Data Centers Software Debugging Linux Firmware IBM Hardware Management Console
+17 more
InfiniBand Networking Hardware Python (Programming Language) Linux System Administration Networking Basics Remote Direct Memory Access Ansible Prometheus Software Engineering Diagnostic Tools Scripting Graphics Processing Unit (GPU) High Performance Computing Grafana Kubernetes Low Latency Golang

Job description

We are looking for an HPC Engineer to join our team to deploy, operate, and support NVLink/NVSwitch platforms across large data centre environments. This role is a strong fit for engineers who enjoy production troubleshooting, hardware-adjacent systems work, automation, observability, and learning specialized infrastructure deeply. You will be responsible for troubleshooting Linux, networking, hardware, firmware, performance, and stability issues in production, while building automation to improve runbooks, dashboards, alerts, and lifecycle workflows. Additionally, you will participate in rotating on-call shifts, lead incident responses, conduct root cause analyses, and collaborate cross-functionally across CoreWeave to ensure reliable workflows scale effectively as our global fleet grows.

Requirements

  • Strong Linux system administration and engineering troubleshooting skills.
  • Solid grasp of networking fundamentals and common diagnostic/troubleshooting tools.
  • Hands-on production debugging experience using logs, metrics, and command-line interfaces.
  • Technical experience troubleshooting server, network, GPU, or data centre hardware.
  • Practical scripting or automation experience using Python, Go, Bash, or similar languages.
  • Clear written and verbal communication, documentation skills, and readiness to participate in an on-call rotation.
  • High curiosity to deeply learn specialized GPU interconnect technologies such as NVLink, NVSwitch, and InfiniBand.

Preferred:

  • Experience with Ansible or other infrastructure-as-code and configuration automation tooling.
  • Kubernetes application development or live platform operations experience.
  • Familiarity with modern observability systems, including Grafana, Prometheus, PromQL, or similar stack components.
  • Experience managing large fleet operations across Linux systems, network devices, GPUs, or infrastructure components.
  • Deep understanding of InfiniBand, RDMA, HPC networking, or low-latency/high-bandwidth fabrics.
  • Experience with BMC, Redfish, IPMI, firmware lifecycle management, or hardware management APIs.
  • Exposure to NVLink, NVSwitch, NVIDIA GPU platforms, NVUE, SONiC, or specialized network operating systems.

Wondering if you’re a good fit?

We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams even if you aren’t a 100% skill or experience match. Here are a few qualities we’ve found compatible with our team. If some of this describes you, we’d love to talk.

  • You love to dive headfirst into production troubleshooting, hardware-adjacent systems work, and bringing robust automation to infrastructure at scale.
  • You’re curious about specialized GPU interconnect technologies, high-bandwidth platforms, and continuous system optimization.
  • You’re an expert in driving assigned work to completion with clear communication, thoughtful prioritisation, and keeping operations running smoothly.

Benefits & conditions

At CoreWeave, we work hard, have fun, and move fast! We’re in an exciting stage of hyper-growth that you will not want to miss out on. We’re not afraid of a little chaos, and we’re constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values:

  • Be Curious at Your Core
  • Act Like an Owner
  • Empower Employees
  • Deliver Best-in-Class Client Experiences
  • Achieve More Together

We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and enables the development of innovative solutions to complex problems. As we get set for takeoff, the organisation’s growth opportunities are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us!

The base salary range for this role is £79,000 to £131,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).

To fulfill our obligation to protect client data, successful applicants offered employment with CoreWeave will be required to complete a basic criminal record check, conducted in compliance with GDPR. Employment offers are conditional upon receiving satisfactory check results.

What We Offer

In addition to a competitive salary, we offer a variety of benefits to support your needs, including:

  • Family-level Medical Insurance
  • Family-level Dental Insurance
  • Generous Pension Contribution
  • Life Assurance at 4x Salary
  • Critical Illness Cover
  • Employee Assistance Programme
  • Tuition Reimbursement
  • Work culture focused on innovative disruption

About the company

CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com., CoreWeave is building and operating some of the largest GPU infrastructure in the world. The Metal Net team owns the high-bandwidth GPU interconnect platforms that make large-scale AI and HPC workloads possible, including NVLink and NVSwitch-based systems. We deploy, operate, troubleshoot, and improve these platforms across our global data centre footprint to provide a powerful alternative to traditional hyperscalers.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all