Senior Systems Engineer, Virtualization

Coreweave, Inc.
United States
10 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$182,000.0 - $242,000.0
Working hours
Regular working hours
Job source

Tech stack

Abstraction Layers Artificial Intelligence C++ (Programming Language) Profiling Software Debugging Linux File Systems Distributed Systems Memory Management Firmware Hardware Interface Design Hypervisor
+16 more
Kernel-Based Virtual Machine Linux Kernel PCI Express Performance Tuning Quick EMUlator (QEMU) System Programming Virtual Machines Virtualization Technology Loadable Kernel Module Graphics Processing Unit (GPU) Build Management Perf (Linux) AI Platforms Kubernetes Bare Metal Hardware Infrastructure

Job description

HAVOCK builds the software stack that bridges AI workloads and bare metal. We own the operating system, virtualization, runtime, and hardware interfaces that allow thousands of GPU servers to securely execute customer workloads at hyperscale.

This team focuses on the execution layer beneath Kubernetes. We build the systems that provide strong workload isolation, efficient GPU sharing, and high-performance execution across containers and lightweight virtual machines. Our work spans Linux, KVM, container runtimes, GPU drivers, and Kubernetes, ensuring customers can safely run demanding AI workloads on shared infrastructure without sacrificing performance., As a Senior Software Engineer on HAVOCK’s Runtime & Virtualization team, you’ll design and build the execution environment that powers CoreWeave’s AI platform. This is fundamentally a Linux systems engineering role where you’ll work across the Linux kernel, KVM/QEMU, container runtimes, GPU drivers, and Kubernetes to solve problems that don’t have off-the-shelf solutions.

You’ll develop secure sandboxed runtimes for GPU workloads, extend virtualization technologies to support new hardware capabilities, optimize the interaction between Linux, hypervisors, and NVIDIA GPUs, and build the tooling that helps engineers understand what’s happening across the entire software stack.

The work spans multiple abstraction layers. One day you might be debugging a kernel memory-management issue affecting VFIO device passthrough; the next you might be improving container startup latency, extending KubeVirt to support new GPU workflows, or building eBPF tooling to diagnose production networking and scheduling problems. We value engineers who enjoy understanding how systems behave from the hardware up rather than treating infrastructure as a black box.

Some of what you’ll work on:

  • Design secure execution environments using containerd, runc, gVisor, Kata Containers, KubeVirt, and KVM/QEMU.
  • Build GPU-aware runtime infrastructure supporting VFIO, Kata, NVIDIA GPU Operator, and PCIe passthrough for multi-tenant AI workloads.
  • Improve Linux kernel and hypervisor performance through optimization of scheduling, memory management, I/O, NUMA locality, and virtualization primitives.
  • Debug complex interactions across Linux, KVM, GPU drivers, firmware, and Kubernetes when workloads don’t behave as expected.
  • Develop observability and debugging tooling using eBPF, perf, tracepoints, and kernel tracing infrastructure.
  • Improve container and VM startup performance, resource isolation, and runtime efficiency for latency-sensitive AI inference and training workloads.
  • Extend virtualization infrastructure supporting virtio devices, IOMMU, SR-IOV, mediated devices, nested virtualization, and hardware passthrough.
  • Profile production systems and build performance analysis tooling to identify bottlenecks across kernels, hypervisors, container runtimes, storage, networking, and GPUs.
  • Collaborate with security, platform, networking, and GPU infrastructure teams to define the next generation of runtime isolation and workload execution.

Requirements

  • 5+ years building production systems software, platform infrastructure, virtualization, or Linux-based distributed systems.
  • Strong Linux systems knowledge, including namespaces, cgroups, scheduling, memory management, filesystems, networking, and process lifecycle.
  • Experience with virtualization technologies such as KVM, QEMU, VFIO, virtio, Kata Containers, KubeVirt, Firecracker, or gVisor.
  • Experience building or operating Kubernetes platforms and container runtimes at scale.
  • Strong systems programming skills in Go, Rust, C/C++, or a combination thereof.
  • Comfortable debugging production failures that span hardware, operating systems, container runtimes, virtualization, and distributed infrastructure.
  • Experience profiling and optimizing system performance using tools such as perf, eBPF, ftrace, bpftrace, flame graphs, or crash analysis.

Preferred:

  • Linux kernel development or kernel module experience.
  • Experience debugging kernel panics, crash dumps, memory corruption, or driver issues using kdump, crash, drgn, or gdb.
  • Familiarity with NVIDIA GPU drivers, GPU Operator, CDI, and GPU virtualization.
  • Experience contributing to Linux, KVM, QEMU, Kata Containers, gVisor, containerd, or Kubernetes.
  • Understanding of PCIe, IOMMU, DMA, NUMA, interrupts, and modern server hardware architecture.
  • Experience building systems that safely execute untrusted workloads in shared environments.

Benefits & conditions

The base salary range for this role is $182,000 to $242,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility)., In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include:

  • Medical, dental, and vision insurance - 100% paid for by CoreWeave
  • Company-paid Life Insurance
  • Voluntary supplemental life insurance
  • Short and long-term disability insurance
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement
  • Ability to Participate in Employee Stock Purchase Program (ESPP)
  • Mental Wellness Benefits through Spring Health
  • Family-Forming support provided by Carrot
  • Paid Parental Leave
  • Flexible, full-service childcare support with Kinside
  • 401(k) with a generous employer match
  • Flexible PTO
  • Catered lunch each day in our office and data center locations
  • A casual work environment
  • A work culture focused on innovative disruption

About the company

CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · WWC 2021

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · WWC Europe 2026

3:46 min

The history of abstractions and hardware virtualization

Edoardo Dusi · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all