Staff Infrastructure Engineer - Virtualization

TensorWave Inc.
Las Vegas, NV, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Proxmox Artificial Intelligence Systems Engineering Cloud Computing Computer Networks Data Centers Linux Distributed Data Store Distributed Systems Memory Management Kernel-Based Virtual Machine Network Layer
+16 more
Networking Basics PCI Express Performance Tuning Quick EMUlator (QEMU) Remote Direct Memory Access Ansible Systems Integration Virtual Local Area Networks Virtualization Technology Weka Ceph (Software) Offline Storage Kubernetes Infrastructure Automation Frameworks Bare Metal Terraform

Job description

We are building large-scale, high-performance infrastructure to power next-generation AI workloads. Our platform operates across multiple data centers and supports GPU-intensive environments with demanding requirements around performance, isolation, and scalability., We are looking for a Staff Infrastructure Engineer to lead the design and evolution of our virtualization platform. This role will own how we build, scale, and operate hypervisor infrastructure as we transition from traditional virtualization platforms toward a more flexible, CSP-aligned architecture based on KVM/QEMU and modern Linux primitives., * Design and implement a scalable virtualization platform capable of supporting high-density compute and GPU workloads

  • Lead the evolution from existing platforms (e.g., Proxmox) toward KVM/QEMU-based architectures
  • Define standards for VM lifecycle management (provisioning, scheduling, migration), performance isolation and resource allocation, failure domains and resilience strategies
  • Optimize virtualization for high-performance workloads, including NUMA alignment, CPU pinning and scheduling, PCIe topology awareness, GPU passthrough and device assignment
  • Partner closely with networking and storage teams to integrate high-throughput, networking (e.g., SR-IOV, RDMA), distributed and local storage systems
  • Build and improve automation for hypervisor deployment and configuration, image pipelines, cluster scaling and lifecycle management
  • Troubleshoot deep system-level performance issues across compute, memory, storage, and network layers
  • Contribute to long-term platform architecture and infrastructure strategy, All offers of employment are contingent upon verification of identity and authorization to work in United States, as required by law.

Background Checks

Where permitted by law, employment may be contingent upon the successful completion of a job-related background check.

Data Privacy Notice

By submitting an application, you acknowledge that TensorWave may collect, use, and retain your personal information for recruiting and employment-related purposes in accordance with applicable data privacy laws.

Requirements

Do you have experience in Virtualization?, * 7+ years of experience in infrastructure, systems engineering, or platform engineering

  • Deep experience with Linux-based virtualization, including:
  • KVM/QEMU
  • libvirt or similar tooling
  • Strong understanding of:
  • CPU scheduling and NUMA architectures
  • Memory management and performance tuning
  • Storage I/O paths and performance characteristics
  • Experience designing and operating virtualization platforms at scale (hundreds+ hosts)
  • Solid networking fundamentals, including:
  • Linux networking (bridges, bonding, VLANs)
  • High-performance networking concepts
  • Experience with infrastructure automation (e.g., Ansible, Terraform, or similar)
  • Strong troubleshooting skills across distributed systems, * Experience in cloud or CSP environments (public or private)
  • Familiarity with:
  • GPU workloads and passthrough (VFIO)
  • SR-IOV and advanced NIC features
  • Experience integrating virtualization with:
  • Kubernetes platforms
  • Bare metal provisioning systems (e.g., MAAS)
  • Exposure to distributed storage systems (e.g., Ceph, Weka, or similar)
  • Experience working in high-performance or low-latency environments

Benefits & conditions

Pulled from the full job description

  • Parental leave
  • 401(k)
  • Health insurance
  • Paid time off
  • Vision insurance
  • Health savings account
  • Dental insurance, * Stock Options
  • 100% paid Medical, Dental, and Vision insurance for Employees
  • Company Health Savings Account Contributions
  • 100% paid Short Term and Long Term Disability Insurance for Employees
  • Life and Voluntary Supplemental Insurance Options
  • Other Insurance Options, such as Pet & Legal Insurance
  • Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support
  • Flexible Spending Account
  • 401(k)
  • Employee Assistance Program
  • Flexible PTO
  • Paid Holidays
  • Parental Leave
  • Other In-Office Perks

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:10 min

Instrumenting supported infrastructure components for automated telemetry scraping

Mathias Palmersheim Mathias Palmersheim · Europe 2026 Virtual

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · WWC Europe 2026

58 sec

Securing uncontaminated AI training data and mandating Linux authentication

Chris Heilmann +1 · LIVE

Videos

See all

Related articles

See all