Senior Infrastructure Engineer - GPU Compute

Boundless Inc.
United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
1 year minimum
Compensation
$200,000.0 - $250,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Bash Shell Nvidia CUDA Data Centers DevOps Distributed Systems Github General-Purpose Computing on Graphics Processing Units Network Topologies Python (Programming Language) Linux System Administration Machine Learning
+15 more
PCI Express Ansible System Programming TypeScript Pulumi Scripting Kubernetes Infrastructure Automation Frameworks Bare Metal Slurm Terraform Docker Network Optimization Golang Programming Languages

Job description

Boundless is coordinating GPU compute at scale as it becomes a leader in AI. As a Senior Infrastructure Engineer (GPU Compute), you’ll build and operate the compute fabric that powers our AI inference workloads - a large, heterogeneous, globally distributed GPU fleet spanning consumer cards (including RTX 5090) and datacenter hardware. Your job is to keep that fleet full, fast, cheap, and always on: orchestrating workloads across regions and providers, squeezing every bit of performance out of the hardware, and driving down cost per GPU-hour. This role rewards engineers who want to go deep on bare-metal and GPU optimization.

You should be comfortable operating with a high degree of autonomy, navigating ambiguity, and defaulting to a strong bias for action.

What You’ll Do

GPU Fleet Orchestration: Operate a heterogeneous, multi-region GPU fleet (consumer + datacenter, including RTX 5090) using tools like SkyPilot, Kubernetes/k3s, and cloud + on-prem providers. Build the patterns that let us schedule inference workloads across the entire fleet reliably.

Compute Scheduling & Utilization: Maximize GPU utilization across inference workloads. Own workload placement across spot, on-prem, and cloud capacity, keeping the “always-on inference substrate” saturated and economical.

Bare-Metal & GPU Optimization: Go deep on GPU performance - PCIe P2P, ReBAR, NUMA topology (e.g. EPYC SP5), CUDA/driver tuning, memory configuration, and network topology - to push throughput per node.

Requirements

  • 5+ years of infrastructure/DevOps experience operating large-scale production systems
  • Deep expertise in Kubernetes, Docker, and container orchestration at scale
  • Strong Linux systems administration skills
  • Proficiency in infrastructure-as-code tools (Terraform, Ansible, Pulumi)
  • Track record of managing mission-critical, high-throughput systems
  • Strong infrastructure-as-code background in heterogeneous environments
  • Proficiency in at least one common scripting or programming language (Python, Bash, TypeScript, Go, etc.)
  • Comfort navigating ambiguity with a strong bias for action

Nice to Have

  • Experience with GPU computing infrastructure (CUDA, bare-metal optimization, kernel tuning)
  • Experience operating ML training or other large-scale distributed compute infrastructure
  • Experience with GPU fleet orchestration (SkyPilot, Ray, Slurm)
  • Familiarity with fleet access and networking tooling (Tailscale, Teleport)
  • Knowledge of network optimization and topology design
  • Experience with multi-region, globally distributed systems
  • Proficiency in Rust or low-level systems programming
  • Experience with on-premises data center operations

Additional Requirements

  • Candidates must include a public GitHub profile in their application.
  • The GitHub profile should demonstrate a minimum of 1 year of activity/history.
  • Applications that do not include a GitHub profile, or show insufficient activity, will not be considered.

Benefits & conditions

Pulled from the full job description Paid time off Vision insurance Dental insurance, At Boundless, we take care of our people, because building the future of AI compute starts with an empowered team. Here’s what you can expect when you join us:

  • Competitive salary + equity allocation
  • Health, dental, vision (for U.S. employees; region-adjusted globally)
  • Flexible PTO
  • Professional development and conference travel budget
  • Remote-first with regular off-sites and a high-trust, high-velocity team environment

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:15 min

Transitioning to hybrid clouds amid GPU scarcity

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all