> Markdown version of [/jobs/ext/2590930-senior-infrastructure-engineer-gpu-compute](https://www.wearedevelopers.com/jobs/ext/2590930-senior-infrastructure-engineer-gpu-compute). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Infrastructure Engineer - GPU Compute - **Company:** Boundless Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $200,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Nvidia CUDA, Data Centers, DevOps, Distributed Systems, Github, General-Purpose Computing on Graphics Processing Units, Network Topologies, Python (Programming Language), Linux System Administration, Machine Learning, PCI Express, Ansible, System Programming, TypeScript, Pulumi, Scripting, Kubernetes, Infrastructure Automation Frameworks, Bare Metal, Slurm, Terraform, Docker, Network Optimization, Golang, Programming Languages - **Published:** August 2, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=b2610df12d837d34 ## About the Role * 5+ years of infrastructure/DevOps experience operating large-scale production systems * Deep expertise in Kubernetes, Docker, and container orchestration at scale * Strong Linux systems administration skills * Proficiency in infrastructure-as-code tools (Terraform, Ansible, Pulumi) * Track record of managing mission-critical, high-throughput systems * Strong infrastructure-as-code background in heterogeneous environments * Proficiency in at least one common scripting or programming language (Python, Bash, TypeScript, Go, etc.) * Comfort navigating ambiguity with a strong bias for action Nice to Have * Experience with GPU computing infrastructure (CUDA, bare-metal optimization, kernel tuning) * Experience operating ML training or other large-scale distributed compute infrastructure * Experience with GPU fleet orchestration (SkyPilot, Ray, Slurm) * Familiarity with fleet access and networking tooling (Tailscale, Teleport) * Knowledge of network optimization and topology design * Experience with multi-region, globally distributed systems * Proficiency in Rust or low-level systems programming * Experience with on-premises data center operations Additional Requirements * Candidates must include a public GitHub profile in their application. * The GitHub profile should demonstrate a minimum of 1 year of activity/history. * Applications that do not include a GitHub profile, or show insufficient activity, will not be considered. ## Description Boundless is coordinating GPU compute at scale as it becomes a leader in AI. As a Senior Infrastructure Engineer (GPU Compute), you'll build and operate the compute fabric that powers our AI inference workloads - a large, heterogeneous, globally distributed GPU fleet spanning consumer cards (including RTX 5090) and datacenter hardware. Your job is to keep that fleet full, fast, cheap, and always on: orchestrating workloads across regions and providers, squeezing every bit of performance out of the hardware, and driving down cost per GPU-hour. This role rewards engineers who want to go deep on bare-metal and GPU optimization. You should be comfortable operating with a high degree of autonomy, navigating ambiguity, and defaulting to a strong bias for action. What You'll Do GPU Fleet Orchestration: Operate a heterogeneous, multi-region GPU fleet (consumer + datacenter, including RTX 5090) using tools like SkyPilot, Kubernetes/k3s, and cloud + on-prem providers. Build the patterns that let us schedule inference workloads across the entire fleet reliably. Compute Scheduling & Utilization: Maximize GPU utilization across inference workloads. Own workload placement across spot, on-prem, and cloud capacity, keeping the "always-on inference substrate" saturated and economical. Bare-Metal & GPU Optimization: Go deep on GPU performance - PCIe P2P, ReBAR, NUMA topology (e.g. EPYC SP5), CUDA/driver tuning, memory configuration, and network topology - to push throughput per node. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)