> Markdown version of [/jobs/ext/2879461-senior-software-engineer-gpu-cluster-infrastructure](https://www.wearedevelopers.com/jobs/ext/2879461-senior-software-engineer-gpu-cluster-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer, GPU Cluster Infrastructure - **Company:** Far AI, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $150,000.0 - $275,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, C++ (Programming Language), Computer Clusters, Computer Networks, Software Debugging, Software Design Documents, Linux, Distributed Computing Environment, Distributed Data Store, Fault Tolerance, InfiniBand, Python (Programming Language), Node.Js, Role-Based Access Control, Ansible, Prometheus, Runbook, Weka, Ceph (Software), Graphics Processing Unit (GPU), Pytorch, Kubernetes, Bare Metal, Slurm, Terraform - **Published:** September 13, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a6c17340c7fa481c ## About the Role * You have 3+ years in systems or infrastructure engineering on production Linux, running GPU, HPC, or large-scale batch platforms, and you've owned at least one system from design through operation. * You've run production Kubernetes for GPU workloads with a batch layer on top (Slurm, Kueue, Volcano, or similar), including quotas, priority and preemption, and node health. * You've owned infrastructure as code and observability for a production fleet, provisioning with Terraform or Ansible, deploying with Helm and ArgoCD, and monitoring with Prometheus, or their equivalents. * You're a strong programmer in at least one language that infrastructure is commonly written in, such as Python, Go, Rust, or C++, and your automation and services are maintained as shared code. * You write clearly for engineers, researchers, and providers, whether it's a design doc, an incident summary, or an escalation. ## Description Foundations is FAR.AI's infrastructure and engineering team. Our remit is broad: we run the compute platform, build the tools and frameworks researchers work in, automate research workflows, and help teams scale experiments well past what they'd manage alone. Our job is to accelerate the research. We do so by working directly with researchers through embedded engagements and day-to-day consulting, and building systems that can scale with the organization as it grows. Foundations is growing quickly, and our infrastructure portfolio is growing fastest. We run FAR.AI's research on a mix of bare-metal and managed Kubernetes GPU clusters from multiple providers. We rent the hardware and operate the platform ourselves. The fleet has grown from dozens to hundreds of GPUs this year and it's continuing to grow quickly: we're adding providers, taking on users beyond our own researchers, and moving experiments onto frontier open-weight models. A large amount of research is now being done by AI agents working directly on the cluster, which is driving updates to our platform infrastructure and security. Running it well now takes dedicated specialists, so we're standing up an infrastructure sub-team that owns the cluster fleet: adding capacity, designing and managing the networking and storage under it, infrastructure as code, and the security posture, plus some of the platform layer above it. It works directly with research teams as their needs change. About the Role You'd work across the whole infrastructure stack, from scheduling to storage to monitoring to security, and bring real depth in at least one part of it. We're particularly interested in experience with large-scale pre-training and post-training infrastructure and the network fabric under it, cluster security and sandboxing, distributed storage systems, and batch scheduling for large GPU clusters. Expertise in an adjacent area is also a good fit. In frontier AI research, working out the infrastructure is often part of the science. You'd work directly with researchers and other engineers to keep our large-scale experiments performant and fault-tolerant. We're also hiring a Tech Lead Manager, GPU Cluster Infrastructure for this team. If leading a small team while staying hands-on sounds like you, take a look there instead. What you'll do * Operate the Kubernetes GPU fleet day to day. You handle node lifecycle, upgrades, driver and image rollouts, staged changes with safe rollback, and capacity planning. * Own batch scheduling and multi-tenancy, including queues, quotas, priorities, preemption, gang scheduling, and fair share across research teams. * Design and run the storage under the fleet, from high-performance shared filesystems for datasets and checkpoints to object storage tiers, quotas, and backups. * Keep multi-node training runs fault-tolerant. You own node health and automated draining, debug NCCL and fabric problems, track down stragglers and flaky GPUs, and build the checkpoint and restart patterns. * Harden the platform, covering identity and access, network policy, secrets, workload isolation, and sandboxing for the AI agents that run on the cluster. * Bring new capacity online. You acceptance-test providers on fabric, NCCL, and storage throughput, hold them to their SLAs, and integrate new clusters into the platform with infrastructure as code. * Work directly with research teams on their infrastructure problems and turn the recurring ones into platform fixes. Share the on-call rotation, runbooks, and postmortems., * Distributed training infrastructure: multi-node PyTorch and NCCL debugging, the NVIDIA node stack (drivers, GPU Operator, DCGM), InfiniBand or RoCE fabrics, topology-aware placement. * Distributed storage: VAST, Weka, Lustre, Ceph, or object storage at scale; checkpoint I/O. * Cluster security: admission control, RBAC, node and container hardening, sandboxed runtimes (gVisor, Kata, Firecracker), and isolating autonomous agents on shared infrastructure. * Scheduler internals: Kubernetes scheduler plugins or custom controllers, gang scheduling, fair-share and quota, and the utilization, fairness, and latency tradeoffs between them. * Multi-provider platforms: scheduling and storage across clusters at different providers so users see one system, including clusters with no shared network and uneven data locality. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Discover the open source trio you didn’t expect: .NET and PostgreSQL on Linux](https://www.wearedevelopers.com/videos/2042-discover-the-open-source-trio-you-didn-t-expect-net-and-postgresql-on-linux) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)