> Markdown version of [/jobs/ext/2881120-tech-lead-manager-gpu-cluster-infrastructure](https://www.wearedevelopers.com/jobs/ext/2881120-tech-lead-manager-gpu-cluster-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Tech Lead Manager, GPU Cluster Infrastructure - **Company:** Far AI, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Software Debugging, Software Design Documents, Linux, Distributed Computing Environment, Distributed Data Store, Fault Tolerance, InfiniBand, Python (Programming Language), Node.Js, Role-Based Access Control, Ansible, Prometheus, Weka, Ceph (Software), Pytorch, Kubernetes, Storage Technologies, Slurm, Terraform - **Published:** September 13, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=6316aa1d82f6277e ## About the Role * You've led engineers as a manager, tech lead, or project lead, setting technical direction, scoping work, and giving feedback. You want to manage people directly; prior direct reports aren't required. * You have 5+ years in systems or infrastructure engineering on production Linux, running GPU, HPC, or large-scale batch platforms, and you've owned systems from design through operation. You have depth in at least one of scheduling, storage, networking, security, or GPU systems, and enough breadth to review designs in the rest. * You've run production Kubernetes for GPU workloads with a batch layer on top (Slurm, Kueue, Volcano, or similar), including quotas, priority and preemption, and node health. * You've owned infrastructure as code and observability for a production fleet, provisioning with Terraform or Ansible, deploying with Helm and ArgoCD, and monitoring with Prometheus, or their equivalents. * You're a strong programmer in at least one language that infrastructure is commonly written in, such as Python, Go, Rust, or C++, and your automation and services are maintained as shared code. * You write clearly for engineers, researchers, and providers, whether it's a roadmap, a design doc, an incident summary, or an escalation. ## Description You'd be the infrastructure sub-team's lead and one of its engineers, with at least half your time on technical work. You'd own the platform's technical direction and roadmap, deciding what we build and how we run it in collaboration with the research teams and the rest of Foundations, and you'd hire and grow the team. In frontier AI research, working out the infrastructure is often part of the science. You'd work directly with researchers and other engineers to keep our large-scale experiments performant and fault-tolerant. We're also hiring a Software Engineer, GPU Cluster Infrastructure for this team. If you want the hands-on work without the management responsibilities, take a look there instead. What you'll do * Set the platform's technical direction and own its roadmap. You decide which systems we run, how we schedule and store across providers, and what we measure, from utilization and queue wait to failure rates. * Stay in the work. You own architecture and the scheduling and storage design, and you debug the failures that cross layers, such as node health, GPU and fabric faults, and multi-node job hangs. * Hire and grow a small team of senior engineers. You set priorities and ownership, scope projects, and give regular feedback and coaching. * Set security direction for a shared cluster where AI agents run experiments, covering identity and access, workload isolation, and sandboxing. * Decide how we operate, from the on-call rotation and incident response to postmortems, fault tolerance, and observability, and take part in it alongside the team. * Be the escalation point for research teams and for providers, and turn recurring problems into platform fixes., * Distributed training infrastructure: multi-node PyTorch and NCCL debugging, the NVIDIA node stack (drivers, GPU Operator, DCGM), InfiniBand or RoCE fabrics, topology-aware placement. * Distributed storage: VAST, Weka, Lustre, Ceph, or object storage at scale; checkpoint I/O. * Cluster security: admission control, RBAC, node and container hardening, sandboxed runtimes (gVisor, Kata, Firecracker), and isolating autonomous agents on shared infrastructure. * Scheduler internals: Kubernetes scheduler plugins or custom controllers, gang scheduling, fair-share and quota, and the utilization, fairness, and latency tradeoffs between them. * Multi-provider platforms: scheduling and storage across clusters at different providers so users see one system, including clusters with no shared network and uneven data locality. * Greenfield team-building: you've hired senior engineers and stood up on-call, incident, and review practices for a new team before. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Discover the open source trio you didn’t expect: .NET and PostgreSQL on Linux](https://www.wearedevelopers.com/videos/2042-discover-the-open-source-trio-you-didn-t-expect-net-and-postgresql-on-linux) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Should Tech Managers Be Developers First? Pros and Cons](https://www.wearedevelopers.com/magazine/327-should-tech-managers-be-developers-first-pros-and-cons) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [From developer to manager – what does it take to become an engineering manager?](https://www.wearedevelopers.com/magazine/42-from-developer-to-manager-what-does-it-take-to-become-an-engineering-manager)