> Markdown version of [/jobs/ext/2652885-cluster-engineer](https://www.wearedevelopers.com/jobs/ext/2652885-cluster-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cluster Engineer - **Company:** STN, inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Application Performance Management, Computing Platforms, Microsoft Azure, Bash Shell, Computer Clusters, Profiling, Nvidia CUDA, Network Congestion, Data Files, Linux, File Systems, Distributed Computing Environment, Distributed Data Store, Memory Management, Ethernet, Firmware, Network Topologies, InfiniBand, Python (Programming Language), Linux System Administration, Metadata, PCI Express, Remote Direct Memory Access, Ansible, Prometheus, AI Infrastructure, Scripting, Graphics Processing Unit (GPU), Google Cloud, High Performance Computing, Pytorch, Large Language Models, Grafana, Parallel Computation, Kubernetes, Infrastructure Automation Frameworks, Storage Technologies, Low Latency, Bare Metal, Slurm, TensorRT, Terraform, Nvme - **Published:** August 5, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=383fef568d0c642b ## About the Role * 7+ years designing or operating large-scale Linux infrastructure. * 5+ years supporting production GPU clusters for AI or HPC workloads. * Demonstrated experience building multi-node GPU training environments from the ground up. * Deep expertise with distributed PyTorch training. * Extensive experience troubleshooting and optimizing NCCL communications. * Strong understanding of distributed AI communication patterns, including: + AllReduce + ReduceScatter + AllGather, * Strong understanding of GPU memory management, including: + KV Cache + Activation checkpointing + Tensor Parallelism + Pipeline Parallelism + Data Parallelism * Experience optimizing LLM inference throughput, including: + Tokens/sec optimization + Batch sizing + Continuous batching + KV cache tuning + Memory bandwidth optimization * Experience tuning CUDA, NCCL, UCX, and MPI for maximum distributed performance. * Expert-level Linux systems administration skills. * Experience with Slurm workload manager. * Experience using Pyxis and Enroot for containerized GPU workloads. * Strong scripting skills using Python and Bash. Technical Expertise AI Frameworks * PyTorch * CUDA * NCCL * Triton (preferred) * TensorRT-LLM (preferred) Cluster Scheduling * Slurm * Pyxis * Enroot GPU Networking Strong understanding of: * InfiniBand * RoCE v2 * RDMA * GPUDirect RDMA * GPUDirect Storage * UCX * MPI * Network topology optimization * Congestion control * QoS * ECN/PFC * High-speed Ethernet (200/400/800 GbE) Storage Experience designing or tuning storage for AI workloads, including: * Parallel file systems * Distributed storage * Object storage * NVMe * Checkpoint optimization * Dataset staging * GPUDirect Storage * Storage bandwidth optimization * Metadata performance Performance Engineering Experience with: * NCCL benchmarking * Multi-node scaling analysis * GPU utilization optimization * Communication/computation overlap * NUMA optimization * CPU affinity * PCIe topology * GPU topology (NVLink/NVSwitch) * Memory bandwidth analysis * End-to-end performance profiling Preferred Qualifications * Experience deploying AI workloads on Kubernetes. * Experience with NVIDIA GPU Operator. * Experience with Kubernetes batch scheduling (Volcano, Kueue, Run:ai, etc.). * Experience with distributed inference platforms such as vLLM, TensorRT-LLM, or SGLang. * Experience with NVIDIA DGX SuperPOD or similar large-scale GPU deployments. * Familiarity with MLPerf benchmarking. * Experience deploying monitoring solutions such as Prometheus, Grafana, and DCGM Exporter. * Experience automating infrastructure using Ansible, Terraform, or similar tools. * Experience working in cloud GPU environments (AWS, Azure, GCP) in addition to bare metal. ## Description We are seeking a highly experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI training and inference workloads. This is a deeply technical role focused on maximizing cluster efficiency, scalability, and performance across the entire AI stack-from GPU hardware and high-speed networking to distributed training frameworks and inference optimization. The ideal candidate has built GPU clusters from the ground up, tuned distributed training environments, optimized large-scale inference deployments, and understands how every layer of the infrastructure contributes to application performance. Responsibilities * Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads. * Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency. * Optimize inference clusters for maximum token generation throughput, low latency, and high GPU utilization. * Build and support production AI infrastructure running hundreds to thousands of GPUs. * Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers. * Perform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance. * Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS. * Configure and tune distributed AI software stacks including: + PyTorch + NCCL + CUDA + UCX + MPI + Slurm + Pyxis/Enroot * Optimize GPU scheduling and resource allocation for both training and inference environments. * Develop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases. * Identify performance regressions and troubleshoot distributed training issues at scale. * Optimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O. * Work closely with ML engineers to improve training scalability and inference efficiency. * Create automation to deploy, validate, benchmark, and monitor GPU clusters. * Evaluate emerging AI infrastructure technologies and recommend improvements to platform architecture. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Discover the open source trio you didn’t expect: .NET and PostgreSQL on Linux](https://www.wearedevelopers.com/videos/2042-discover-the-open-source-trio-you-didn-t-expect-net-and-postgresql-on-linux) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)