> Markdown version of [/jobs/ext/2824247-gpu-infrastructure-lead-systems-integrator](https://www.wearedevelopers.com/jobs/ext/2824247-gpu-infrastructure-lead-systems-integrator). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # GPU Infrastructure Lead - Systems Integrator - **Company:** Hamilton Barnes - **Location:** Greater London, UK - **Experience:** Expert - **Salary:** £150,000.0 - **Contract:** Permanent contract - **Skills:** Build Automation, Automation of Tests, Computer Clusters, Nvidia CUDA, Data Centers, Software Debugging, Firmware, Network Topologies, InfiniBand, Performance Tuning, Regression Testing, Software Deployment, Graphics Processing Unit (GPU), Hardware Infrastructure, Service Stack - **Published:** September 10, 2026 - **Apply:** https://www.collegerecruiter.com/job/2840613450-gpu-infrastructure-lead-systems-integrator ## About the Role * Experience deploying GPU clusters at varying scales, from smaller node deployments to environments with hundreds or thousands of GPUs, across both air-cooled and liquid-cooled infrastructure. DLC experience is highly advantageous. * Strong experience troubleshooting complex infrastructure issues, including cabling and topology errors, firmware mismatches, failing optics, thermal throttling, and production performance challenges. * Experience automating infrastructure-as-code deployments, monitoring and alerting stacks, and hardware acceptance and regression testing. * Deep familiarity with the NVIDIA technology stack, including HGX platforms, NVLink/NVSwitch, CUDA-level debugging, and NCCL performance tuning. * Strong leadership skills with the ability to set technical direction, solve complex problems, and guide engineering teams. * Excellent communication skills with the ability to engage effectively with non-infrastructure stakeholders, including traders, lawyers, capacity providers, and regulators. ## Description Join a pioneering compute infrastructure technology company building the platforms and tools that power the rapidly evolving AI compute market. The organisation develops financial and settlement infrastructure for compute, including technology that connects buyers with GPU capacity, while building automated systems to validate, benchmark, and certify large-scale GPU clusters. The successful candidate will take end-to-end ownership of the GPU infrastructure function, developing the frameworks and automation used to validate, benchmark, and certify large-scale GPU clusters. Combining hands-on engineering with technical leadership, they will build deployment and testing tooling, establish infrastructure standards, shape technical strategy alongside the founders, and ultimately build and lead the team responsible for the function., * Design and run cluster validation and certification: performance benchmarks, interconnect testing (NCCL, InfiniBand/RoCE), thermal and power verification, availability monitoring against SLAs. * Build automation for cluster deployment, health checks, and continuous testing so certification scales without headcount scaling with it. * Set the technical standards for what "deliverable compute" means: node configs, network topologies, storage, cooling envelopes. * Work directly with capacity providers (neoclouds, data centre operators) during onboarding, from site walkthroughs to acceptance testing. * Feed what you learn on the ground back into our contract specs, index methodology, and product roadmap. * Hire and lead the infrastructure engineering team as we grow. ## Related Videos - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [The weekly developer show: Boosting Python with CUDA, CSS Updates & Navigating New Tech Stacks](https://www.wearedevelopers.com/videos/1293-the-weekly-developer-show-boosting-python-with-cuda-css-updates-navigating-new-tech-stacks) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 157: CUDA in Python, Gemini Code Assist and Back-dooring LLMs](https://www.wearedevelopers.com/magazine/557-dev-digest-157-cuda-in-python-gemini-code-assist-and-back-dooring-llms) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)