> Markdown version of [/jobs/ext/2260217-solutions-architect](https://www.wearedevelopers.com/jobs/ext/2260217-solutions-architect). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Solutions Architect - **Company:** WNTD Ltd - **Location:** London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computer Clusters, Nvidia CUDA, Information Technology Consulting, Data Centers, Linux, DevOps, Firmware, InfiniBand, Network Layer, Network Architecture, User-Centered Design, Ceph (Software), High Performance Computing, Performance Testing, Hardware Testing, Kubernetes, Information Technology, Slurm, Hardware Infrastructure, Physical Design, Oracle Cloud Infrastructure, Data Pipelines, Docker - **Published:** August 26, 2026 - **Apply:** https://www.collegerecruiter.com/job/2815045391-solutions-architect ## About the Role * Proven experience architecting and delivering NVIDIA GPU clusters at scale (AI/ML/HPC environments). * Strong hands-on understanding of GPU interconnects (NVLink/NVSwitch) and DGX/HGX/SuperPod architectures. * Deep knowledge of InfiniBand and high-performance networking architectures. * Experience with cluster orchestration: Kubernetes, Slurm, PBS, or similar. * Familiarity with AI/ML workload requirements, CUDA, Docker/OCI containers, and NVIDIA software stacks (NCCL, CUDA Toolkit). * Comfort with Linux systems engineering, hardware validation, and troubleshooting across compute/network layers. Soft Skills * Strong communication skills, with the ability to bridge engineering and business discussions. * Comfortable owning architecture decisions and delivering executive-ready documentation. * Ability to work autonomously while coordinating with multi-disciplinary teams. * Problem-solver with strong critical-thinking abilities and a delivery-focused mindset. * Experience with hyperscaler-class deployments or multi-megawatt datacenter environments. * Work with NVIDIA Base Command Manager or similar cluster management tooling. * Exposure to data pipelines, storage systems (Lustre, GPUDirect Storage, Ceph), or AI workflow platforms. * Certifications such as NVIDIA Certified Associate/Expert, Kubernetes certifications (CKA/CKS), or related vendor accreditations. ## Description Talent Solutions Delivery Lead at WNTD | Data Centre & AI Technology Infrastructure Job Specification: Solution Architect - NVIDIA Cluster (End-to-End Design & Validation) Travel: Occasional travel to datacenter sites outside the UK Engagement: Contract Inside IR35 Department: Engineering/Advanced Compute Role Overview We are seeking a highly skilled Solution Architect with deep experience designing, validating, and delivering end-to-end NVIDIA GPU clusters in enterprise and hyperscale environments. This individual will own the full lifecycle of architectural design-from requirements gathering through implementation oversight and performance validation. They will work closely with engineering, networking, DevOps, security, and datacenter operations teams to ensure high-performance, scalable, and resilient GPU infrastructure for AI, HPC, and ML workloads. The role is primarily London-based one day per week, with occasional international travel required to support datacenter design reviews, deployment validation, or site acceptance testing. Key Responsibilities * Lead the architecture of NVIDIA GPU clusters leveraging technologies such as H100/H200, NVLink, NVSwitch, DGX, HGX, or SuperPod-class designs. * Produce high-level and low-level designs (HLD/LLD), including compute, network, storage, and power/cooling considerations. * Validate hardware and platform selections, ensuring alignment with customer requirements and scalability goals. * Design fabric architectures including InfiniBand (200/400Gb), RoCE, and high-performance east-west traffic patterns. * Define and execute validation test plans for GPU cluster performance, resilience, networking throughput, and workload behaviour. * Oversee integration of GPU nodes, networking, and storage systems into the existing datacenter environment. * Collaborate with DevOps/Platform teams to validate cluster orchestration (Kubernetes, Slurm, Bright Cluster Manager, or equivalents). * Validate firmware, drivers, NCCL, CUDA libraries, and container environments for production readiness. Deployment & Delivery Oversight * Provide technical leadership across the full deployment lifecycle. * Partner with datacenter operations to ensure correct rack layouts, cabling, airflow and power design. * Support delivery teams during build-out phases, ensuring the design is executed correctly. * Participate in factory acceptance tests (FAT), site acceptance tests (SAT), and operational readiness reviews. Stakeholder Collaboration * Work closely with internal and external teams including network engineering, platform engineering, procurement, and vendors such as NVIDIA, Mellanox, Supermicro, Dell, or HPE. * Provide technical guidance to customers, partners, and cross-functional engineering teams. * Communicate complex architectural concepts clearly to both technical and non-technical audiences. Documentation & Governance * Produce detailed architecture documents, diagrams, acceptance criteria, and operational runbooks. * Ensure security, compliance, and governance standards are built into the design. * Provide knowledge transfer and training sessions to internal teams where required. Required Skills & Experience * Proven experience architecting and delivering NVIDIA GPU clusters at scale (AI/ML/HPC environments). * Strong hands-on understanding of GPU interconnects (NVLink/NVSwitch) and DGX/HGX/SuperPod architectures. * Deep knowledge of InfiniBand and high-performance networking architectures. * Experience with cluster orchestration: Kubernetes, Slurm, PBS, or similar. * Familiarity with AI/ML workload requirements, CUDA, Docker/OCI containers, and NVIDIA software stacks (NCCL, CUDA Toolkit). * Comfort with Linux systems engineering, hardware validation, and troubleshooting across compute/network layers. Soft Skills * Strong communication skills, with the ability to bridge engineering and business discussions. * Comfortable owning architecture decisions and delivering executive-ready documentation. * Ability to work autonomously while coordinating with multi-disciplinary teams. * Problem-solver with strong critical-thinking abilities and a delivery-focused mindset. * Experience with hyperscaler-class deployments or multi-megawatt datacenter environments. * Work with NVIDIA Base Command Manager or similar cluster management tooling. * Exposure to data pipelines, storage systems (Lustre, GPUDirect Storage, Ceph), or AI workflow platforms. * Certifications such as NVIDIA Certified Associate/Expert, Kubernetes certifications (CKA/CKS), or related vendor accreditations. What We Offer * Hybrid working: 1 day per week in London. * Opportunity to design next-generation high-performance GPU infrastructure. * Exposure to cutting-edge AI compute at scale. Job Details * Seniority level: Mid-Senior level * Employment type: Contract * Job function: Information Technology * Industries: Professional Services, IT Services, IT Consulting ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [WWC24 - Ankit Patel - Unlocking the Future Breakthrough Application Performance and Capabilities with NVIDIA](https://www.wearedevelopers.com/videos/920-wwc24-ankit-patel-unlocking-the-future-breakthrough-application-performance-and-capabilities-with-nvidia) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026) - [What Makes WeAreDevelopers World Congress Different From Every Other Tech Event?](https://www.wearedevelopers.com/magazine/701-what-makes-wearedevelopers-world-congress-different-from-every-other-tech-event)