Platform Architect - Nvidia AI/GB300
Oscar Associates (UK) Ltd
London, UK
16 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.reed.co.uk
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
£182,000.0 - £198,900.0
Working hours
Regular working hours
Job source
Tech stack
Artificial Intelligence
Computing Platforms
Bash Shell
Continuous Integration
Data Centers
Linux
InfiniBand
Python (Programming Language)
Remote Direct Memory Access
Ansible
High Performance Computing
IT Architecture
+5 more
Git
Kubernetes
Slurm
Hardware Infrastructure
Terraform
Requirements
The key focus is NVIDIA HGX GB300 / NVL72 and RTX 6000 series servers, so proven experience designing and deploying NVIDIA GPU infrastructure is essential
You’ll translate complex technical requirements into HLDs, LLDs and platform architecture, working across infrastructure, networking, storage, security, applications and data-centre teams. Key Requirements
- Strong Platform / Infrastructure Architecture experience across compute, storage, networking and Linux.
- Expert-level Kubernetes architecture experience.
- Strong Slurm and Run experience - essential.
- Proven experience with GPU/HPC environments and large-scale AI platforms.
- Hands-on experience with NVIDIA HGX GB300 / NVL72, including NVLink, NVSwitch and Grace Blackwell architecture.
- Experience with NVIDIA RTX 6000 series GPU servers.
- Strong understanding of GPU workload scheduling, partitioning and sharing, including MIG, vGPU and time-slicing
- Strong understanding of InfiniBand, RoCE, Spectrum-X, GPUDirect RDMA/Storage and high-performance AI fabrics.
- Experience with Terraform, Ansible, Python/shell, Git and CI/CD
Highly Desirable
- Experience with NVIDIA AI Factory / GB300 OR GB200 reference architectures.
- Understanding of rack-scale liquid cooling, 100kW+ rack power densities, CDUs and data centre infrastructure.
- Experience working with colocation/facilities teams on high density GPU deployments.
- Knowledge of DGX BasePOD / SuperPOD architectures.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.reed.co.uk
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
CS
Christina Schaireiter
3 months ago
LM
Luis Minvielle
7 Cloud Computing Trends Coming in 2025 for Developers
over 2 years ago
MH
Michael Hunger
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?
8 months ago
ER
Erin Rifkin
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
about 1 year ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
22 days ago