Senior Solutions Engineer

LJB & Co
London, UK
1 day ago
Apply on www.totaljobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Part-time / full-time
Experience level
Expert
Experience required
5 years minimum
Compensation
£260,000.0 - £312,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Systems Engineering Cloud Computing Computer Clusters Data Centers InfiniBand Machine Learning Performance Tuning Remote Direct Memory Access AI Infrastructure High Performance Computing Kubernetes
+3 more
Bare Metal Slurm Hardware Infrastructure

Job description

This is a senior technical position where you’ll be responsible for taking complex customer requirements and turning them into practical, scalable infrastructure solutions.

You’ll be working on some seriously advanced environments, designing large-scale NVIDIA GPU platforms, high-speed networking, storage and cloud infrastructure for AI and machine-learning workloads.

The role would suit someone who enjoys being both hands-on technically and customer-facing, rather than someone who wants to sit purely on the engineering side.

What You’ll Be Doing

  • Designing large-scale GPU and AI infrastructure for enterprise customers.
  • Developing technical architecture, system designs and infrastructure proposals.
  • Creating detailed HLDs, LLDs and Bills of Materials.
  • Designing GPU cluster architectures around NVIDIA platforms.
  • Working with NVLink, NVSwitch, InfiniBand and RoCE.
  • Designing high-performance networking and storage environments.
  • Supporting both bare-metal and Kubernetes-based deployments.
  • Working with technologies such as Slurm, Kubernetes, NVIDIA GPU Operator and NCCL.
  • Supporting technical sales opportunities and major customer engagements.
  • Running technical workshops with senior customer stakeholders.
  • Supporting RFP/RFI responses and complex technical proposals.
  • Helping design and validate Proof-of-Concept environments.
  • Troubleshooting performance across GPU, networking and storage infrastructure.
  • Working closely with engineering teams, customers and technology partners., * Kubernetes
  • Slurm
  • NVIDIA GPU Operator
  • NCCL
  • RDMA / GPUDirect
  • High-performance storage
  • GPU cluster networking
  • HLD / LLD production
  • Technical pre-sales and customer-facing architecture

Why This Role?

This is an opportunity to work at the cutting edge of AI infrastructure, rather than traditional cloud or data centre projects.

You’ll be working on high-performance GPU environments supporting AI, machine learning and HPC workloads, with exposure to some of the latest NVIDIA technologies and large-scale infrastructure deployments.

Requirements

We’re looking for a Senior Solutions Engineer with a strong background in GPU, AI or HPC infrastructure to join an exciting and rapidly growing technology business., You’ll ideally have 5+ years’ experience within one or more of the following:

  • Solution Architecture
  • Technical Pre-Sales
  • Systems Engineering
  • HPC
  • AI / Machine Learning Infrastructure
  • GPU Infrastructure
  • Data Centre Infrastructure

You should have strong practical knowledge of NVIDIA GPU environments and understand how large GPU clusters are designed, connected and deployed.

Experience with Blackwell, HGX, DGX, NVLink/NVSwitch, InfiniBand or RoCE would be highly advantageous.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.totaljobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues · World Congress 2025

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

3:05 min

Acquiring Mellanox to build cohesive AI factories

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all