Principal - AI & HPC Data Centre Compute

Epam
London, UK
17 days ago
Apply on www.apply4u.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Data Centers Distributed Systems InfiniBand Network Planning and Design Performance Tuning Software Engineering Systems Integration AI Infrastructure Pytorch Large Language Models Parallel Computation
+5 more
AI Platforms Kubernetes Slurm Machine Learning Operations Data Pipelines

Job description

We’re looking for a Director - AI & HPC Data Centre Compute to join our team in London, United Kingdom in a hybrid working mode. In this senior leadership role, you will drive EPAM’s AI and HPC data centre strategy, leading engagements that optimise compute infrastructure across full-stack environments-from accelerators and networking to schedulers, containers and AI platforms. You will address cost, scalability and power constraints through software engineering, platform optimisation and systems integration rather than hardware procurement, helping clients deliver performance and efficiency at scale. Responsibilities Advise hyperscalers, neoclouds and enterprises on compute strategies for AI and HPC, including power, cooling and network design Optimise AI training and inference workloads (LLMs, multimodal and scientific AI) across distributed clusters for cost, throughput and latency targets Lead HPC cluster design and orchestration using Slurm, Kubernetes and parallel processing

Requirements

models (MPI) Enhance GPU utilisation by addressing bottlenecks across compute, memory and data pipelines; collaborate with energy teams on power-aware scheduling Architect scalable AI platforms integrating MLOps frameworks, automation tools and reusable assets Develop repeatable AI/HPC offerings, contribute to solution roadmaps and establish strategic partnerships with major ecosystem players Track industry trends in AI factories, GPU economics and liquid cooling technologies, representing EPAM at industry events and thought leadership forums Requirements 12+ years of experience in HPC, AI infrastructure, accelerated computing or distributed systems architecture Deep knowledge of GPU architectures, AI workloads, networking and large-scale cluster operations Hands-on expertise with Slurm, Kubernetes and performance optimisation across multi-node environments Proven ability to advise senior stakeholders and influence technical strategy at C-level Demonstrated experience in pre-sales solutioning and shaping complex technology engagements Nice to have Exposure to NVIDIA ecosystems (DGX, HGX, SuperPOD) or alternative accelerators (AMD or similar) Familiarity with InfiniBand/RoCE networking, PyTorch, high-density rack design, liquid cooling or GPU-as-a-Service deployments We offer EPAM Employee Stock Purchase Plan (ESPP) Protection benefits including life assurance, income protection and critical illness cover Private medical insurance and dental care Employee Assistance Program Competitive group pension plan Cyclescheme, Techscheme and season ticket loans Various perks such as free Wednesday lunch in-office, on-site massages and regular social events Learning and development opportunities including in-house training and coaching, professional certifications, and courses If otherwise eligible, participation in the discretionary annual bonus program If otherwise eligible and hired into a qualifying level, participation in the discretionary Long-Term Incentive (LTI) Program #J-18808-Ljbffr

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.apply4u.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues · World Congress 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

3:23 min

The AI workload technology stack and its components

Lerna Ekmekcioglu Lerna Ekmekcioglu · Europe 2026 Virtual

Videos

See all

Related articles

See all