> Markdown version of [/jobs/ext/2920080-principal-ai-hpc-data-centre-compute](https://www.wearedevelopers.com/jobs/ext/2920080-principal-ai-hpc-data-centre-compute). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal - AI & HPC Data Centre Compute - **Company:** Epam - **Location:** London, UK - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Data Centers, Distributed Systems, InfiniBand, Network Planning and Design, Performance Tuning, Software Engineering, Systems Integration, AI Infrastructure, Pytorch, Large Language Models, Parallel Computation, AI Platforms, Kubernetes, Slurm, Machine Learning Operations, Data Pipelines - **Published:** September 15, 2026 - **Apply:** https://www.apply4u.co.uk/jobs/principal-ai-hpc-data-centre-compute/47166480 ## About the Role models (MPI) Enhance GPU utilisation by addressing bottlenecks across compute, memory and data pipelines; collaborate with energy teams on power-aware scheduling Architect scalable AI platforms integrating MLOps frameworks, automation tools and reusable assets Develop repeatable AI/HPC offerings, contribute to solution roadmaps and establish strategic partnerships with major ecosystem players Track industry trends in AI factories, GPU economics and liquid cooling technologies, representing EPAM at industry events and thought leadership forums Requirements 12+ years of experience in HPC, AI infrastructure, accelerated computing or distributed systems architecture Deep knowledge of GPU architectures, AI workloads, networking and large-scale cluster operations Hands-on expertise with Slurm, Kubernetes and performance optimisation across multi-node environments Proven ability to advise senior stakeholders and influence technical strategy at C-level Demonstrated experience in pre-sales solutioning and shaping complex technology engagements Nice to have Exposure to NVIDIA ecosystems (DGX, HGX, SuperPOD) or alternative accelerators (AMD or similar) Familiarity with InfiniBand/RoCE networking, PyTorch, high-density rack design, liquid cooling or GPU-as-a-Service deployments We offer EPAM Employee Stock Purchase Plan (ESPP) Protection benefits including life assurance, income protection and critical illness cover Private medical insurance and dental care Employee Assistance Program Competitive group pension plan Cyclescheme, Techscheme and season ticket loans Various perks such as free Wednesday lunch in-office, on-site massages and regular social events Learning and development opportunities including in-house training and coaching, professional certifications, and courses If otherwise eligible, participation in the discretionary annual bonus program If otherwise eligible and hired into a qualifying level, participation in the discretionary Long-Term Incentive (LTI) Program #J-18808-Ljbffr ## Description We're looking for a Director - AI & HPC Data Centre Compute to join our team in London, United Kingdom in a hybrid working mode. In this senior leadership role, you will drive EPAM's AI and HPC data centre strategy, leading engagements that optimise compute infrastructure across full-stack environments-from accelerators and networking to schedulers, containers and AI platforms. You will address cost, scalability and power constraints through software engineering, platform optimisation and systems integration rather than hardware procurement, helping clients deliver performance and efficiency at scale. Responsibilities Advise hyperscalers, neoclouds and enterprises on compute strategies for AI and HPC, including power, cooling and network design Optimise AI training and inference workloads (LLMs, multimodal and scientific AI) across distributed clusters for cost, throughput and latency targets Lead HPC cluster design and orchestration using Slurm, Kubernetes and parallel processing ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) ## Related Articles - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)