> Markdown version of [/jobs/ext/2488758-senior-ai-platform-engineer](https://www.wearedevelopers.com/jobs/ext/2488758-senior-ai-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior AI Platform Engineer - **Company:** IQVIA LLC - **Location:** London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Profiling, Nvidia CUDA, Information Engineering, Distributed Computing Environment, Distributed Systems, Graph Database, Machine Learning, Performance Tuning, Management of Software Versions, High Performance Computing, Pytorch, Large Language Models, AI Platforms, Kubernetes, Infrastructure Automation Frameworks, Data Lineage, Slurm, Machine Learning Operations, TensorRT, Hardware Infrastructure, Nim (Programming Language), Automation Anywhere - **Published:** August 4, 2026 - **Apply:** https://dejobs.org/x/x/DBF870F003164732831FD91DC88EF631/job/ ## About the Role * Significant experience designing, building, and operating large-scale AI, machine learning, or distributed computing platforms in enterprise environments. * Deep understanding of LLM architectures and their interaction with GPU infrastructure, including CUDA, cuDNN, NCCL, kernel-level acceleration libraries, and distributed training frameworks such as PyTorch. * Strong knowledge of distributed training and inference strategies, including tensor, pipeline, data, and expert parallelism approaches. * Experience optimising LLM inference workloads using technologies such as vLLM, TensorRT-LLM, NVIDIA NIM, SGLang, or similar high-performance serving frameworks. * Expertise in model optimisation techniques including quantisation, mixed precision training and inference (FP8, GPTQ, AWQ, LoRA), and performance tuning for large-scale model deployment. * Advanced experience profiling, troubleshooting, and optimising GPU workloads using tools such as NVIDIA Nsight, DCGM, and related ecosystem technologies. * Strong background in AWS cloud services, high-performance computing, distributed systems, containerised environments, and infrastructure automation. * Experience with workload orchestration technologies such as Slurm, Kubernetes, Ray, or equivalent distributed compute frameworks. * Demonstrated success bridging research and production environments, enabling rapid experimentation while maintaining operational excellence, governance, security, and reliability. * Proven ability to lead complex cross-functional initiatives, influence technical direction, and communicate effectively with engineering, research, product, and executive stakeholders. ## Description The Senior AI Platform Engineer is responsible for defining and delivering the infrastructure strategy underpinning IQVIA's Large Language Model (LLM) programmes. This role provides technical leadership across compute, data, model lifecycle management, evaluation frameworks, and platform engineering, ensuring research innovations can be successfully transformed into secure, scalable, and production-ready AI solutions., Acting as a key technical leader and cross-functional integrator, the Senior AI Platform Engineer partners with research, product, infrastructure, data engineering, and MLOps teams to design and operate the platforms required to train, evaluate, deploy, and govern large-scale AI systems across IQVIA products and healthcare use cases., * Own the AI platform and infrastructure roadmap, leading the planning and execution of LLM initiatives and translating research requirements into scalable engineering solutions. * Partner with centralised infrastructure teams to design and deliver high-performance compute environments across AWS and on-premises platforms, including GPU infrastructure, Slurm clusters, and migration from ad hoc research workflows. * Optimise LLM training and inference workloads, supporting research and product teams in maximising performance, scalability, and reliability across the infrastructure stack. * Establish and maintain model and data lifecycle capabilities, including dataset versioning, lineage tracking, reproducibility standards, and integration with model registries. * Lead the evolution of knowledge graph infrastructure, driving technology selection, migration strategies, performance optimisation, and integration with AI workflows. * Serve as the primary technical coordination point across AI Research, Data Engineering, MLOps, Product, and Infrastructure teams, resolving dependencies and prioritising activities critical to delivery. * Provide technical leadership for vendor selection, procurement, and technology partnerships, advising on compute architectures, GPU specifications, AI platforms, and integration approaches. * Define platform engineering standards, governance, and best practices while mentoring engineers and promoting operational excellence across AI and platform teams. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Architecting the Future: Leveraging AI, Cloud, and Data for Business Success](https://www.wearedevelopers.com/videos/1096-architecting-the-future-leveraging-ai-cloud-and-data-for-business-success) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)