> Markdown version of [/jobs/ext/2049779-platform-engineer](https://www.wearedevelopers.com/jobs/ext/2049779-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Platform Engineer - **Company:** Era4 - **Location:** UK - **Salary:** £93,221.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Systems Engineering, Ubuntu (Operating System), CentOS, Nvidia CUDA, Distributed Systems, Ethernet, InfiniBand, Linux Kernel, Machine Learning, Performance Tuning, Role-Based Access Control, Red Hat Enterprise Linux, Ansible, Data Logging, Graphics Processing Unit (GPU), Cloud Platform System, System Availability, Grafana, AI Platforms, Kubernetes, Slurm, Kibana, Splunk, Data Pipelines - **Published:** August 14, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5838475149 ## About the Role * Experience supporting HPE PCAI or other AI/HPC infrastructure and platforms. * System administration experience with OS's like RHEL/CentOS, Ubuntu, tuning Linux kernel. * Proficiency with Ansible, Nvidia and CUDA toolkits, Kubernetes and container orchestration. * Understanding of automation, monitoring and security with GPU as a service. * Extensive experience in system engineering, platform operations or SRE. * Experience with GPU resource allocation (across instances, GPUs count and time). * Advanced networking skills with High performance networking, troubleshooting and fine tuning. * Familiarity with cloud-based platforms, APIs, and distributed systems. * Understanding of AI/ML concepts and tooling (model training, inference, data pipelines basics). * Experience with monitoring/logging tools (e.g., Grafana, Kibana, Splunk). * Excellent communication skills to interface with both customers and internal / vendor teams. * Good understanding of tools requirements for ML engineers and data scientists, and how to optimise the experience. ## Description We are looking for Platform Engineer (HPC & AI) who can assist in shaping our new Platform team, this role will be customer facing, involve technical troubleshooting, and collaboration with vendor engineering teams to ensure seamless AI platform operations. Responsibilities: * Designing, deploying, and managing large-scale HPC and GPU-accelerated clusters, including NVIDIA based compute environments. * Implementing and administering HPC scheduling and resource-management systems (e.g., Slurm), including GPU partitioning, workload scheduling, and capacity planning. * Architecting and optimising InfiniBand and Ethernet network topologies. * Ensuring high availability and resilience through failover strategies, planned maintenance coordination, and proactive risk mitigation. * Automating provisioning, configuration, monitoring, and operational workflows across multi-vendor HPC hardware and software stacks. * Monitoring real-time performance and leading troubleshooting efforts across compute, storage, interconnect, drivers, and node failures, engaging vendor support for critical issues. * Incident response: node failure management, network issues, driver issues, troubleshooting common issues and then working with vendor support to resolve any critical issues. * Security and access control: Manage user permissions, RBAC, security hardening, data protection. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [WWC24 - Ankit Patel - Unlocking the Future Breakthrough Application Performance and Capabilities with NVIDIA](https://www.wearedevelopers.com/videos/920-wwc24-ankit-patel-unlocking-the-future-breakthrough-application-performance-and-capabilities-with-nvidia) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)