Platform Engineer

Era4
UK
21 days ago
Apply on www.adzuna.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
£93,221.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Systems Engineering Ubuntu (Operating System) CentOS Nvidia CUDA Distributed Systems Ethernet InfiniBand Linux Kernel Machine Learning Performance Tuning
+14 more
Role-Based Access Control Red Hat Enterprise Linux Ansible Data Logging Graphics Processing Unit (GPU) Cloud Platform System System Availability Grafana AI Platforms Kubernetes Slurm Kibana Splunk Data Pipelines

Job description

We are looking for Platform Engineer (HPC & AI) who can assist in shaping our new Platform team, this role will be customer facing, involve technical troubleshooting, and collaboration with vendor engineering teams to ensure seamless AI platform operations.

Responsibilities:

  • Designing, deploying, and managing large-scale HPC and GPU-accelerated clusters, including NVIDIA based compute environments.

  • Implementing and administering HPC scheduling and resource-management systems (e.g., Slurm), including GPU partitioning, workload scheduling, and capacity planning.

  • Architecting and optimising InfiniBand and Ethernet network topologies.

  • Ensuring high availability and resilience through failover strategies, planned maintenance coordination, and proactive risk mitigation.

  • Automating provisioning, configuration, monitoring, and operational workflows across multi-vendor HPC hardware and software stacks.

  • Monitoring real-time performance and leading troubleshooting efforts across compute, storage, interconnect, drivers, and node failures, engaging vendor support for critical issues.

  • Incident response: node failure management, network issues, driver issues, troubleshooting common issues and then working with vendor support to resolve any critical issues.

  • Security and access control: Manage user permissions, RBAC, security hardening, data protection.

Requirements

  • Experience supporting HPE PCAI or other AI/HPC infrastructure and platforms.

  • System administration experience with OS’s like RHEL/CentOS, Ubuntu, tuning Linux kernel.

  • Proficiency with Ansible, Nvidia and CUDA toolkits, Kubernetes and container orchestration.

  • Understanding of automation, monitoring and security with GPU as a service.

  • Extensive experience in system engineering, platform operations or SRE.

  • Experience with GPU resource allocation (across instances, GPUs count and time).

  • Advanced networking skills with High performance networking, troubleshooting and fine tuning.

  • Familiarity with cloud-based platforms, APIs, and distributed systems.

  • Understanding of AI/ML concepts and tooling (model training, inference, data pipelines basics).

  • Experience with monitoring/logging tools (e.g., Grafana, Kibana, Splunk).

  • Excellent communication skills to interface with both customers and internal / vendor teams.

  • Good understanding of tools requirements for ML engineers and data scientists, and how to optimise the experience.

Benefits & conditions

You’ll be joining a mission-driven start-up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next-generation company operates at scale.

Diversity & Inclusion:

Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Note:

We appreciate this is a relatively new skill set and we are open to candidates who may not tick all the boxes but are willing to learn and develop their skillset.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

Videos

See all

Related articles

See all