Infrastructure Specialist - Linux, HPC, Infrastructure & Storage

Alexander Ash
London, UK
15 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Artificial Intelligence Backup Devices Bash Shell Computer Clusters Information Technology Consulting Linux Disaster Recovery Monitoring of Systems Python (Programming Language) NetApp Applications Windows PowerShell Ansible
+10 more
Tensorflow Server Virtualization Virtualization Technology High Performance Computing Pytorch Information Technology Data Management Slurm Veeam Terraform

Job description

We are partnering with a leading organization to hire Skilled Infrastructure Platform Specialists to support their advanced IT infrastructure, HPC, and enterprise storage environments. This is an exciting opportunity for candidates with a strong technical background in server, storage, and high-performance computing environments.

Key Responsibilities

  • Install, configure, and maintain physical and virtual servers, including Dell and Nvidia platforms.
  • Design, deploy, and manage enterprise storage (SAN/NAS) and backup environments, ensuring performance, scalability, and compliance.
  • Support HPC and AI compute infrastructures, GPU clusters, and model training platforms.
  • Monitor system health, performance metrics, and proactively resolve incidents to maintain uptime and meet SLA targets.
  • Collaborate with cross-functional teams to deliver seamless operations and support strategic projects.
  • Implement security best practices, patching, and disaster recovery processes.
  • Document procedures, configurations, and incident resolutions for compliance and knowledge sharing.
  • Lead troubleshooting efforts, root cause analysis, and operational improvements, providing guidance to junior engineers.

Required Skills & Experience

  • Strong technical expertise across server, storage, networking, and HPC infrastructure.
  • Experience with Dell EMC, NetApp, HPE, Nvidia, and virtualization technologies.
  • Linux and Windows system administration experience.
  • Knowledge of backup, disaster recovery, and monitoring tools.
  • Scripting and automation skills (Python, PowerShell, Bash, Ansible, Terraform).
  • Experience with workload schedulers (e.g., SLURM, PBS) and AI/ML frameworks (TensorFlow, PyTorch) is a plus.
  • Excellent analytical, troubleshooting, and collaboration skills.

Qualifications

  • Relevant certifications (ITIL, Dell EMC, NetApp, NVIDIA DLI, Veeam, or equivalent) are highly desirable.

Additional Details

  • Seniority level: Mid-Senior level
  • Employment type: Full-time
  • Job function: Information Technology
  • Industries: IT Services and IT Consulting, IT System Operations and Maintenance
  • Location: London, England, United Kingdom

Requirements

  • Strong technical expertise across server, storage, networking, and HPC infrastructure.
  • Experience with Dell EMC, NetApp, HPE, Nvidia, and virtualization technologies.
  • Linux and Windows system administration experience.
  • Knowledge of backup, disaster recovery, and monitoring tools.
  • Scripting and automation skills (Python, PowerShell, Bash, Ansible, Terraform).
  • Experience with workload schedulers (e.g., SLURM, PBS) and AI/ML frameworks (TensorFlow, PyTorch) is a plus.
  • Excellent analytical, troubleshooting, and collaboration skills., * Relevant certifications (ITIL, Dell EMC, NetApp, NVIDIA DLI, Veeam, or equivalent) are highly desirable.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · World Congress 2025

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:29 min

Tech infrastructure capacity and AI product innovations

Videos

See all

Related articles

See all