> Markdown version of [/jobs/ext/2252435-infrastructure-specialist-linux-hpc-infrastructure-storage](https://www.wearedevelopers.com/jobs/ext/2252435-infrastructure-specialist-linux-hpc-infrastructure-storage). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Infrastructure Specialist - Linux, HPC, Infrastructure & Storage - **Company:** Alexander Ash - **Location:** London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Backup Devices, Bash Shell, Computer Clusters, Information Technology Consulting, Linux, Disaster Recovery, Monitoring of Systems, Python (Programming Language), NetApp Applications, Windows PowerShell, Ansible, Tensorflow, Server Virtualization, Virtualization Technology, High Performance Computing, Pytorch, Information Technology, Data Management, Slurm, Veeam, Terraform - **Published:** August 26, 2026 - **Apply:** https://www.collegerecruiter.com/job/2815149895-infrastructure-specialist--linux-hpc-infrastructure--storage ## About the Role * Strong technical expertise across server, storage, networking, and HPC infrastructure. * Experience with Dell EMC, NetApp, HPE, Nvidia, and virtualization technologies. * Linux and Windows system administration experience. * Knowledge of backup, disaster recovery, and monitoring tools. * Scripting and automation skills (Python, PowerShell, Bash, Ansible, Terraform). * Experience with workload schedulers (e.g., SLURM, PBS) and AI/ML frameworks (TensorFlow, PyTorch) is a plus. * Excellent analytical, troubleshooting, and collaboration skills., * Relevant certifications (ITIL, Dell EMC, NetApp, NVIDIA DLI, Veeam, or equivalent) are highly desirable. ## Description We are partnering with a leading organization to hire Skilled Infrastructure Platform Specialists to support their advanced IT infrastructure, HPC, and enterprise storage environments. This is an exciting opportunity for candidates with a strong technical background in server, storage, and high-performance computing environments. Key Responsibilities * Install, configure, and maintain physical and virtual servers, including Dell and Nvidia platforms. * Design, deploy, and manage enterprise storage (SAN/NAS) and backup environments, ensuring performance, scalability, and compliance. * Support HPC and AI compute infrastructures, GPU clusters, and model training platforms. * Monitor system health, performance metrics, and proactively resolve incidents to maintain uptime and meet SLA targets. * Collaborate with cross-functional teams to deliver seamless operations and support strategic projects. * Implement security best practices, patching, and disaster recovery processes. * Document procedures, configurations, and incident resolutions for compliance and knowledge sharing. * Lead troubleshooting efforts, root cause analysis, and operational improvements, providing guidance to junior engineers. Required Skills & Experience * Strong technical expertise across server, storage, networking, and HPC infrastructure. * Experience with Dell EMC, NetApp, HPE, Nvidia, and virtualization technologies. * Linux and Windows system administration experience. * Knowledge of backup, disaster recovery, and monitoring tools. * Scripting and automation skills (Python, PowerShell, Bash, Ansible, Terraform). * Experience with workload schedulers (e.g., SLURM, PBS) and AI/ML frameworks (TensorFlow, PyTorch) is a plus. * Excellent analytical, troubleshooting, and collaboration skills. Qualifications * Relevant certifications (ITIL, Dell EMC, NetApp, NVIDIA DLI, Veeam, or equivalent) are highly desirable. Additional Details * Seniority level: Mid-Senior level * Employment type: Full-time * Job function: Information Technology * Industries: IT Services and IT Consulting, IT System Operations and Maintenance * Location: London, England, United Kingdom ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Discover the open source trio you didn’t expect: .NET and PostgreSQL on Linux](https://www.wearedevelopers.com/videos/2042-discover-the-open-source-trio-you-didn-t-expect-net-and-postgresql-on-linux) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)