HPC Support Engineer

Hays plc
London, UK
2 days ago
Apply on www.hays.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Bash Shell Computer Clusters Compilers Nvidia CUDA Continuous Integration Linux File Systems Monitoring of Systems Python (Programming Language) OpenCL Package Management Systems
+12 more
Ansible High Performance Computing Grafana Parallel Computation HybridCloud Git Kubernetes Infrastructure Automation Frameworks Slurm Splunk Docker Elk Stack

Job description

Your new company

You will join a prestigious, research-led higher education institution in London, supporting academics and researchers undertaking complex computational research and AI-driven projects. The organisation is continuing to invest in its high-performance computing capabilities, providing an opportunity to work with advanced research infrastructure and emerging technologies.

Your new role

As a Research HPC Support Engineer, you will support, maintain and enhance the organisation’s high-performance computing environment. This includes HPC and GPU clusters, research storage, high-speed networking and hybrid cloud infrastructure.

You will administer and optimise Linux-based clusters, manage scientific software stacks and help researchers make effective use of HPC resources. You will also improve automation, monitoring and deployment processes using technologies such as Python, Bash, Ansible, Docker, Kubernetes, Grafana and CI/CD tooling.

Working closely with researchers and technical teams, you will:

  • Administer and support HPC clusters, GPU systems and specialist research hardware
  • Manage workload scheduling through technologies such as Slurm or OpenPBS
  • Maintain scientific software, compilers and libraries using tools such as Spack or EasyBuild
  • Automate the lifecycle of compute and GPU nodes
  • Troubleshoot complex Linux, storage, networking and performance issues
  • Support containerised and cloud-based research workloads
  • Help researchers optimise applications for parallel computing and GPU acceleration
  • Deliver technical guidance, documentation and training

What you’ll need to succeed

You will need strong hands-on experience supporting high-performance computing environments, alongside:

  • Advanced Linux cluster administration experience
  • Experience with HPC schedulers such as Slurm or OpenPBS
  • Knowledge of parallel computing and GPU technologies such as CUDA or OpenCL
  • Strong scripting and automation skills using Python, Bash or Ansible
  • Experience with Docker, Kubernetes, Git and CI/CD practices
  • Knowledge of monitoring and observability tools such as Grafana, ELK Stack or Splunk
  • Experience using HPC software build frameworks such as Spack or EasyBuild
  • Knowledge of scientific software stacks, compilers and libraries
  • Experience managing large-scale storage and file systems
  • The ability to work directly with researchers and translate technical requirements into practical solutions

Experience with cloud-based HPC, infrastructure-as-code tools, research computing environments or higher education would be advantageous.

What you’ll get in return

This is an opportunity to work within a highly regarded research environment where your expertise will directly support innovative computational research and AI applications.

You will work with modern HPC and GPU technologies, contribute to the future design of research computing services and collaborate with researchers tackling complex, data-intensive challenges. The position also provides scope to explore emerging technologies and influence continuous improvements across the HPC environment.

What you need to do now

If you’re interested in this role, click ‘apply now’ to forward an up-to-date copy of your CV, or call us now.

If this job isn’t quite right for you, but you are looking for a new position, please contact us for a confidential discussion about your career. #4835409 - Charlie

Requirements

  • Advanced Linux cluster administration experience
  • Experience with HPC schedulers such as Slurm or OpenPBS
  • Knowledge of parallel computing and GPU technologies such as CUDA or OpenCL
  • Strong scripting and automation skills using Python, Bash or Ansible
  • Experience with Docker, Kubernetes, Git and CI/CD practices
  • Knowledge of monitoring and observability tools such as Grafana, ELK Stack or Splunk
  • Experience using HPC software build frameworks such as Spack or EasyBuild
  • Knowledge of scientific software stacks, compilers and libraries
  • Experience managing large-scale storage and file systems
  • The ability to work directly with researchers and translate technical requirements into practical solutions

Experience with cloud-based HPC, infrastructure-as-code tools, research computing environments or higher education would be advantageous.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.hays.co.uk
Prepare application

Good distractions

Loading talks and stories from around this role…