> Markdown version of [/jobs/ext/3325427-hpc-operations-engineer-24-7-linux-compute-automation](https://www.wearedevelopers.com/jobs/ext/3325427-hpc-operations-engineer-24-7-linux-compute-automation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # HPC Operations Engineer - 24/7 Linux Compute & Automation - **Company:** Jump Trading - **Location:** Greater London, UK - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Software Documentation, Linux, General Parallel File Systems, Monitoring of Systems, Python (Programming Language), Linux System Administration, Remote Direct Memory Access, High Performance Computing, Bug Reporting, Performance Monitor, Slurm, Programming Languages - **Published:** September 5, 2026 - **Apply:** https://www.collegerecruiter.com/job/2859580437-hpc-operations-engineer--247-linux-compute--automation ## About the Role We are looking for an adaptable hands-on individual, passionate about the details and nuances of managing Linux HPC environments at scale, and eager to tackle complex and unpredictable operational work as their primary job function., * A desire for operational work as primary job function * At least 2+ years of professional experience with Linux systems administration * High performance computing (HPC), including parallel filesystems (e.g., Lustre, GPFS), batch systems (e.g., Slurm, Grid Engine), and high-performance network interconnects experience is a plus, but not required * High proficiency with at least one programming/scripting language (e.g., Go, Python, C) and ability to learn additional languages quickly * A compulsion to perform root cause analysis * Strong verbal and written communication skills, including the ability to communicate effectively and efficiently with both coworkers and third-party vendors * Strong collaboration skills with a willingness to undertake tasks of various technologies and complexities * Ability to independently manage complex projects and multiple workstreams * Strong sense of urgency * Willingness to perform regular operational maintenance work during evenings and weekends and as needed * Ability to work effectively in a busy, open floor plan office environment * Reliable and predictable availability ## Description * Provide front-line operational support for 24/7 Linux HPC compute, storage, and interconnects. Technologies involved include RDMA fabrics, parallel filesystems, HPC batch schedulers, FUSE filesystems, internal Jump software, multi-vendor hardware, cybersecurity requirements, a challenging and unpredictable client workload, and high user expectations * Solve problem reports and questions posed by members of Jump's research community, escalating as needed and managing the entire problem lifecycle * Respond to alerts in a timely fashion * Participate in large, coordinated maintenance operations, including during evenings and weekends * Work on global projects across a wide range of infrastructure * Write code for diagnosing, resolving, and triaging difficult problems and automating frequently performed tasks * Collaborate with team members and across teams to write code and testing infrastructures spanning both new and existing codebases in multiple programming languages * Manage relationships with outside vendors, including traveling both domestically and internationally to meet with current and potential vendors * Implement and support performance monitoring and fault monitoring systems * Develop and improve systems and user documentation * Develop and monitor the tools used to maintain a production computing environment * Provide operational support as primary job function * Adhere to all company cybersecurity and IT policies, including performing all work using only approved hardware and software * Participate in an on-call rotation * Other tasks as assigned or needed * Work from company office an average of 5 days a week ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [Best Companies to work for in London: Top 25 Companies in 2023](https://www.wearedevelopers.com/magazine/187-best-companies-to-work-for-in-london-top-25-companies-in-2023) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)