HPC Systems Engineer

Susquehanna International Group, LLP
United States
19 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Microsoft Windows Algorithmic Trading Bash Shell Information Systems Databases Linux General Parallel File Systems Python (Programming Language) Performance Tuning Tcpdump Wireshark Diagnostic Tools
+3 more
Graphics Processing Unit (GPU) Information Technology Slurm

Job description

As a member of our Platform Development team, you will be instrumental in building and optimizing high-performance trading systems, research compute clusters, databases, support systems, and more. You will heavily utilize Linux and Windows internals while working on servers in our HPC environment.

What You’ll Do:

Automate and Evolve: Contribute to our library of home-grown tools, written primarily in Python and Bash, to automate monitoring, and maintenance, allowing you to focus on performance related issues and projects.

Collaborate: Work closely with Strategy Developers, Quantitative Researchers, and trade-supporting application teams to translate complex problems into scalable solutions. Coordinate with IT infrastructure teams, including storage and networking, to identify and implement the best solutions.

Optimize: Tune operating systems and batch workflows for performance. Dive deep on root-cause analysis of systems issues. Integrate all of these solutions into our systems effectively and efficiently.

Comprehensive HPC Environment Management: Oversee all aspects of our HPC environment, including the scheduler, parallel filesystems, GPUs, and interconnects.

High throughput storage: Implement and optimize high-performance storage solutions, including Lustre, VAST, and GPFS, to efficiently support and enhance cluster performance.

Capacity Planning and Design: Develop strategies to ensure optimal resource allocation and scalability, using analytics to forecast needs and design efficient, reliable systems.

Troubleshooting and Tuning: Utilize monitoring and diagnostic tools to quickly pinpoint failures, streamline troubleshooting processes, and ensure the timely recovery of disrupted workflows.

Requirements

  • A Bachelor’s degree in Engineering, Computer Science, Information Systems, or a related discipline.
  • 5-7 years of progressive experience building Linux and/or Windows based HPC based platforms.
  • Familiarity with kernel-level and I/O subsystem tweaks and tools such as sysctl, strace, tcpdump, and netstat. Bonus points for equivalent Windows knowledge (registry, procmon, wireshark, tshark).
  • Recent hands-on experience with automation in Python or other tools.
  • Experience administering Lustre, GPFS, VAST, or other parallel filesystems.
  • Understanding of resource schedulers like HTCondor, SLURM, or similar.

About the company

Susquehanna is a global quantitative trading firm powered by scientific rigor, curiosity, and innovation. Our culture is intellectually driven and highly collaborative, bringing together researchers, engineers, and traders to design and deploy impactful strategies in our systematic trading environment. To meet the unique challenges of global markets, Susquehanna applies machine learning and advanced quantitative research to vast datasets in order to uncover actionable insights and build effective strategies. By uniting deep market expertise with cutting-edge technology, we excel in solving complex problems and pushing boundaries together.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Introduction to eBPF as a secure virtual machine

Ayesha Kaleem · WWC 2023

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · WWC Europe 2026

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:04 min

Insights on transitioning from supercomputing to technical education

Andrew Holway · LIVE

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · WWC Europe 2026

Videos

See all

Related articles

See all