HPC Systems Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+3 more
Job description
As a member of our Platform Development team, you will be instrumental in building and optimizing high-performance trading systems, research compute clusters, databases, support systems, and more. You will heavily utilize Linux and Windows internals while working on servers in our HPC environment.
What You’ll Do:
Automate and Evolve: Contribute to our library of home-grown tools, written primarily in Python and Bash, to automate monitoring, and maintenance, allowing you to focus on performance related issues and projects.
Collaborate: Work closely with Strategy Developers, Quantitative Researchers, and trade-supporting application teams to translate complex problems into scalable solutions. Coordinate with IT infrastructure teams, including storage and networking, to identify and implement the best solutions.
Optimize: Tune operating systems and batch workflows for performance. Dive deep on root-cause analysis of systems issues. Integrate all of these solutions into our systems effectively and efficiently.
Comprehensive HPC Environment Management: Oversee all aspects of our HPC environment, including the scheduler, parallel filesystems, GPUs, and interconnects.
High throughput storage: Implement and optimize high-performance storage solutions, including Lustre, VAST, and GPFS, to efficiently support and enhance cluster performance.
Capacity Planning and Design: Develop strategies to ensure optimal resource allocation and scalability, using analytics to forecast needs and design efficient, reliable systems.
Troubleshooting and Tuning: Utilize monitoring and diagnostic tools to quickly pinpoint failures, streamline troubleshooting processes, and ensure the timely recovery of disrupted workflows.
Requirements
- A Bachelor’s degree in Engineering, Computer Science, Information Systems, or a related discipline.
- 5-7 years of progressive experience building Linux and/or Windows based HPC based platforms.
- Familiarity with kernel-level and I/O subsystem tweaks and tools such as sysctl, strace, tcpdump, and netstat. Bonus points for equivalent Windows knowledge (registry, procmon, wireshark, tshark).
- Recent hands-on experience with automation in Python or other tools.
- Experience administering Lustre, GPFS, VAST, or other parallel filesystems.
- Understanding of resource schedulers like HTCondor, SLURM, or similar.
About the company
Susquehanna is a global quantitative trading firm powered by scientific rigor, curiosity, and innovation. Our culture is intellectually driven and highly collaborative, bringing together researchers, engineers, and traders to design and deploy impactful strategies in our systematic trading environment. To meet the unique challenges of global markets, Susquehanna applies machine learning and advanced quantitative research to vast datasets in order to uncover actionable insights and build effective strategies. By uniting deep market expertise with cutting-edge technology, we excel in solving complex problems and pushing boundaries together.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
7 Cloud Computing Trends Coming in 2025 for Developers
Data Science & more: The Lopez dilemma
9 Ways to Make Money Hacking
The Best X (Twitter) Accounts for Developers