Software Engineer, HPC Scheduling

NorthMark Strategies
Dallas, United States
27 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Computing Platforms Software Quality Directed Acyclic Graph (Directed Graphs) Data Structures Software Debugging Distributed Systems Event-Driven Programming Job Scheduling PostgreSQL Linux System Administration Software Architecture
+13 more
Queueing Systems Prometheus Software Engineering Software Systems Data Logging High Performance Computing Grafana Procedural Programming Kubernetes Apache Kafka Non-relational Database Slurm Golang

Job description

  • Designing and developing high-quality software solutions using procedural programming languages, with a focus on Golang
  • Building and maintaining highly scalable, highly available and globally distributed systems to support large-scale research workloads
  • Managing and optimising data interactions across relational and non-relational databases, particularly PostgreSQL
  • Developing and operating containerised applications within Kubernetes, ensuring effective orchestration and workload scheduling
  • Supporting, tuning and troubleshooting Linux-based systems as part of our core compute platform
  • Applying core networking knowledge to help debug, optimise and enhance platform connectivity and performance
  • Independently diagnosing and resolving complex technical issues across infrastructure and software layers
  • Applying solid software architecture principles, computer science fundamentals and data structure knowledge to guide design decisions and code quality
  • Driving continuous improvement by contributing to CI/CD pipelines and engineering best practices
  • Staying up to date with emerging technologies and approaches, and applying new knowledge across disciplines

Requirements

The HPC Scheduling team develops and manages a large high-performance compute (HPC) platform to enable the business to conduct complex research at scale. We are seeking a highly motivated person to join our team to help us continue to push the envelope running batch workloads on Kubernetes.

The ideal candidate will have an active interest in Kubernetes and batch computing, a broad range of experience with software engineering and development, as well as experience managing large-scale infrastructure and complex tooling environments., * Experience with developing Kubernetes components, such as controllers and operators

  • Experience with event-driven programming and message queues, such as apache Kafka and Pulsar
  • Experience of high-performance computing, Kubernetes, or DAG (Directed Acyclic Graph) workflows
  • Experience of running systems at scale using a cloud provider, ideally AWS
  • Use of operational and runtime tools and practices, including monitoring and logging with systems such as Prometheus and Grafana
  • Experience of operating or using job scheduling systems, such as SLURM

It is impossible to list every requirement for, or responsibility of, any position. Similarly, we cannot identify all the skills a position may require since job responsibilities and the Company’s needs may change over time. Therefore, the above job description is not comprehensive or exhaustive. The Company reserves the right to adjust, add to or eliminate any aspect of the above description. The Company also retains the right to require all employees to undertake additional or different job responsibilities when necessary to meet business needs.

Must be legally authorized to work in the United States without the need for employer sponsorship, now or at any time in the future.

Benefits & conditions

  • Company-Paid Benefits: 100% Employer-Paid Medical in our High Deductible Health Plan, Dental and Vision benefits for employees and their families, 16 weeks of Paid Parental Leave, Employee Assistance Program, Life insurance, Short-Term Disability and Long-Term Disability
  • 401(k): Company will match 100% of your contributions up to 6%
  • Optional Employee-Paid Benefits: Medical insurance in our PPO plan and a variety of other benefits such as Health Savings Accounts (with Company Contribution!), Flexible Spending Accounts, Supplemental Life Insurance, Wellhub and more.
  • Time Off: 25 days of Paid Time Off plus 12 company holidays

About the company

NorthMark Compute & Cloud (NMC ) is backed by dedicated leadership and investment, with a clear mission as it operates at the bleeding edge of technology. Its goal is to scale and enhance the high-performance computing (HPC) and cloud infrastructure that supports its clients’ research, production, and delivery, enabling breakthroughs that shape the industries of tomorrow. Its engineers build critical infrastructure to eliminate friction in scientific research, simulations, analysis, and decision-making, accelerating discovery and driving faster innovation.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

Videos

See all

Related articles

See all