Machine Learning Engineer

G P U Nuclear, Inc
United States
8 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Build Automation Computer Engineering Software Debugging Linux Distributed Computing Environment Distributed Systems Python (Programming Language) Machine Learning Performance Tuning Kubernetes Slurm Machine Learning Operations

Job description

We’re looking for a Senior Machine Learning Engineer to join our team during an exciting phase of growth. In this role, you’ll be responsible for building and operating the core systems that power large-scale ML training and inference across TensorWave’s GPU platform, working closely with cross-functional partners to support business objectives while upholding our standards for excellence, collaboration, and impact.

What You’ll Do

  • Design, operate, and improve ML infrastructure systems supporting distributed training and inference workloads
  • Build reliable, repeatable workload execution and orchestration patterns across shared GPU environments
  • Troubleshoot performance, reliability, and scalability issues across the ML stack
  • Partner with ML, systems, and platform teams to improve developer experience and operational efficiency

Requirements

  • Bachelor of Science in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience
  • Expertise supporting production ML systems using SLURM and Kubernetes
  • Strong understanding of GPU-accelerated workloads and distributed systems concepts
  • Solid Linux fundamentals and experience debugging infrastructure-level issues
  • Ability to build automation and tooling - Python, Go, etc., * Experience working across schedulers, orchestration platforms, or cluster managers
  • Familiarity with large-scale GPU environments or HPC-style systems
  • Experience improving infrastructure reliability, utilization, or performance at scale

Benefits & conditions

  • 100% paid Medical, Dental, and Vision insurance for Employees

  • Company Health Savings Account Contributions

  • 100% paid Short Term and Long Term Disability Insurance for Employees

  • Life and Voluntary Supplemental Insurance Options

  • Other Insurance Options, such as Pet & Legal Insurance

  • Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support

  • Flexible Spending Account

  • 401(k)

  • Employee Assistance Program

  • Flexible PTO

  • Paid Holidays

  • Parental Leave

  • Other In-Office Perks

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · World Congress 2022

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

Videos

See all

Related articles

See all