ML Training Infrastructure Engineer

Dyna Robotics
Redwood City, United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Computer Clusters Compilers Program Optimization Profiling Distributed Computing Environment Distributed Systems Memory Management Protocol Buffers Job Scheduling Data Processing Graphics Processing Unit (GPU)
+5 more
Pytorch Kubernetes Slurm Machine Learning Operations TensorRT

Job description

As a ML Training Infrastructure Engineer, you will architect and build the systems that turn our multi-cloud GPU fleet into a training engine our researchers love. Your charter is singular and broad: own training infrastructure end-to-end so that every GPU is busy, every run is reproducible, and every researcher’s next experiment is one command away., * Scale Distributed Training: Architect and own the infrastructure for large-scale GPU clusters. You’ll implement sharding, activation checkpointing, and memory optimization (ZeRO, FSDP) to enable the training of massive multimodal models.

  • Optimize Researcher Ergonomics: Build a research codebase and job scheduling system (Kubernetes/SLURM) that prioritizes fast iteration, automated retries, and seamless failure recovery.
  • High-Performance Data Handling: Design high-throughput pipelines to ingest and transform terabytes of multimodal robot data (video, proprioception, 3D signals), ensuring dataloaders never starve the GPUs.
  • Production Inference: Build low-latency inference pipelines for real-time robot control. You’ll apply quantization, distillation, and model compilation (TensorRT, Triton) to move models from the lab to the physical world.
  • Deep Systems Profiling: Dive into the weeds of GPU utilization, I/O bottlenecks, and memory fragmentation to squeeze every bit of performance out of our expanding compute fleet.

Requirements

  • 7+ Years of Engineering: With a track record of leading technical projects in high-performance computing (HPC) or ML infrastructure.
  • ML Systems Mastery: Deep experience with PyTorch and distributed training frameworks (DeepSpeed, Accelerate). You understand the nuances of mixed precision and gradient accumulation.
  • Infrastructure Expertise: Hands-on experience managing cloud GPU environments (GCP/AWS) and container orchestration (Kubernetes).
  • Low-Level Intuition: A fundamental understanding of distributed systems, including race conditions, memory management, and NCCL/inter-node communication.
  • Ownership Mindset: You don’t just “deploy” code; you design, build, and operate systems end-to-end to unblock fast-moving research.

Bonus Points For

  • Experience with Robotics Data Formats (MCAP, Protobuf) or multimodal models (VLAs).
  • Deep ML systems experience: custom kernels (Triton), compilers, or runtime optimization.
  • Experience as a founding or early-stage infrastructure hire.

About the company

Join us to shape the next frontier of AI-driven robotics!

Dyna Robotics makes general-purpose robots powered by a proprietary embodied AI foundation model that generalizes and self-improves across varied environments with commercial-grade performance. Dyna’s robots have been deployed at customers across multiple industries. Its frontier model has the top generalization and performance in the industry.

Dyna Robotics was founded by repeat founders Lindon Gao and York Yang, who sold Caper AI for $350 million, and former DeepMind research scientist Jason Ma. The company has raised over $140M, backed by top investors, including CRV and First Round. We’re positioned to redefine the landscape of robotic automation., At Dyna Robotics, we build technology for the real world, which requires a team as diverse as the environments our robots inhabit. We are an equal opportunity employer committed to technical rigor and mutual respect.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · World Congress 2025

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · World Congress 2024

Videos

See all

Related articles

See all