ML Infrastructure Engineer

Finoit Inc
Redwood City, CA, United States
about 1 month ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Computer Clusters Distributed Computing Environment Distributed Systems Memory Management Job Scheduling Machine Learning Performance Tuning Google Cloud Pytorch Containerization
+5 more
Kubernetes Slurm Machine Learning Operations TensorRT Data Pipelines

Job description

  • Design and scale distributed ML training infrastructure for large GPU clusters.
  • Build and optimize training pipelines using PyTorch, DeepSpeed, and distributed training frameworks.
  • Develop and maintain job scheduling systems using Kubernetes and/or SLURM.
  • Create high-throughput data pipelines for large-scale multimodal datasets.
  • Optimize GPU utilization, memory efficiency, and overall system performance.
  • Build low-latency inference pipelines for production ML deployments.

Requirements

  • 7+ years of experience in ML Infrastructure, HPC, or Distributed Systems.
  • Strong experience with PyTorch, DeepSpeed, FSDP, ZeRO, or similar distributed training frameworks.
  • Hands-on experience with Kubernetes, cloud platforms (AWS/Google Cloud Platform), and containerized environments.
  • Strong understanding of distributed systems, GPU optimization, NCCL, memory management, and performance tuning.
  • Experience building scalable ML infrastructure from development through production., * Experience with multimodal AI, robotics data pipelines, Triton, TensorRT, custom ML kernels, or ML compiler/runtime optimization.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · World Congress 2025

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · World Congress 2024

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all