ML Platform Engineer - GPU Infrastructure

Optimal Inc.
Warren, MI, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Systems Engineering Microsoft Azure Bash Shell Computer Engineering Continuous Integration Linux DevOps Distributed Systems Monitoring of Systems Python (Programming Language)
+12 more
Machine Learning Performance Tuning Prometheus Azure Machine Learning Scripting Grafana Software Troubleshooting Containerization Kubernetes Information Technology Hardware Infrastructure Docker

Job description

Support team by designing, implementing, and maintaining the automation and ML workload enablement layer of the GPU cluster platform. This role focuses on optimizing GPU compute environments for AI/ML training and Isaac Sim simulation workloads, integrating GPU jobs into CI/CD pipelines, standardizing runtime environments, and supporting reliable storage and artifact management., Support GPU cluster platforms for AI/ML and simulation workloads Optimize GPU compute environments for ML training and Isaac Sim execution Integrate GPU workload execution into CI/CD pipelines Standardize runtime environments using containers and automation tools Manage storage, artifacts, and workload outputs Troubleshoot and improve platform reliability, scalability, and performance Collaborate with ML, infrastructure, and engineering teams

Requirements

Do you have experience in Tooling?, Do you have a Master’s degree?, 3+ years of experience in ML Platform Engineering, DevOps, Infrastructure Engineering, or related field Bachelor’s or Master’s degree in Systems Engineering, Computer Science, Computer Engineering, or related discipline, Experience with Linux, Kubernetes, Docker, and GPU infrastructure Knowledge of CI/CD tools and automation scripting (Python/Bash) Experience supporting AI/ML workloads and distributed systems Familiarity with NVIDIA GPU technologies and containerized environments Strong troubleshooting and performance optimization skills

Preferred Skills Experience with Isaac Sim or simulation workloads Exposure to cloud platforms (AWS, Azure, or GCP) Knowledge of monitoring and observability tools such as Grafana or Prometheus

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all