Software Engineer - ML Platform
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+2 more
Job description
As an ML Platform Engineer at Avride, you’ll own critical pieces of the ML stack: workflow orchestration, distributed execution, resource governance, performance.You will shape how ML teams across the company run experiments and train models at scale. You will build the abstractions and services that make training workloads reliable, cost-efficient, and fast, helping ML teams run at scale on Kubernetes with strong reliability and excellent developer experience., * Build and scale our ML compute platform on Kubernetes, using Argo Workflows for training, evaluation, and data processing orchestration
- Design and implement core platform capabilities, including a Ray-based internal SDK for distributed execution, and multi-tenant resource governance - scheduling, priorities, quotas, and policy enforcement across GPU, CPU, memory, and IO
- Improve end-to-end training throughput and platform efficiency by optimizing data access patterns, caching, and removing bottlenecks in storage, network, and resource contention
- Work directly with ML teams to debug complex workload issues, drive root-cause analysis, and turn recurring problems into platform-level fixes
- Evaluate, integrate and extend open-source tooling (Argo Workflows, Ray, Kubernetes ecosystem) to meet evolving platform needs
Requirements
- Strong proficiency in Python or Go; C++ is a plus
- Track record of designing and building scalable, maintainable systems and services
- Experience operating production services end-to-end: APIs, reliability practices, observability
- Deep knowledge of Kubernetes: how scheduling, resource management, controllers, and pod lifecycle actually behave under pressure
- Solid Linux and systems debugging skills: performance investigation, networking, storage/IO
- Ability to troubleshoot complex production issues across logs, metrics, and traces and drive them to resolution
Nice to have
- Experience with Argo Workflows, Ray, MLflow, or comparable distributed ML tooling
- Hands-on experience building or operating large-scale ML training systems: GPU scheduling, distributed training, training data pipelines
- Track record of optimizing resource usage and performance in distributed environments
LI-MS1
Candidates are required to be authorized to work in the U.S. The employer is not offering relocation, sponsorship, and remote work options are not available.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Fully Remote Software Engineer Jobs
Find a Developer Job: 12 Best Job Sites For Developers
Is Software Engineering Over-Saturated?