cloud orchestration engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+3 more
Job description
We’re looking for an cloud orchestration engineer to build the operational backbone that keeps vLLM running reliably at massive scale. You’ll design the systems for cluster management, deployment automation, and production monitoring that enable teams worldwide to serve AI models without friction. You’ll ensure that vLLM deployments are observable, debuggable, and recoverable, turning operational complexity into infrastructure that just works.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
- Strong experience with Kubernetes and container orchestration at scale.
- Experience designing and implementing custom Kubernetes operators.
- Proficiency in Python/Rust/Go and infrastructure-as-code tools (Terraform, Helm, etc).
- Experience managing GPU clusters and debugging hardware issues.
- Ability to work across cloud platforms (AWS, GCP, Azure) and on-premise infrastructure.
Preferred qualifications:
- Experience with ML-specific orchestration tools (Ray, Slurm).
- Knowledge of GPU scheduling, multi-tenancy, and resource optimization.
- Familiarity with vLLM deployment patterns and configuration.
- Track record of improving operational reliability for ML systems.
Bonus points if you have:
- Experience deploying inference systems on large-scale GPU (1,000+) clusters.
Benefits & conditions
- Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
- Visa sponsorship: We sponsor visas on a case-by-case basis.
- Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
About the company
Inferact’s mission is to grow vLLM as the world’s AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware-a position that took years to build.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
7 Cloud Computing Trends Coming in 2025 for Developers
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production