> Markdown version of [/jobs/ext/2719133-cloud-orchestration-engineer](https://www.wearedevelopers.com/jobs/ext/2719133-cloud-orchestration-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # cloud orchestration engineer - **Company:** VLLM LLC - **Location:** San Francisco, CA, United States - **Salary:** $200,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Computer Clusters, Python (Programming Language), Google Cloud, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Deployment Automation, Slurm, Machine Learning Operations, Hardware Infrastructure, Terraform, Hardware Debugging - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/member-of-technical-staff-cloud-orchestration-inferact-8115119 ## About the Role * Bachelor's degree or equivalent experience in computer science, engineering, or similar. * Strong experience with Kubernetes and container orchestration at scale. * Experience designing and implementing custom Kubernetes operators. * Proficiency in Python/Rust/Go and infrastructure-as-code tools (Terraform, Helm, etc). * Experience managing GPU clusters and debugging hardware issues. * Ability to work across cloud platforms (AWS, GCP, Azure) and on-premise infrastructure. Preferred qualifications: * Experience with ML-specific orchestration tools (Ray, Slurm). * Knowledge of GPU scheduling, multi-tenancy, and resource optimization. * Familiarity with vLLM deployment patterns and configuration. * Track record of improving operational reliability for ML systems. Bonus points if you have: * Experience deploying inference systems on large-scale GPU (1,000+) clusters. ## Description We're looking for an cloud orchestration engineer to build the operational backbone that keeps vLLM running reliably at massive scale. You'll design the systems for cluster management, deployment automation, and production monitoring that enable teams worldwide to serve AI models without friction. You'll ensure that vLLM deployments are observable, debuggable, and recoverable, turning operational complexity into infrastructure that just works. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)