> Markdown version of [/jobs/ext/2547559-machine-learning-system-scheduling-engineer-graduate-applied-machine-learning-2027-start](https://www.wearedevelopers.com/jobs/ext/2547559-machine-learning-system-scheduling-engineer-graduate-applied-machine-learning-2027-start). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning System Scheduling Engineer Graduate (Applied Machine Learning) - 2027 Start - **Company:** BYTEDANCE INC. - **Location:** San Jose, CA, United States - **Experience:** Starter - **Salary:** $128,000.0 - $256,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Cloud Engineering, Cloud Storage, Nvidia CUDA, Computer Programming, Data Centers, Linux, Microprocessors, Distributed Data Store, Distributed Systems, Python (Programming Language), Machine Learning, Open Source Technology, Remote Direct Memory Access, Tensorflow, Mesos, AI Infrastructure, Graphics Processing Unit (GPU), High Performance Computing, Apache Yarn, Pytorch, Amazon Virtual Private Cloud (VPC), Containerization, Kubernetes, Information Technology, Machine Learning Operations, Docker, Programming Languages - **Published:** August 6, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=7f3d267960a01093 ## About the Role * Individuals who are completing or have recently completed a Bachelor's or Master's degree in Computer Science or a related discipline. * Proficient in one or two programming languages in a Linux environment, such as Go, Java, or Python. * Solid foundation in computer science and programming, familiarity with common algorithms and data structures, and good coding habits. * Familiar with at least one mainstream machine learning framework, such as TensorFlow, PyTorch, or an internally developed framework. * Familiar with Kubernetes architecture and ecosystem, as well as container technologies such as Docker, container, and Kata; * Understands distributed system principles and has participated in the design, development, or maintenance of large-scale distributed systems., * Practical experience in large-scale cluster online/offline resource scheduling; source-level understanding of one or more open-source schedulers such as Kubernetes, Volcano, YARN, or Mesos; familiarity with containerization and lightweight virtualization technologies. * Deep understanding and practical experience in scheduling topics such as multi-tenant quota governance, preemption, elasticity, fragmentation, tidal scheduling, co-location, and QoS; strong analytical and modeling ability for complex problems; GPU scheduling experience is preferred. * Experience in at least one of the following areas: CUDA, RDMA, AI infrastructure, hardware and software co-design, high-performance computing, machine learning hardware architecture such as GPUs, accelerators and networking, ML for systems, or distributed storage. * Hands-on experience in cloud-native machine learning systems is preferred., Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment ## Description * Design and develop resource scheduling systems for machine learning workloads, supporting Volcano Ark and machine learning platform products. * Optimize orchestration and scheduling of heterogeneous compute resources, including GPUs, CPUs, and other accelerators, as well as storage resources such as cloud storage and networking resources such as VPC and RDMA, across multiple data centers and clusters. * Support scheduling requirements for offline training, online inference, and other workloads under strict multi-tenant isolation, improving overall resource utilization and efficiency. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [6 Emerging Technologies We’ll Learn About in 2025](https://www.wearedevelopers.com/magazine/381-6-emerging-technologies-we-ll-learn-about-in-2025)