Machine Learning System Scheduling Engineer Graduate (Applied Machine Learning) - 2027 Start

BYTEDANCE INC.
San Jose, CA, United States
28 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Starter
Experience required
1 year minimum
Compensation
$128,000.0 - $256,000.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Cloud Engineering Cloud Storage Nvidia CUDA Computer Programming Data Centers Linux Microprocessors Distributed Data Store Distributed Systems Python (Programming Language) Machine Learning
+16 more
Open Source Technology Remote Direct Memory Access Tensorflow Mesos AI Infrastructure Graphics Processing Unit (GPU) High Performance Computing Apache Yarn Pytorch Amazon Virtual Private Cloud (VPC) Containerization Kubernetes Information Technology Machine Learning Operations Docker Programming Languages

Job description

  • Design and develop resource scheduling systems for machine learning workloads, supporting Volcano Ark and machine learning platform products.
  • Optimize orchestration and scheduling of heterogeneous compute resources, including GPUs, CPUs, and other accelerators, as well as storage resources such as cloud storage and networking resources such as VPC and RDMA, across multiple data centers and clusters.
  • Support scheduling requirements for offline training, online inference, and other workloads under strict multi-tenant isolation, improving overall resource utilization and efficiency.

Requirements

  • Individuals who are completing or have recently completed a Bachelor’s or Master’s degree in Computer Science or a related discipline.
  • Proficient in one or two programming languages in a Linux environment, such as Go, Java, or Python.
  • Solid foundation in computer science and programming, familiarity with common algorithms and data structures, and good coding habits.
  • Familiar with at least one mainstream machine learning framework, such as TensorFlow, PyTorch, or an internally developed framework.
  • Familiar with Kubernetes architecture and ecosystem, as well as container technologies such as Docker, container, and Kata;
  • Understands distributed system principles and has participated in the design, development, or maintenance of large-scale distributed systems., * Practical experience in large-scale cluster online/offline resource scheduling; source-level understanding of one or more open-source schedulers such as Kubernetes, Volcano, YARN, or Mesos; familiarity with containerization and lightweight virtualization technologies.
  • Deep understanding and practical experience in scheduling topics such as multi-tenant quota governance, preemption, elasticity, fragmentation, tidal scheduling, co-location, and QoS; strong analytical and modeling ability for complex problems; GPU scheduling experience is preferred.
  • Experience in at least one of the following areas: CUDA, RDMA, AI infrastructure, hardware and software co-design, high-performance computing, machine learning hardware architecture such as GPUs, accelerators and networking, ML for systems, or distributed storage.
  • Hands-on experience in cloud-native machine learning systems is preferred., Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment

Benefits & conditions

Pulled from the full job description

  • Paid parental leave
  • Parental leave
  • Health insurance
  • 401(k) matching
  • Vision insurance
  • Dental insurance
  • Paid sick time, Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

About the company

Volcano Ark is an all-in-one large model service platform launched by Volcano Engine. It is a leading platform in China’s large model market by product capability and market share. The platform provides end-to-end services including model inference, evaluation, fine-tuning, AI application development, and a plugin ecosystem. Volcano Ark hosts Doubao and leading industry large models, and supports enterprise AI adoption through stable, secure, and trusted solutions as well as professional algorithm and technical services.

Data AML is ByteDance’s machine learning platform team. It provides training and inference systems for recommendation, advertising, computer vision, speech, and NLP scenarios across products such as Douyin, Toutiao, and Xigua Video. The team also supports internal business teams with large-scale machine learning compute, explores general and innovative algorithms for business problems, and offers core machine learning and recommendation system capabilities to external enterprise customers through Volcano Engine.

We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth., Founded in 2012, ByteDance’s mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content., Inspiring creativity is at the core of ByteDance’s mission. Our innovative products are built to help people authentically express themselves, discover and connect - and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an “Always Day 1” mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us., ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:29 min

Evaluating orchestration tools for distributed container deployments

Aleksandr Kalikov · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

5:28 min

Navigating machine learning operations and maturity level frameworks

Julian Joseph · LIVE

Videos

See all

Related articles

See all