> Markdown version of [/jobs/ext/2635928-machine-learning-engineer-graduate-aml-engine-orchestration-2027-start](https://www.wearedevelopers.com/jobs/ext/2635928-machine-learning-engineer-graduate-aml-engine-orchestration-2027-start). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer Graduate (AML-Engine-Orchestration) - 2027 Start - **Company:** BYTEDANCE INC. - **Location:** Seattle, WA, United States - **Experience:** Starter - **Salary:** $121,600.0 - $243,200.0 - **Contract:** Internship / Graduate position - **Skills:** Artificial Intelligence, C++ (Programming Language), Profiling, Computer Networks, Concurrent Computing, Data Structures, Linux, Disaster Recovery, Distributed Systems, Statistical Hypothesis Testing, Python (Programming Language), Machine Learning, Online Service Provider, Open Source Technology, Performance Tuning, Software Engineering, Autoscaling, Build Management, Kubernetes, Information Technology, Free and Open-Source Software, Machine Learning Operations, Decoding - **Published:** August 10, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=1a3bec1ecbd603ac ## About the Role * Individuals who are completing or have recently completed a Bachelor's or Master's degree in Computer Science, Software Engineering, Artificial Intelligence, or a related technical field. * Proficiency in at least one of Go, C++, or Python, with a solid foundation in data structures, algorithms, and software engineering principles. * Familiarity with Linux and a foundational understanding of operating systems, computer networks, concurrent programming, and distributed systems. * Strong hands-on and exploratory abilities, with a willingness to investigate systems through source code, metrics, logs, profiling, and experiments. * A systematic and quantitative approach to problem solving, with the ability to define measurements, test hypotheses, and validate system improvements. * Demonstrated ownership and collaboration through coursework, research, internships, open-source contributions, or other engineering projects., * Experience with Kubernetes, container runtimes, resource scheduling, quota management, multi-tenant systems, or FinOps. * Contributions to open-source infrastructure projects such as Kubernetes, Volcano, Koordinator, or OpenKruise. * Experience with model serving systems such as vLLM, SGLang, Triton, KServe, or Ray Serve, or an understanding of KV Cache, Continuous Batching, Prefill/Decode disaggregation, or model parallelism. * Experience with online services, gateways, traffic management, autoscaling, performance optimization, or highly available distributed systems. * Experience with GPU/NPU programming, heterogeneous resource scheduling, model distribution, or inference performance analysis., Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment ## Description The Data-AML-Engine Orchestration team builds large-scale machine learning infrastructure that powers online model serving across ByteDance products, including TikTok. We develop the orchestration, scheduling, and resource management systems that connect heterogeneous compute infrastructure with production ML workloads. You will work on systems that directly affect GPU utilization, serving latency and availability, infrastructure reliability, and MLE productivity. Depending on your background and interests, you may focus on one or more of the following areas., 1. Design and build foundational orchestration capabilities for machine learning platforms, including Kubernetes Operators, container runtimes, and lifecycle management for jobs, services, and stateful workloads. 2. Build multi-tenant resource and quota systems that support priorities, preemption, fair sharing, elasticity, and cross-cluster scheduling. Improve GPU utilization and cost efficiency through resource pooling and FinOps. 3. Build lifecycle orchestration for online model serving, including model and image distribution, deployment, upgrades, rollback, autoscaling, multi-cluster operation, and disaster recovery. 4. Build serving orchestration and traffic management capabilities for disaggregated serving clusters, including topology-aware scheduling, KV Cache affinity, intelligent request routing, and QoS/SLA management. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Challenges and Solutions for Efficient, Large-Scale Video Analysis](https://www.wearedevelopers.com/videos/2022-challenges-and-solutions-for-efficient-large-scale-video-analysis) - [Decode Your People: Using PCM to Build High-Performance Teams](https://www.wearedevelopers.com/videos/100194-decode-your-people-using-pcm-to-build-high-performance-teams) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)