> Markdown version of [/jobs/ext/3186087-senior-software-engineer-compute-infrastructure-orchestration-scheduling](https://www.wearedevelopers.com/jobs/ext/3186087-senior-software-engineer-compute-infrastructure-orchestration-scheduling). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer - Compute Infrastructure (Orchestration & Scheduling) - **Company:** BYTEDANCE INC. - **Location:** Seattle, WA, United States - **Experience:** Expert - **Salary:** $207,480.0 - **Contract:** Permanent contract - **Skills:** Clean Code Principles, Java (Programming Language), Artificial Intelligence, Amazon Web Services, Microsoft Azure, C++ (Programming Language), Computer Engineering, Data Centers, Job Scheduling, Python (Programming Language), Online Service Provider, Open Source Technology, Mesos, Azure Machine Learning, Software Systems, Rust (Programming Language), Computer Networking Systems, Google Cloud, Apache Yarn, Large Language Models, AI Platforms, Kubernetes, Information Technology, Machine Learning Operations, Serverless Computing, Docker, Golang, Programming Languages - **Published:** September 1, 2026 - **Apply:** https://www.gamesjobsdirect.com/job/bytedance/senior-software-engineer--compute-infrastructure-orchestration-scheduling/353965 ## About the Role Minimum Qualifications - B.S./M.S, degree in Computer Science, Computer Engineering or a related area with 3+ years of relevant industry experience; Ph.D. degree and strong publication records can be an exception. - Solid understanding of at least one of the following fields: Unix/Linux environments, distributed and parallel systems, high-performance networking systems, developing large scale software systems - Proven experience designing, architecting and building cloud and ML infrastructure related but not limited to resource management, allocation, job scheduling and monitoring. - Familiarity with container and orchestration technologies such as Docker and Kubernetes. - Proficiency in at least one major programming language such as Python, Go, C++, Rust, and Java. Preferred Qualifications - Experience in one large scale cluster management systems, e.g., Kubernetes, Ray, Yarn, or Mesos - Experience in large scale resource efficiency management and job scheduling development - Project experience in application scaling, workload co-location, and isolation enhancement - Experience with a public cloud provider (AWS, Azure and GCP), and their ML services (e.g., AWS SageMaker, Azure ML, GCP Vertex AI). - Great communication skills and the ability to work well within a team and across engineering teams. - Passionate about system efficiency, quality, performance and scalability ## Description About the Team The Compute Infrastructure - Orchestration & Scheduling team uses Kubernetes and Serverless technologies to build a large, reliable, and efficient compute infrastructure. This infrastructure powers hundreds of large-scale clusters globally, with over millions of online containers and offline jobs daily, including AI and LLM workloads. The team is dedicated to building cutting-edge, industry-leading infrastructure that empowers AI innovation, ensuring high performance, scalability, and reliability to support the most demanding AI/LLM workloads.The team is also dedicated to open-sourcing key infrastructure technologies, including projects in the K8s portfolio such as kubewharf, Serverless initiatives like Ray on K8s, and LLM inference control plan project AiBrix. At ByteDance, as we expand and innovate, powering global platforms like TikTok and various AI/ML & LLM initiatives, we face the challenge of enhancing resource cost efficiency on a massive scale within our rapidly growing compute infrastructure. We're seeking talented software engineers excited to optimize our infrastructure for AI & LLM models. Your expertise can drive solutions to better utilize computing resources (including CPU, GPU, power, etc.), directly impacting the performance of all our AI services and helping us build the future of computing infrastructure. Also, with the goal of growing compute infrastructure in overseas regions, including North America, Europe, and Asia Pacific, you will have the opportunities of working closely with leaders from ByteDance's global business units to ensure that we continue to scale and optimize our infrastructure globally. Responsibilities - Engineer hyper-scale cluster management: Enhance Kubernetes-based cluster platforms to deliver exceptional performance, scalability, and resilience-powering resource management across ByteDance's massive global infrastructure. - Innovate on core scheduling capabilities: Design and maintain a truly unified scheduling that powers diverse workloads (Containers & VMs, online services, offline computing, AI/ML, CPU/GPU workloads, etc) in a massive-scale resource pool. - Develop an intelligent scheduling system: Leverage AI models to optimize workload performance and resource utilization across heterogeneous resources-including CPU, GPU, memory, network, and power across global data centers. - Lead Infrastructure for Next-Gen ML Workloads: Design and drive the evolution of compute platforms purpose-built for fast, reliable, and cost-effective ML and LLM training/inference. - Deliver Quality and Innovation: Write high-quality, maintainable code, and stay at the forefront of open-source and research advancements in AI, ML, systems, and Serverless technologies. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)