Production System Engineer Graduate (Server Management) - 2027 Start

BYTEDANCE INC.
San Jose, CA, United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Internship / Graduate position
Employment type
Full-time (> 32 hours)
Experience level
Starter
Compensation
$76,000.0 - $128,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Data Analysis Systems Engineering Automation of Tests Bash Shell Ubuntu (Operating System) Nvidia CUDA Computer Maintenance Computer Programming Computer Engineering Data Centers Dynamic Host Configuration Protocol
+30 more
Debian Linux Linux DevOps Domain Name System (DNS) Monitoring of Systems Python (Programming Language) Knowledge Management Automation of Marketing Networking Basics Routing Open Source Technology Reliability Engineering Ansible Server Administration SQL Databases TCP/IP Virtual Local Area Networks AI Infrastructure Data Logging Scripting High Performance Computing Git Kubernetes Infrastructure Automation Frameworks Information Technology Hardware Infrastructure Restful APIs Docker Golang Programming Languages

Job description

The Server Management DevOps team is responsible for the end-to-end lifecycle management of servers across our self-built data centers in the United States and Europe.

Our scope covers the complete server lifecycle, including new hardware introduction, data center delivery, production operations, hardware maintenance, configuration changes, capacity migration, asset decommissioning, data sanitization, and hardware reuse.

The team serves as a central coordination point between multiple functions, including:

  • Hardware New Product Introduction (NPI)
  • Server and data center operations
  • Field maintenance and infrastructure management
  • Hardware vendors and service providers
  • Supply chain and asset management
  • Infrastructure platform and automation engineering teams

Our goal is to ensure that server infrastructure operates reliably, efficiently, and compliantly at scale throughout its entire lifecycle., As a Production Systems Engineer, you will work alongside experienced infrastructure engineers on real production challenges involving large-scale server fleets, GPU infrastructure, automation platforms, and AI-assisted operational tools. You will contribute to building automation, improving operational efficiency, troubleshooting production issues, and supporting the lifecycle management of servers deployed across ByteDance’s global data centers., * Server Infrastructure Operations: Assist with the deployment, validation, monitoring, maintenance, and lifecycle management of large-scale server fleets, including CPU and GPU servers.

  • Automation Development: Develop scripts, tools, and automation solutions using Python, Bash, Go, or other programming languages to reduce manual operational work and improve infrastructure efficiency.
  • Linux Systems: Work with Linux-based production environments and help troubleshoot operating system, hardware, storage, networking, and performance-related issues.
  • GPU and AI Infrastructure: Gain exposure to modern AI infrastructure and GPU server platforms, and contribute to operational tooling, validation, monitoring, or reliability improvements.
  • Monitoring and Data Analysis: Analyze server health, hardware failures, operational metrics, and infrastructure data to identify trends, risks, and opportunities for improvement.
  • AI for Infrastructure Operations: Explore opportunities to apply AI and large language models to infrastructure troubleshooting, automation, knowledge management, and operational decision-making.
  • Strong analytical and troubleshooting skills with the ability to learn unfamiliar technologies quickly.
  • Good communication skills and the ability to collaborate effectively in cross-functional engineering teams.

Requirements

We are looking for a motivated Production Systems Engineer who is passionate about Linux systems, server hardware, automation, AI infrastructure, and large-scale data center operations., * Individuals who are completing or have recently completed a Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, Information Technology, or a related technical field.

  • Experience in systems engineering, infrastructure operations, DevOps, Site Reliability Engineering, or related technical roles, or equivalent hands-on project experience.
  • Solid understanding of Linux system (Debian or Ubuntu preferred) administration and troubleshooting.
  • Programming or scripting experience in Python, Bash, Go, or another modern programming language.
  • Understanding of operating systems, computer architecture, networking fundamentals, and storage systems.

Preferred Qualifications

  • Experience developing automation tools or infrastructure software using Python, Bash, Go, or similar languages.
  • Working with server hardware, PC building, homelabs, or data center infrastructure.
  • Experience working with NVIDIA GPU platforms, AI infrastructure, CUDA, or high-performance computing environments.
  • Applying networking fundamentals, including TCP/IP, DNS, DHCP, VLANs, and routing.
  • Hands-on experience through internships, research, open-source projects, homelabs, technical competitions, or personal engineering projects.
  • Experience with infrastructure monitoring, observability, logging, or telemetry platforms.
  • Experience with Git, Docker, Kubernetes, REST APIs, SQL, or infrastructure automation frameworks such as Ansible., Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment

Benefits & conditions

Pulled from the full job description Paid parental leave Parental leave Health insurance 401(k) matching Vision insurance Dental insurance Paid sick time, Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · World Congress 2022

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all