Software Engineer, Platform & Infrastructure

Riot Games
Los Angeles, CA, United States
9 days ago
Apply on www.riotgames.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$161,500.0 - $227,000.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Computer Clusters Continuous Integration Programming Tools Distributed Computing Environment Distributed Systems Python (Programming Language) Machine Learning Azure Machine Learning Software Engineering Computer Networking Systems
+8 more
Cloud Platform System Kubernetes Information Technology Deployment Automation Machine Learning Operations Software Version Control Data Pipelines Microservices

Job description

Platform and infrastructure engineers at Riot build the foundational systems that enable teams to develop, deploy, and operate systems at global scale. They partner across disciplines with other engineers, data scientists, designers, and product teams to ensure reliable, scalable, and secure infrastructure underpins every ML capability that reaches players., As a Senior Platform & Infrastructure Engineer on the Riot Technology team, you will design, build, and operate the core infrastructure and ML platforms. Your focus will be on the computing and orchestration platforms that power large-scale distributed training of agents (e.g. via RL, IL, and other techniques), simulation environments, and policy evaluation, as well as the CI/CD, infrastructure-as-code, observability, and developer tooling that keep these systems production-grade. You will close critical infrastructure gaps across the team’s stack, driving improvements to standards, automation, and operational maturity. You will operate independently on multi-month work efforts and begin to influence technical direction beyond your immediate team. You will report to the Manager of Machine Learning., * Build and operate Kubernetes, multi-node GPU clusters, and networking infrastructure for distributed ML bot training and large-scale policy evaluation.

  • Design infrastructure for running simulation environments at scale, enabling parallel rollouts, data collection, training, and evaluation.
  • Build CI/CD, deployment automation, artifact management, and infrastructure-as-code across cloud environments.
  • Improve platform reliability, cost efficiency, performance, reproducibility, auditability, and operational maturity.
  • Build observability, monitoring, alerting, health indicators, and SLO-aligned dashboards for infrastructure and ML workloads.
  • Develop internal APIs, control planes, templates, and developer tooling for distributed training and evaluation workflows.
  • Support MLOps workflows including automated training pipelines, model artifact management, experiment tracking, and reproducible ML lifecycle operations.
  • Build security and governance controls, manage production incidents, drive root-cause remediation, mentor engineers, and support recruiting for platform roles.

Requirements

  • Bachelor’s degree in Computer Science or a related field, or equivalent practical experience.
  • 3+ years of software engineering experience, with meaningful experience in infrastructure, platform engineering, or SRE roles.
  • Experience operating distributed systems in production and keeping them healthy under real load.
  • Strong experience with Kubernetes, AWS or GCP, infrastructure-as-code, CI/CD, deployment automation, and production tooling.
  • Experience with GPU compute infrastructure, including scheduling, multi-node orchestration, and resource optimization for long-running training workloads.
  • Proficiency in Python and solid understanding of networking, microservices, core infrastructure services, and distributed systems fundamentals., * Familiarity with MLOps workflows such as model versioning, pipeline orchestration, experiment tracking, artifact management, and reproducible ML workflows.
  • Experience with distributed training or HPC frameworks, inference serving, systems languages, high-performance networking, Unreal/client-server architecture, AI-assisted development tools
  • Passion for games and player experience.

For this role, you’ll find success through craft expertise, a collaborative spirit, and decision-making that prioritizes the delight of players. We will be looking at your past studies, experience, and your personal relationship with games. If you embody player empathy and care about players’ experiences, this could be your role!

About the company

Riot focuses on work/life balance, shown by our open paid time off policy and other perks such as flexible work schedules. We offer medical, dental, and life insurance, parental leave for you, your spouse/domestic partner, and children, and a 401k with company match. Check out our benefits pages for more information.

At Riot Games, we put players first. That mission drives every decision in our quest to create games and experiences that make it better to be a player. Whether you’re working directly on a new player-facing experience or you’re supporting the company as a whole, everyone at Riot is part of our mission. And just like in our games, we’re better when we work together. Our goal is to create collaborative teams where you are empowered to bring your unique perspective everyday. If that sounds like the kind of place you want to work, we’re looking forward to your application.

It’s our policy to provide equal employment opportunity for all applicants and members of Riot Games, Inc. Riot Games makes reasonable accommodations for handicapped and disabled Rioters and does not unlawfully discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, handicap, veteran status, marital status, criminal history, or any other category protected by applicable federal and state law. We consider for employment all qualified applicants, including those with criminal histories, in a manner consistent with applicable federal, state and local law, including the California Fair Chance Act, the City of Los Angeles Fair Chance Initiative for Hiring Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, the San Francisco Fair Chance Ordinance, and the Washington Fair Chance Act.

Per the Los Angeles County Fair Chance Ordinance, the following core duties may create a basis for disqualifying candidates with relevant criminal histories:

  • Safeguarding confidential and sensitive Company data
  • Communication with others, including Rioters and third parties such as vendors, and/or players, including minors
  • Accessing Company assets, secure digital systems, and networks
  • Ensuring a safe interactive environment for players and other Rioters

These duties are directly related to essential operations, safety, trust, and compliance obligations within our organization. Please note that job duties may evolve based on business needs and additional responsibilities may be assigned as necessary to maintain operational efficiency and security.

  • (Los Angeles Only) Base salary range between $161,500.00 - $227,000.00 USD + incentive compensation + equity + 401K with company match + medical, dental, vision, and life insurance + short and long-term disability + open PTO.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.riotgames.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

3:07 min

Transitioning architecture to microservices at Netflix

Steve Upton Steve Upton · World Congress 2022

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff · World Congress 2024

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · World Congress 2022

Videos

See all

Related articles

See all