DevOps Engineer - ML & Data Infrastructure

High 5 Games, LLC
United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Batch Processing BigTable BigQuery Cloud Computing Code Review Data Governance Data Infrastructure Data Systems DevOps Data Flow Control Fraud Prevention and Detection
+21 more
Groovy Monitoring of Systems Python (Programming Language) Machine Learning Scrum Methodology Ansible Management of Software Versions Datadog Data Logging Scripting Google Cloud Multi-Agent Systems Reliability of Systems Containerization Kubernetes Google Cloud Functions Machine Learning Operations Terraform Data Pipelines Docker Jenkins

Job description

We’re looking for a DevOps Engineer to help design, build, and optimize the cloud infrastructure powering our machine learning operations. You’ll play a key role in scaling AI models from research to production - ensuring smooth deployments, real-time monitoring, and rock-solid reliability across our Google Cloud Platform (GCP) environment.

You’ll work hand-in-hand with data scientists, ML engineers, and other DevOps experts to automate workflows, enhance performance, and keep our AI systems running seamlessly for millions of players worldwide. We’re also looking for someone with strong leadership and team management capabilities who can mentor engineers, coordinate initiatives, and help drive operational excellence across the team.

What You’ll Do:

  • Manage, configure, and automate cloud infrastructure using tools such as Terraform and Ansible.
  • Implement CI/CD pipelines for ML models and data workflows, focusing on automation, versioning, rollback, and monitoring with tools like Vertex AI, Jenkins, and DataDog.
  • Build and maintain scalable data and feature pipelines for both real-time and batch processing using BigQuery, BigTable, Dataflow, Composer, Pub/Sub, and Cloud Run.
  • Set up infrastructure for model monitoring and observability - detecting drift, bias, and performance issues using Vertex AI Model Monitoring and custom dashboards.
  • Optimize inference performance, improving latency and cost-efficiency of AI workloads.
  • Ensure overall system reliability, scalability, and performance across the ML/Data platform.
  • Define and implement infrastructure best practices for deployment, monitoring, logging, and security.
  • Troubleshoot complex issues affecting ML/Data pipelines and production systems.
  • Ensure compliance with data governance, security, and regulatory standards, especially for real-money gaming environments.
  • Lead and mentor DevOps engineers, helping guide technical decisions and operational processes.
  • Support sprint planning, task prioritization, and cross-functional coordination across infrastructure and platform initiatives.
  • Conduct code reviews, share best practices, and contribute to building a high-performing engineering culture.
  • Collaborate closely with ML, Data, Product, and Security teams to align infrastructure strategy with business objectives.

Requirements

Do you have experience in Stakeholder relationship building?, * 5+ years of experience as a DevOps Engineer, ideally with a focus on ML and Data infrastructure.

  • Experience leading projects, mentoring engineers, or managing technical teams.
  • Strong hands-on experience with Google Cloud Platform (GCP) - especially BigQuery, Dataflow, Vertex AI, Cloud Run, and Pub/Sub.
  • Proficiency with Terraform (and bonus points for Ansible).
  • Solid grasp of containerization (Docker, Kubernetes) and orchestration platforms like GKE.
  • Experience building and maintaining CI/CD pipelines, preferably with Jenkins.
  • Strong understanding of monitoring and logging best practices for cloud and data systems.
  • Scripting experience with Python, Groovy, or Shell.
  • Familiarity with AI orchestration frameworks (LangGraph or LangChain) is a plus.
  • Strong communication, collaboration, and stakeholder management skills.
  • Bonus points if you’ve worked in gaming, real-time fraud detection, or AI-driven personalization systems.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

4:55 min

Discussing modern Java language evolution and syntax innovations

Daniel Strmečki · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

4:17 min

Generating static microsites for technical documentation using DocToolchain

Johannes Dienst · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all