Cloud MLOps Engineer

Insight Global
Austin, TX, United States
19 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Cloud Engineering Continuous Integration Python (Programming Language) Machine Learning Robotic Automation Software Software Deployment Data Streaming
+7 more
Management of Software Versions Scripting Kubernetes Apache Kafka Azure AKS Slurm Machine Learning Operations

Job description

We are seeking a Cloud MLOps Engineer to build and operate the cloud infrastructure that powers machine learning for a humanoid robotics platform. This role sits at the intersection of ML research, production systems, and end-user applications, with a strong focus on robot telemetry data, model lifecycle management, and production deployment.

You will enable researchers and applied ML engineers to reliably train, evaluate, and deploy models at scale, while ensuring telemetry-driven insights flow from robots in the real world back into continuous learning systems.

What You’ll Do

Design, deploy, and maintain cloud-native MLOps platforms supporting large-scale ML training, evaluation, and inference workloads

Operate Kubernetes-based infrastructure (self-managed or managed services such as GKE, EKS, or AKS) for ML workloads and data applications

Build and maintain end-to-end ML pipelines that bridge research workflows with production systems

Support robot telemetry ingestion, processing, and analytics, enabling model feedback loops from deployed humanoid robots

Integrate and operate ML tooling such as MLflow, Weights & Biases, Slurm, or similar systems for experiment tracking, scheduling, and reproducibility

Enable model deployment to production, including CI/CD for models, versioning, monitoring, and rollback strategies

Partner closely with ML researchers, perception, controls, and applications teams to productionize models safely and efficiently

Implement observability across ML systems, including model performance, data drift, and system health

Improve reliability, scalability, and security of cloud ML infrastructure supporting real-world robotic systems

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global’s Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.

Requirements

Strong experience with cloud platforms: AWS, GCP, and/or Azure

Hands-on experience operating Kubernetes or managed Kubernetes services in production

Experience building or maintaining MLOps platforms supporting training and inference

Familiarity with ML experiment tracking and orchestration tools (e.g., MLflow, Weights & Biases, Slurm, Ray, Kubeflow, or similar)

Experience deploying ML models into production-facing applications or services

Strong understanding of CI/CD, infrastructure-as-code, and automation

Proficiency in Python; experience with Bash or another scripting language

Ability to collaborate effectively across research and engineering teams Experience working with robotics or real-time telemetry data

Familiarity with streaming data systems (e.g., Kafka, Pub/Sub, Kinesis)

Experience supporting GPU workloads in cloud or Kubernetes environments

Exposure to edge-cloud ML deployment or fleet-based systems

Prior work in robotics, autonomy, or embodied AI environments

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer · WWC 2023

5:01 min

Container hosting options available on Microsoft Azure

Federico Fregosi · WWC 2022

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · WWC Europe 2026

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

2:46 min

Evaluating managed container services and migration strategies

Adam Bien · WWC 2021

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · WWC Europe 2026

Videos

See all

Related articles

See all