> Markdown version of [/jobs/ext/1009481-cloud-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/1009481-cloud-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloud Infrastructure Engineer - **Company:** Gatik Carrier - **Location:** Mountain View, CA, United States - **Experience:** Expert - **Salary:** $180,000.0 - $240,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Bash Shell, Cloud Computing, Cloud Engineering, Computer Clusters, Continuous Integration, Information Engineering, Data Infrastructure, DevOps, Distributed Computing Environment, Distributed Systems, Identity and Access Management, InfiniBand, Python (Programming Language), Role-Based Access Control, Prometheus, Pytorch, Grafana, Apache Spark, Gaussian, AI Platforms, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, ONNX (Open Neural Network Exchange) Format, Apache Kafka, Machine Learning Operations, Virtual Agents, Terraform, Data Pipelines - **Published:** June 9, 2026 - **Apply:** https://www.dice.com/job-detail/1d46cae3-85c0-46dc-8b1e-3e4798e1b994 ## About the Role * Experience: 5+ years in Cloud Infrastructure, DevOps, or MLOps supporting high-scale compute environments. * Kubernetes Mastery: Deep expertise in K8s, Helm, and container orchestration. * Orchestration & Tooling: Strong background in Apache Airflow, Argo Workflows, MLFlow, and Terraform. * Distributed Systems: Practical experience supporting frameworks like Ray and PyTorch Distributed. * Core Skills: Proficiency in Python, Bash scripting, and a solid understanding of IAM/RBAC. Bonus Qualifications * Distributed Training Expertise: Deep understanding of FSDP, and DeepSpeed. * AI Agent Orchestration: Experience building Agentic Workflows (LangGraph, AutoGen) for infrastructure automation or data curation. * Advanced Protocols: Familiarity with Model Context Protocol (MCP) to connect AI agents with infrastructure tools. ## Description We are seeking a Senior Cloud Infrastructure Engineer to architect and manage the large-scale compute and data infrastructure powering our autonomous driving stack. While researchers develop perception, planning, and world models, your mission is to build the high-performance systems and pipelines that make their work possible. You will be the backbone of our AI platform, ensuring that multi-GPU clusters, distributed training frameworks, and automated workflows are scalable, resilient, and cost-effective. This role is onsite 5 days a week at our Mountain View, CA office! What you'll do * Cloud-Native Orchestration & Kubernetes + Advanced K8s Management: Architect and maintain mission-critical Kubernetes clusters optimized for heavy GPU/TPU workloads. + GPU Scheduling: Implement and optimize Kubernetes-native GPU scheduling (NVIDIA GPU Operator) to ensure maximum hardware utilization. + Infrastructure as Code: Drive the "Everything as Code" philosophy using Terraform, Helm, and cloud-native tools. + Self-Healing Infrastructure: Deploy Autonomous AI Agents (LangGraph, CrewAI) to monitor cluster health and enable automated triage of hardware failures and NCCL timeouts. * Data Engineering & CI/CD Pipelines + Autonomy Data Pipelines: Build large-scale pipelines using Apache Airflow, Kafka, and Spark to process raw sensor data into training-ready formats. + GitOps: Implement robust GitOps workflows using ArgoCD, Gitlab CI/CD to automate the deployment of both infrastructure and model artifacts. + Observability: Maintain deep visibility into infrastructure health and model serving performance using Prometheus, Grafana, and OpenTelemetry. + Agentic DevOps & CI/CD: Develop agent-driven workflows to optimize the developer experience, such as automated PR reviewers for Terraform and AI agents that proactively suggest Kubernetes resource-limit adjustments based on model training telemetry. * Model Management & Lifecycle (MLOps) + Experiment & Model Tracking: Design and maintain MLFlow and feature store integrations to provide a robust system of record for every model iteration. + Workflow Automation: Build complex, automated model lifecycles using Airflow and Kubernetes to streamline the transition from training to simulation. + High-Performance Serving: Support the deployment of models into simulation and production environments using Triton Inference Server, Ray Serve, and ONNX Runtime. * Distributed Training & ML Systems Support + Training Systems Support: Enable researchers to scale models (VLA, World Models) across multi-node setups using PyTorch Distributed (TorchElastic), Ray Train, and Horovod. + Networking Optimization: Optimize low-level communication (e.g., NCCL tuning, InfiniBand, or RoCE v2) to minimize latency for 3D Gaussian Splatting (3DGS) and large-scale training. + Hardware-Aware Orchestration: Partner with researchers to fine-tune performance across multi-node GPU clusters for FSDP and DeepSpeed workloads. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)