> Markdown version of [/jobs/ext/1901603-ml-cloud-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/1901603-ml-cloud-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML & Cloud Infrastructure Engineer - **Company:** Gritt Robotics Inc. - **Location:** South San Francisco, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Microsoft Azure, C++ (Programming Language), Cloud Computing, Cloud Engineering, Python (Programming Language), Performance Tuning, Systems Development Life Cycle, Tensorflow, Software Engineering, Parquet, Graphics Processing Unit (GPU), Data Ingestion, Pytorch, Kubernetes, Information Technology, Data Management, Machine Learning Operations, Data Pipelines, Docker - **Published:** July 31, 2026 - **Apply:** https://www.careerbuilder.com/job-details/ml-cloud-infrastructure-engineer-south-san-francisco-ca--17b7d8ba-7a0d-4830-823b-7fa8be98346a ## About the Role * Degree in computer science or related engineering disciplines (or equivalent experience). * 4+ years of experience deploying high-performance ML pipelines in production. * Proficient in Python and comfortable with C++/Go. * Experience with ML frameworks like PyTorch. * Experience with IO and data-loading workflows, including formats like Parquet, HDF5, TFRecord etc. * Experience with deploying on cloud platforms like AWS, GCP or Azure. * Experience with tooling like Docker, Kubernetes, and Airflow. * Should be comfortable taking ownership of tasks with light supervision. * Must have excellent problem-solving skills. * Legally authorized to work in the United States. Skills: Artificial Intelligence (AI), Cloud Architecture, Cloud Computing, Computer Science, Construction, Cross-Functional, Data Management, Docker, GPU (Graphics Processing Unit), Input/Output, Machine Tool, Metrics, Performance Tuning/Optimization, Reporting Dashboards, Robotics, Scalable System Development, Software Engineering, Software Evaluation, Startup, Team Player, Verification Plans ## Description We're looking for an experienced ML & Cloud Infrastructure Engineer to join our team. As an early member, you will play a pivotal role in architecting scalable cloud infrastructure for our AI and data pipelines. You'll need to thrive in a fast-paced startup environment where you'll wear multiple hats and have a direct impact on our product's evolution. Ideally, you have a proven track record of developing and deploying high-performance ML and cloud pipelines in production, and you're passionate about pushing the boundaries of what's possible in robotics with AI. What you'll get to work on * Develop and deploy scalable AI training and validation pipelines in the cloud. * Spin up distributed pipelines for data ingestion, pre-processing, training and evaluation. * Deploy monitoring and CI/CD pipelines. * Enable large-scale evaluation of AI models via cloud-based metrics. * Enable large-scale evaluation of autonomy software and models via simulations in the cloud. * Optimize performance, I/O and GPU utilization. * Build tooling and dashboards for rapid experimentation, orchestration and visualization. * Work with other teams to integrate cloud tooling into workflows. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [13 AI Tools for Developers](https://www.wearedevelopers.com/magazine/302-13-ai-tools-for-developers) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)