ML & Cloud Infrastructure Engineer

Gritt Robotics Inc.
South San Francisco, CA, United States
about 1 month ago
Apply on www.careerbuilder.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Airflow Amazon Web Services Microsoft Azure C++ (Programming Language) Cloud Computing Cloud Engineering Python (Programming Language) Performance Tuning Systems Development Life Cycle Tensorflow Software Engineering
+10 more
Parquet Graphics Processing Unit (GPU) Data Ingestion Pytorch Kubernetes Information Technology Data Management Machine Learning Operations Data Pipelines Docker

Job description

We’re looking for an experienced ML & Cloud Infrastructure Engineer to join our team. As an early member, you will play a pivotal role in architecting scalable cloud infrastructure for our AI and data pipelines. You’ll need to thrive in a fast-paced startup environment where you’ll wear multiple hats and have a direct impact on our product’s evolution. Ideally, you have a proven track record of developing and deploying high-performance ML and cloud pipelines in production, and you’re passionate about pushing the boundaries of what’s possible in robotics with AI.

What you’ll get to work on

  • Develop and deploy scalable AI training and validation pipelines in the cloud.

  • Spin up distributed pipelines for data ingestion, pre-processing, training and evaluation.
  • Deploy monitoring and CI/CD pipelines.
  • Enable large-scale evaluation of AI models via cloud-based metrics.
  • Enable large-scale evaluation of autonomy software and models via simulations in the cloud.
  • Optimize performance, I/O and GPU utilization.
  • Build tooling and dashboards for rapid experimentation, orchestration and visualization.
  • Work with other teams to integrate cloud tooling into workflows.

Requirements

  • Degree in computer science or related engineering disciplines (or equivalent experience).

  • 4+ years of experience deploying high-performance ML pipelines in production.
  • Proficient in Python and comfortable with C++/Go.
  • Experience with ML frameworks like PyTorch.
  • Experience with IO and data-loading workflows, including formats like Parquet, HDF5, TFRecord etc.
  • Experience with deploying on cloud platforms like AWS, GCP or Azure.
  • Experience with tooling like Docker, Kubernetes, and Airflow.

  • Should be comfortable taking ownership of tasks with light supervision.

  • Must have excellent problem-solving skills.

  • Legally authorized to work in the United States.

Skills: Artificial Intelligence (AI), Cloud Architecture, Cloud Computing, Computer Science, Construction, Cross-Functional, Data Management, Docker, GPU (Graphics Processing Unit), Input/Output, Machine Tool, Metrics, Performance Tuning/Optimization, Reporting Dashboards, Robotics, Scalable System Development, Software Engineering, Software Evaluation, Startup, Team Player, Verification Plans

About the company

Gritt is an intelligent system that combines robotics and AI to build the infrastructure that pulls society forward. Gritt deploys via simple attachments to common equipment found on construction sites and autonomously performs labor-intensive tasks, verification, and planning. Gritt systems are already building critical infrastructure in the harshest outdoor environments, starting with large-scale solar. The founding team includes experts in robotics and AI from Carnegie Mellon, Stanford, and MIT. Gritt is backed by Obvious Ventures, Union Square Ventures, First Round Capital, Climactic, Congruent Ventures, and other leading firms.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all