Senior Fleet Software Engineer

Rhoda ai
Mountain View, CA, United States
24 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services C++ (Programming Language) Data Transport Utility Software Debugging Linux on Embedded Systems Fault Tolerance Python (Programming Language) Prometheus Software Engineering Real Time Systems Grafana Kubernetes
+4 more
Low Latency Deployment Automation Data Pipelines Docker

Job description

You’ll build the software and infrastructure that lets a small team operate and monitor a growing fleet of deployed robots. The observability, alerting, and automation you build are what the team watches during a shift and what on-call responds to when something goes wrong. You’ll contribute to the fleet’s operational software end to end, from the telemetry we collect on every robot to the dashboards, pipelines, and tooling that act on it, and work closely with the operations and response teams so the fleet gets easier to run as it scales., * Build and own fleet observability: the metrics, logs, traces, and dashboards that give the team full visibility into live robots

  • Design telemetry and data pipelines that reliably move robot data to the cloud for monitoring, debugging, and model training
  • Build fleet-health dashboards and reports that make performance and regressions easy to spot, * Build alerting and on-call tooling that catches issues fast and routes them to the right responder with the right context
  • Automate diagnostics and incident capture so responders can start debugging instead of gathering data
  • Automate manual, error-prone work and reduce operational toil

Deployment and release infrastructure

  • Contribute to CI/CD pipelines that deliver code reliably from development to the fleet
  • Work with software teams to automate over-the-air (OTA) software and firmware updates across the fleet, with staged rollout, monitoring, and rollback
  • Build provisioning and configuration tooling to keep the fleet consistent and reproducible

Requirements

  • 3+ years of software engineering, with real ownership of internal tooling, infrastructure, or reliability systems
  • Strong proficiency in Python plus at least one systems language (Go, C++, or Rust)
  • Experience building and operating CI/CD pipelines and deployment automation
  • Experience with a major cloud (AWS or GCP), containers, and orchestration (Docker, Kubernetes)
  • Solid distributed-systems fundamentals across data transport, monitoring, and fault tolerance
  • Ability to debug across the full stack, from a device on the network to a service in the cloud, * Background in robotics, autonomous vehicles, or other latency- or safety-critical domains
  • Experience with observability stacks (e.g., Prometheus/Grafana, OpenTelemetry, Foxglove)
  • Experience with OTA or fleet deployment and safe-rollout patterns such as staged rollout and auto-rollback
  • Familiarity with ROS/ROS2, edge or embedded Linux, or low-latency data transport for real-time systems
  • Experience building tooling for an operations, on-call, or field team

About the company

At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We’ve raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:37 min

Example technology stack for robotics software development

Falk-Moritz Schaefer · World Congress 2022

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:05 min

Measuring system availability utilizing Prometheus and straightforward PromQL

Alexander Schwartz Alexander Schwartz · World Congress 2025

13:07 min

Configuring application observability with Micrometer and Prometheus

Aleksandr Kalikov · LIVE

4:47 min

Automating frontend performance metrics with Google Lighthouse

Miki Lombardi · JS Congress

Videos

See all

Related articles

See all