Senior DevOps/MLOps Engineer

Leverege LLC
United States
25 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Computer Vision Microsoft Azure Bash Shell Continuous Integration Software Debugging Linux DevOps Elasticsearch Identity and Access Management Python (Programming Language) Key Management
+17 more
PostgreSQL Node.Js Redis Prometheus Data Streaming Scripting Google Cloud Grafana RTSP Kubernetes Sentry Cloudflare Machine Learning Operations Video Streaming Terraform Data Pipelines Golang

Job description

You will own the infrastructure that runs Leverege’s VisionAI in production, in the cloud and at the edge, and build the systems that let us operate it at scale.

  • Own and operate the GKE platform across our per-customer Google Cloud projects, including provisioning new customer clusters end to end with Terraform, Helm, and GitOps.
  • Build the edge fleet management systems that let us deploy, monitor, update, and roll back software across a growing fleet of on-site edge servers, replacing manual per-server work with automated, auditable processes.
  • Run the MLOps path for computer vision by building repeatable pipelines that get GPU workloads and CV models from the ML team onto inference nodes and the edge fleet reliably.
  • Keep production healthy and observable with strong instrumentation and alerting (Prometheus, Grafana, Sentry), and serve as a primary responder on the on-call rotation.
  • Optimize cost and capacity across GKE and GPU node pools, balancing spend against reliability as the deployment count grows.
  • Harden security and support compliance through least-privilege access, proper secrets management, and SOC 2 evidence for the infrastructure you own.

Who You Are

  • A driver, not a passenger. You take ownership of production and edge systems end to end and act before being asked. No one will look over your shoulder, and you prefer it that way.
  • A fleet thinker. You instinctively design for many machines across many environments, not one server at a time, and you plan for intermittent connectivity, remote updates, and things going wrong far from your keyboard.
  • Calm and methodical under pressure. When production breaks or an edge site goes dark, you debug systematically and communicate clearly rather than thrashing.
  • Direct and collaborative. You push back on risky changes, escalate straight to the right person regardless of rank, and tell product and ML teams honestly when something is not safe to ship, while staying kind and easy to work with async.
  • An automator by instinct. You would rather build the tool once than do the toil forever, and you improve shared tooling so the whole team moves faster.
  • Comfortable with autonomy and ambiguity. This is a fast-scaling environment where not everything is documented yet. If you need a lot of structure and hand-holding, this is not the right fit., * Computer vision applied to the physical world. We build proprietary vision models that solve operational problems for businesses in automotive, manufacturing, and retail, running on infrastructure you will own.
  • Infrastructure that is core to the business. The edge fleet and platform are how the product reaches customers, so your work is visible and high-impact.
  • Clear product-market fit with room to run. Fortune 500 customers, a growing portfolio of products, and large markets with little direct competition.
  • Fully remote, permanently. Work from anywhere in the US. We have committed to remote for good and have figured out how to make it work.
  • High-trust culture without politics. Smart, kind, driven people. High expectations, but the hard part is the problems, not internal friction.
  • A company scaling from startup to mid-size. Real opportunity to shape how our infrastructure and MLOps practices grow as we scale toward hundreds of deployments.

Requirements

  • 8+ years in DevOps, SRE, platform, or infrastructure engineering, with real production ownership.
  • Expert-level Kubernetes in production (managed Kubernetes such as GKE strongly preferred).
  • Strong Infrastructure-as-Code experience with Terraform and Helm.
  • Deep experience on a major cloud provider (GCP preferred; strong AWS or Azure background with willingness to work primarily in GCP is fine).
  • Experience operating a distributed fleet of remote or edge servers, or comparable experience managing infrastructure across many isolated environments.
  • Hands-on experience running GPU workloads and/or deploying ML models to production (inference serving, model rollout).
  • Solid CI/CD, Linux, networking, and scripting (Python, Node, Go, or Bash) fundamentals.
  • Experience as a primary on-call responder for production systems.

Qualifications (Preferred)

  • Experience with model-serving stacks (Triton) and computer vision or ML data pipelines.
  • GitOps with ArgoCD, and observability with the Prometheus operator and Grafana.
  • Operating stateful services in Kubernetes (CloudNativePG/PostgreSQL, Elasticsearch, Redis) and event streaming (Pub/Sub).
  • Networking and connectivity for distributed fleets (Tailscale, Cloudflare) and camera or video streaming such as RTSP.
  • SOC 2 or similar compliance experience, and familiarity with GCP access controls (IAM, PAM, Secret Manager, External Secrets).
  • Exposure to edge hardware and on-site compute in real deployments.

About the company

Leverege is hiring a Senior DevOps/MLOps Engineer. We build AI-native software that turns cameras into real-time visibility into physical operations, and our computer vision runs everywhere from managed Kubernetes in the cloud to a growing fleet of on-site edge servers in auto service centers, factories, and stores. This role owns the infrastructure that keeps all of it running and builds the systems that let us deploy and manage that edge fleet at scale.

This is a rare chance to join a high-trust, fully remote company and own platform and MLOps work that directly moves the business. If you love Kubernetes, think in terms of fleets rather than single servers, and want to own the GPU and edge infrastructure that gets computer vision to customers reliably, we would love to hear from you!, Leverege builds VisionAI software that helps businesses see what is happening across their physical operations. We connect to cameras, run proprietary computer vision models (often on a small edge appliance on-site), and deliver real-time operational insights to Fortune 500 customers across automotive service, manufacturing, and retail. Our products run on the Leverege Stack across per-customer environments in Google Cloud, with GPU and inference infrastructure supporting our machine learning team and a fleet of edge servers deployed at customer locations.

That footprint is scaling fast, from dozens of deployments toward hundreds, and the infrastructure needs to scale with it. Today, a lot of edge server provisioning and management is manual. We need someone to own the platform, harden it, and build the automation that makes deploying and updating a large edge fleet routine and safe.

You will be a senior individual contributor on the DevOps team. Success means our products ship reliably, our GPU and inference infrastructure keeps up with the ML team, incidents are rare and quickly resolved, and managing hundreds of edge servers feels as controlled as managing one.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

1:30 min

Developer interfaces for interacting with camera hardware

Flo Pachinger · LIVE

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · WWC Europe 2026

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · WWC 2022

Videos

See all

Related articles

See all