Senior DevOps/MLOps Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+17 more
Job description
You will own the infrastructure that runs Leverege’s VisionAI in production, in the cloud and at the edge, and build the systems that let us operate it at scale.
- Own and operate the GKE platform across our per-customer Google Cloud projects, including provisioning new customer clusters end to end with Terraform, Helm, and GitOps.
- Build the edge fleet management systems that let us deploy, monitor, update, and roll back software across a growing fleet of on-site edge servers, replacing manual per-server work with automated, auditable processes.
- Run the MLOps path for computer vision by building repeatable pipelines that get GPU workloads and CV models from the ML team onto inference nodes and the edge fleet reliably.
- Keep production healthy and observable with strong instrumentation and alerting (Prometheus, Grafana, Sentry), and serve as a primary responder on the on-call rotation.
- Optimize cost and capacity across GKE and GPU node pools, balancing spend against reliability as the deployment count grows.
- Harden security and support compliance through least-privilege access, proper secrets management, and SOC 2 evidence for the infrastructure you own.
Who You Are
- A driver, not a passenger. You take ownership of production and edge systems end to end and act before being asked. No one will look over your shoulder, and you prefer it that way.
- A fleet thinker. You instinctively design for many machines across many environments, not one server at a time, and you plan for intermittent connectivity, remote updates, and things going wrong far from your keyboard.
- Calm and methodical under pressure. When production breaks or an edge site goes dark, you debug systematically and communicate clearly rather than thrashing.
- Direct and collaborative. You push back on risky changes, escalate straight to the right person regardless of rank, and tell product and ML teams honestly when something is not safe to ship, while staying kind and easy to work with async.
- An automator by instinct. You would rather build the tool once than do the toil forever, and you improve shared tooling so the whole team moves faster.
- Comfortable with autonomy and ambiguity. This is a fast-scaling environment where not everything is documented yet. If you need a lot of structure and hand-holding, this is not the right fit., * Computer vision applied to the physical world. We build proprietary vision models that solve operational problems for businesses in automotive, manufacturing, and retail, running on infrastructure you will own.
- Infrastructure that is core to the business. The edge fleet and platform are how the product reaches customers, so your work is visible and high-impact.
- Clear product-market fit with room to run. Fortune 500 customers, a growing portfolio of products, and large markets with little direct competition.
- Fully remote, permanently. Work from anywhere in the US. We have committed to remote for good and have figured out how to make it work.
- High-trust culture without politics. Smart, kind, driven people. High expectations, but the hard part is the problems, not internal friction.
- A company scaling from startup to mid-size. Real opportunity to shape how our infrastructure and MLOps practices grow as we scale toward hundreds of deployments.
Requirements
- 8+ years in DevOps, SRE, platform, or infrastructure engineering, with real production ownership.
- Expert-level Kubernetes in production (managed Kubernetes such as GKE strongly preferred).
- Strong Infrastructure-as-Code experience with Terraform and Helm.
- Deep experience on a major cloud provider (GCP preferred; strong AWS or Azure background with willingness to work primarily in GCP is fine).
- Experience operating a distributed fleet of remote or edge servers, or comparable experience managing infrastructure across many isolated environments.
- Hands-on experience running GPU workloads and/or deploying ML models to production (inference serving, model rollout).
- Solid CI/CD, Linux, networking, and scripting (Python, Node, Go, or Bash) fundamentals.
- Experience as a primary on-call responder for production systems.
Qualifications (Preferred)
- Experience with model-serving stacks (Triton) and computer vision or ML data pipelines.
- GitOps with ArgoCD, and observability with the Prometheus operator and Grafana.
- Operating stateful services in Kubernetes (CloudNativePG/PostgreSQL, Elasticsearch, Redis) and event streaming (Pub/Sub).
- Networking and connectivity for distributed fleets (Tailscale, Cloudflare) and camera or video streaming such as RTSP.
- SOC 2 or similar compliance experience, and familiarity with GCP access controls (IAM, PAM, Secret Manager, External Secrets).
- Exposure to edge hardware and on-site compute in real deployments.
About the company
Leverege is hiring a Senior DevOps/MLOps Engineer. We build AI-native software that turns cameras into real-time visibility into physical operations, and our computer vision runs everywhere from managed Kubernetes in the cloud to a growing fleet of on-site edge servers in auto service centers, factories, and stores. This role owns the infrastructure that keeps all of it running and builds the systems that let us deploy and manage that edge fleet at scale.
This is a rare chance to join a high-trust, fully remote company and own platform and MLOps work that directly moves the business. If you love Kubernetes, think in terms of fleets rather than single servers, and want to own the GPU and edge infrastructure that gets computer vision to customers reliably, we would love to hear from you!, Leverege builds VisionAI software that helps businesses see what is happening across their physical operations. We connect to cameras, run proprietary computer vision models (often on a small edge appliance on-site), and deliver real-time operational insights to Fortune 500 customers across automotive service, manufacturing, and retail. Our products run on the Leverege Stack across per-customer environments in Google Cloud, with GPU and inference infrastructure supporting our machine learning team and a fleet of edge servers deployed at customer locations.
That footprint is scaling fast, from dozens of deployments toward hundreds, and the infrastructure needs to scale with it. Today, a lot of edge server provisioning and management is manual. We need someone to own the platform, harden it, and build the automation that makes deploying and updating a large edge fleet routine and safe.
You will be a senior individual contributor on the DevOps team. Success means our products ship reliably, our GPU and inference infrastructure keeps up with the ML team, incidents are rare and quickly resolved, and managing hundreds of edge servers feels as controlled as managing one.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
Fully Remote Software Engineer Jobs
DevOps Engineer Salary [2023]
Is Software Engineering Over-Saturated?