> Markdown version of [/jobs/ext/1314941-senior-devops-mlops-engineer](https://www.wearedevelopers.com/jobs/ext/1314941-senior-devops-mlops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior DevOps/MLOps Engineer - **Company:** Leverege LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Computer Vision, Microsoft Azure, Bash Shell, Continuous Integration, Software Debugging, Linux, DevOps, Elasticsearch, Identity and Access Management, Python (Programming Language), Key Management, PostgreSQL, Node.Js, Redis, Prometheus, Data Streaming, Scripting, Google Cloud, Grafana, RTSP, Kubernetes, Sentry, Cloudflare, Machine Learning Operations, Video Streaming, Terraform, Data Pipelines, Golang - **Published:** July 17, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=fc2c5ef617a021a1 ## About the Role * 8+ years in DevOps, SRE, platform, or infrastructure engineering, with real production ownership. * Expert-level Kubernetes in production (managed Kubernetes such as GKE strongly preferred). * Strong Infrastructure-as-Code experience with Terraform and Helm. * Deep experience on a major cloud provider (GCP preferred; strong AWS or Azure background with willingness to work primarily in GCP is fine). * Experience operating a distributed fleet of remote or edge servers, or comparable experience managing infrastructure across many isolated environments. * Hands-on experience running GPU workloads and/or deploying ML models to production (inference serving, model rollout). * Solid CI/CD, Linux, networking, and scripting (Python, Node, Go, or Bash) fundamentals. * Experience as a primary on-call responder for production systems. Qualifications (Preferred) * Experience with model-serving stacks (Triton) and computer vision or ML data pipelines. * GitOps with ArgoCD, and observability with the Prometheus operator and Grafana. * Operating stateful services in Kubernetes (CloudNativePG/PostgreSQL, Elasticsearch, Redis) and event streaming (Pub/Sub). * Networking and connectivity for distributed fleets (Tailscale, Cloudflare) and camera or video streaming such as RTSP. * SOC 2 or similar compliance experience, and familiarity with GCP access controls (IAM, PAM, Secret Manager, External Secrets). * Exposure to edge hardware and on-site compute in real deployments. ## Description You will own the infrastructure that runs Leverege's VisionAI in production, in the cloud and at the edge, and build the systems that let us operate it at scale. * Own and operate the GKE platform across our per-customer Google Cloud projects, including provisioning new customer clusters end to end with Terraform, Helm, and GitOps. * Build the edge fleet management systems that let us deploy, monitor, update, and roll back software across a growing fleet of on-site edge servers, replacing manual per-server work with automated, auditable processes. * Run the MLOps path for computer vision by building repeatable pipelines that get GPU workloads and CV models from the ML team onto inference nodes and the edge fleet reliably. * Keep production healthy and observable with strong instrumentation and alerting (Prometheus, Grafana, Sentry), and serve as a primary responder on the on-call rotation. * Optimize cost and capacity across GKE and GPU node pools, balancing spend against reliability as the deployment count grows. * Harden security and support compliance through least-privilege access, proper secrets management, and SOC 2 evidence for the infrastructure you own. Who You Are * A driver, not a passenger. You take ownership of production and edge systems end to end and act before being asked. No one will look over your shoulder, and you prefer it that way. * A fleet thinker. You instinctively design for many machines across many environments, not one server at a time, and you plan for intermittent connectivity, remote updates, and things going wrong far from your keyboard. * Calm and methodical under pressure. When production breaks or an edge site goes dark, you debug systematically and communicate clearly rather than thrashing. * Direct and collaborative. You push back on risky changes, escalate straight to the right person regardless of rank, and tell product and ML teams honestly when something is not safe to ship, while staying kind and easy to work with async. * An automator by instinct. You would rather build the tool once than do the toil forever, and you improve shared tooling so the whole team moves faster. * Comfortable with autonomy and ambiguity. This is a fast-scaling environment where not everything is documented yet. If you need a lot of structure and hand-holding, this is not the right fit., * Computer vision applied to the physical world. We build proprietary vision models that solve operational problems for businesses in automotive, manufacturing, and retail, running on infrastructure you will own. * Infrastructure that is core to the business. The edge fleet and platform are how the product reaches customers, so your work is visible and high-impact. * Clear product-market fit with room to run. Fortune 500 customers, a growing portfolio of products, and large markets with little direct competition. * Fully remote, permanently. Work from anywhere in the US. We have committed to remote for good and have figured out how to make it work. * High-trust culture without politics. Smart, kind, driven people. High expectations, but the hard part is the problems, not internal friction. * A company scaling from startup to mid-size. Real opportunity to shape how our infrastructure and MLOps practices grow as we scale toward hundreds of deployments. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Computer Vision from the Edge to the Cloud done easy](https://www.wearedevelopers.com/videos/263-computer-vision-from-the-edge-to-the-cloud-done-easy) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)