> Markdown version of [/jobs/ext/553400-staff-machine-learning-systems-engineer-mlops](https://www.wearedevelopers.com/jobs/ext/553400-staff-machine-learning-systems-engineer-mlops). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Machine Learning Systems Engineer (MLOps) - **Company:** Hims, Inc. - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $210,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon Cloudfront, Amazon S3, Encodings, Continuous Integration, Data Infrastructure, Data Security, Software Design Documents, DevOps, Programming Tools, Identity and Access Management, Python (Programming Language), Key Management, Online Transaction Processing, Open Source Technology, OpenID, Reliability Engineering, Azure Machine Learning, Test Execution Engine, AI Infrastructure, Datadog, Autoscaling, System Availability, Delivery Pipeline, Large Language Models, Apache Spark, Multi-Cloud, AI Platforms, Git Flow, Kubernetes, Build Tools, Machine Learning Operations, Vertica, Terraform, Docker, Databricks - **Published:** June 14, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=730aae26f29482db ## About the Role * 8+ years of professional experience in infrastructure, platform, DevOps, or SRE engineering - with at least 3 years focused on ML/AI systems in production. * Deep, hands-on experience with Kubernetes (ideally EKS) and the cloud-native ecosystem - autoscaling, GitOps, Helm/Kustomize, operating clusters at scale, and general process/job orchestration. * Strong infrastructure-as-code skills (Terraform) and experience designing secure cloud architectures: IAM, OIDC, secrets management, and least-privilege access. * Strong proficiency in Python, with experience building production infrastructure tooling, CLIs, and data/observability pipelines. * 2+ years of experience operating LLM-based systems in production (LLMOps) - inference routing, serving, tracing, and the reliability patterns needed to run them at scale. * Hands-on experience with observability/tracing stacks (Datadog, OpenTelemetry, Langfuse, or equivalent) and metrics/log/trace pipelines. * Experience designing and maintaining CI/CD pipelines, build systems, and developer tooling for fast-moving engineering teams. * A systems-and-operations mindset: you think about failure modes, SLOs, observability, security, and long-term maintainability before shipping. * Experience writing and leading technical design documents (TDDs/RFCs) for infrastructure-scale initiatives. * Strong collaboration skills across engineering, ML, product, security, and clinical teams. * A deep appreciation for safety, privacy, and security - ideally with experience in a regulated domain such as healthcare, fintech, or life sciences. Nice to Have: * Experience with AWS (EKS, Bedrock, S3, CloudFront, IAM) and multi-cloud (GCP/Vertex AI) inference routing. * Experience with Databricks (MLflow, Unity Catalog, Spark, Delta) and data platform access governance. * Experience provisioning LLM observability infrastructure (Langfuse, ClickHouse, OpenTelemetry/OTLP tracing, LogFire) and LLM behavior monitoring. * Experience with Karpenter, cluster autoscaling, and cost optimization for ML compute. * Experience with monorepo build systems (Pants, Bazel) and large-scale CI/CD. * Experience building automated PR-review / convention-enforcement pipelines and developer-workflow standards. * Familiarity with Vertex AI Agent Builder, Vertex AI Model Registry, or GCP managed AI/ML services as a stretch growth area. * Contributions to open-source infrastructure, IaC modules, SDKs, or developer tooling projects. ## Description We're hiring a Staff ML Systems Engineer to design, build, and operate the production infrastructure that powers AI across Hims & Hers. This is a deeply technical, hands-on infrastructure role focused on the systems underneath AI - the Kubernetes platform, CI/CD and GitOps pipelines, infrastructure-as-code, inference and model-serving infrastructure, and the observability and tracing stack that keeps AI services reliable, debuggable, and compliant in production. You won't just deploy models - you'll own the machinery that lets every AI team ship and operate safely. You'll own critical systems like our EKS clusters, deployment and autoscaling infrastructure, IAM and secrets management, LLM tracing/observability pipelines (Langfuse, Datadog, OpenTelemetry), and the developer platform that AI and product engineers rely on daily. You'll partner with ML engineers, product engineers, and clinical teams to ensure our AI systems are reliable, observable, secure, and trustworthy in a regulated healthcare environment. This role is ideal for someone who thinks in systems and infrastructure, cares deeply about reliability, security, and cost, and wants to define how AI runs in production at a company where it directly impacts patient outcomes., * Build and maintain GitOps-based deployment pipelines (Helm/Kustomize overlays, environment promotion) that let teams ship AI services safely and repeatably. * Design ephemeral/preview environments, feature-branched deployments, and nightly release pipelines so teams can validate AI changes in production-like conditions before release. * Drive efficiency and cost management across compute, autoscaling, and inference infrastructure. Build inference and model-serving infrastructure * Operate and scale inference infrastructure and a multi-provider LLM AI gateway (e.g. Bedrock, Vertex, and other providers) - including credentials, rate limits, and failover. * Build reliable serving patterns for LLM-powered workflows: routing, grounding, tool execution, and context assembly at the platform level. * Create reusable infrastructure abstractions and contracts that standardize how AI services are deployed, configured, and consumed across the company. Own observability, tracing, and reliability * Own the LLM/AI observability and tracing stack - provisioning and scaling systems like Langfuse, Datadog (dd-trace), OpenTelemetry tracing (OTLP), and the underlying datastores (e.g. ClickHouse) - so AI behavior is auditable and debuggable in production. * Build analytics and monitoring pipelines that surface latency, error, quality, and regression signals to engineering and clinical stakeholders. * Define SLOs, alerting, on-call runbooks, and incident response for AI infrastructure; lead troubleshooting and continuously raise platform reliability. Scale the AI developer platform and CI/CD * Own and improve the monorepo build system and CI/CD pipelines for AI workloads - including eval workflows, Docker image builds, automated PR checks and convention enforcement, and cross-platform test execution. * Own shared infrastructure tooling, CLIs, and IaC modules (Terraform, Scalr) that AI and product engineers use daily. * Identify and eliminate platform bottlenecks - reducing CI/CD cycle times, build latency, and deployment friction - to improve developer velocity across the Applied AI organization. Drive security, compliance, and governance at the systems level * Build IAM, OIDC, and secrets management as first-class infrastructure - scoped, least-privilege roles, write-only secret rotation, and cross-account access audits. * Encode security-by-default, scope boundaries, and access controls into the platform so AI services are HIPAA-compliant and privacy-first. * Partner with clinical, legal, security, and data platform teams (including Databricks/Unity Catalog access governance) to enforce compliant, auditable data access. Set technical direction and raise the bar * Drive multi-quarter infrastructure initiatives, from cluster and deployment architecture to inference platform, GPU compute strategy, and observability evolution. * Write and lead technical design documents and design reviews, define infrastructure standards and development-workflow conventions, and contribute to technical governance across AI engineering. * Mentor engineers on reliability engineering, infrastructure-as-code, and MLOps best practices, and bridge the gap between prototypes and production-grade systems. ## Related Videos - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Keeping applications secure by evolving OAuth 2.0 and OpenID Connect](https://www.wearedevelopers.com/videos/100152-keeping-applications-secure-by-evolving-oauth-2-0-and-openid-connect) - [Delegating the chores of authenticating users to Keycloak](https://www.wearedevelopers.com/videos/1558-delegating-the-chores-of-authenticating-users-to-keycloak) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)