> Markdown version of [/jobs/ext/3122647-ml-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/3122647-ml-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Infrastructure Engineer - **Company:** CLERA, LLC - **Location:** San Mateo, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Microsoft Azure, C++ (Programming Language), Cloud Computing, Program Optimization, Data Governance, Distributed Systems, Graph Database, Python (Programming Language), Machine Learning, Tensorflow, Prometheus, Search Technologies, Software Deployment, Grafana, Backend, Containerization, Low Latency, Machine Learning Operations, Dynatrace, Docker - **Published:** September 28, 2026 - **Apply:** https://www.juju.com/job/16_e0b97189 ## About the Role * 5 or more years building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments. * Hands-on experience designing and scaling inference-serving systems using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom solutions. * Strong distributed systems fundamentals, including containerization and orchestration with Docker and Kubernetes. * Proficiency with monitoring and observability tooling for production systems, such as Prometheus, Grafana, or distributed tracing frameworks. * Experience deploying and managing ML workloads on cloud platforms (AWS, GCP, or Azure). * Proficiency in at least one systems or backend language: Python, Go, Rust, C++, or Java. * Comfort collaborating across both ML and infrastructure disciplines in a fast-paced environment. * Nice to have: experience with knowledge graphs, semantic search, or graph databases; real-time or low-latency inference systems; agentic or multi-step AI pipelines; enterprise data integration or pipeline infrastructure. ## Description This is a hands-on ML Infrastructure Engineer role at an early-stage enterprise AI company building a context and data governance layer that makes AI agents reliable in production. You will own the inference and model-serving infrastructure end to end, ensuring agents run fast and reliably at increasing concurrency. The work is squarely production-focused with real-world impact across regulated industries like insurance, banking, healthcare, and asset management. What You'll Do * Design, build, and scale inference and model-serving infrastructure from the ground up through production deployment. * Optimize systems for latency, throughput, and reliability under high concurrency. * Collaborate closely with ML and infrastructure teams to ensure seamless integration and surface performance bottlenecks. * Drive solutions to infrastructure challenges across a fast-moving, cross-functional team. ## Related Videos - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)