ML Infrastructure Engineer

CLERA, LLC
San Mateo, CA, United States
3 days ago
Apply on www.juju.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Artificial Intelligence Amazon Web Services Microsoft Azure C++ (Programming Language) Cloud Computing Program Optimization Data Governance Distributed Systems Graph Database Python (Programming Language) Machine Learning
+11 more
Tensorflow Prometheus Search Technologies Software Deployment Grafana Backend Containerization Low Latency Machine Learning Operations Dynatrace Docker

Job description

This is a hands-on ML Infrastructure Engineer role at an early-stage enterprise AI company building a context and data governance layer that makes AI agents reliable in production. You will own the inference and model-serving infrastructure end to end, ensuring agents run fast and reliably at increasing concurrency. The work is squarely production-focused with real-world impact across regulated industries like insurance, banking, healthcare, and asset management. What You’ll Do

  • Design, build, and scale inference and model-serving infrastructure from the ground up through production deployment.
  • Optimize systems for latency, throughput, and reliability under high concurrency.
  • Collaborate closely with ML and infrastructure teams to ensure seamless integration and surface performance bottlenecks.
  • Drive solutions to infrastructure challenges across a fast-moving, cross-functional team.

Requirements

  • 5 or more years building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments.
  • Hands-on experience designing and scaling inference-serving systems using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom solutions.
  • Strong distributed systems fundamentals, including containerization and orchestration with Docker and Kubernetes.
  • Proficiency with monitoring and observability tooling for production systems, such as Prometheus, Grafana, or distributed tracing frameworks.
  • Experience deploying and managing ML workloads on cloud platforms (AWS, GCP, or Azure).
  • Proficiency in at least one systems or backend language: Python, Go, Rust, C++, or Java.
  • Comfort collaborating across both ML and infrastructure disciplines in a fast-paced environment.
  • Nice to have: experience with knowledge graphs, semantic search, or graph databases; real-time or low-latency inference systems; agentic or multi-step AI pipelines; enterprise data integration or pipeline infrastructure.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:18 min

Exploring the tiered architecture of modern machine learning stacks

Kris Howard · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all