ML Infrastructure Engineer
CLERA, LLC
San Mateo, CA, United States
3 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.juju.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source
Tech stack
Java (Programming Language)
Artificial Intelligence
Amazon Web Services
Microsoft Azure
C++ (Programming Language)
Cloud Computing
Program Optimization
Data Governance
Distributed Systems
Graph Database
Python (Programming Language)
Machine Learning
+11 more
Tensorflow
Prometheus
Search Technologies
Software Deployment
Grafana
Backend
Containerization
Low Latency
Machine Learning Operations
Dynatrace
Docker
Job description
This is a hands-on ML Infrastructure Engineer role at an early-stage enterprise AI company building a context and data governance layer that makes AI agents reliable in production. You will own the inference and model-serving infrastructure end to end, ensuring agents run fast and reliably at increasing concurrency. The work is squarely production-focused with real-world impact across regulated industries like insurance, banking, healthcare, and asset management. What You’ll Do
- Design, build, and scale inference and model-serving infrastructure from the ground up through production deployment.
- Optimize systems for latency, throughput, and reliability under high concurrency.
- Collaborate closely with ML and infrastructure teams to ensure seamless integration and surface performance bottlenecks.
- Drive solutions to infrastructure challenges across a fast-moving, cross-functional team.
Requirements
- 5 or more years building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments.
- Hands-on experience designing and scaling inference-serving systems using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom solutions.
- Strong distributed systems fundamentals, including containerization and orchestration with Docker and Kubernetes.
- Proficiency with monitoring and observability tooling for production systems, such as Prometheus, Grafana, or distributed tracing frameworks.
- Experience deploying and managing ML workloads on cloud platforms (AWS, GCP, or Azure).
- Proficiency in at least one systems or backend language: Python, Go, Rust, C++, or Java.
- Comfort collaborating across both ML and infrastructure disciplines in a fast-paced environment.
- Nice to have: experience with knowledge graphs, semantic search, or graph databases; real-time or low-latency inference systems; agentic or multi-step AI pipelines; enterprise data integration or pipeline infrastructure.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.juju.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
almost 3 years ago
LM
Luis Minvielle
What Are Large Language Models?
almost 3 years ago
BB
Benedikt Bischof
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
about 4 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
about 2 months ago
BB
Benedikt Bischof
MLOps – What’s the deal behind it?
almost 4 years ago
ER
Erin Rifkin
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
over 1 year ago