Inference Lead (Machine Learning Platform Engineer Lead Real-Time Inference)
Stellent IT LLC
Charlotte, NC, United States
4 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source
Tech stack
Application Programming Interfaces (APIs)
Cloud Computing
Continuous Integration
Disaster Recovery
Load Testing
Machine Learning
Performance Tuning
Reliability Engineering
Load Balancing
Autoscaling
Kubernetes
Information Technology
+2 more
Low Latency
Microservices
Job description
- Architect low-latency online inference and real-time model-serving solutions.
- Develop scalable APIs, microservices, and deployment patterns for predictive models.
- Implement Kubernetes-based deployment, autoscaling, load balancing, and traffic-management strategies.
- Conduct benchmarking, performance tuning, capacity planning, and load testing.
- Optimize latency, throughput, resource consumption, availability, and cost.
- Define monitoring, alerting, SLOs, runbooks, and incident-response practices.
- Build CI/CD pipelines for repeatable model and service releases.
- Design resilience, failover, rollback, disaster recovery, and graceful-degradation patterns.
- Lead technical reviews and mentor inference and platform engineers.
Requirements
- Online inference and real-time model-serving architecture.
- REST/gRPC APIs and distributed microservices.
- Kubernetes, containers, autoscaling, and traffic management.
- Performance engineering, latency optimization, and load testing.
- Monitoring, SLOs, capacity planning, and production operations.
- CI/CD and progressive-deployment approaches.
- Resilience and high-availability engineering.
- Cloud and on-premises deployment experience., * Degree in computer science, engineering, or a related discipline.
- Experience with enterprise model-serving platforms and inference runtimes.
- Cloud, Kubernetes, SRE, or ML engineering certification.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
CH
Chris Heilmann
about 2 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
26 days ago
DC
Daniel Cranney
Stephan Gillich - Bringing AI Everywhere
almost 2 years ago
LM
Luis Minvielle
What Are Large Language Models?
almost 3 years ago
ER
Erin Rifkin
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
about 1 year ago
LM
Luis Minvielle
How to Become an AI Engineer
almost 3 years ago