Inference Lead (Machine Learning Platform Engineer Lead Real-Time Inference)

Stellent IT LLC
Charlotte, NC, United States
4 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Cloud Computing Continuous Integration Disaster Recovery Load Testing Machine Learning Performance Tuning Reliability Engineering Load Balancing Autoscaling Kubernetes Information Technology
+2 more
Low Latency Microservices

Job description

  • Architect low-latency online inference and real-time model-serving solutions.
  • Develop scalable APIs, microservices, and deployment patterns for predictive models.
  • Implement Kubernetes-based deployment, autoscaling, load balancing, and traffic-management strategies.
  • Conduct benchmarking, performance tuning, capacity planning, and load testing.
  • Optimize latency, throughput, resource consumption, availability, and cost.
  • Define monitoring, alerting, SLOs, runbooks, and incident-response practices.
  • Build CI/CD pipelines for repeatable model and service releases.
  • Design resilience, failover, rollback, disaster recovery, and graceful-degradation patterns.
  • Lead technical reviews and mentor inference and platform engineers.

Requirements

  • Online inference and real-time model-serving architecture.
  • REST/gRPC APIs and distributed microservices.
  • Kubernetes, containers, autoscaling, and traffic management.
  • Performance engineering, latency optimization, and load testing.
  • Monitoring, SLOs, capacity planning, and production operations.
  • CI/CD and progressive-deployment approaches.
  • Resilience and high-availability engineering.
  • Cloud and on-premises deployment experience., * Degree in computer science, engineering, or a related discipline.
  • Experience with enterprise model-serving platforms and inference runtimes.
  • Cloud, Kubernetes, SRE, or ML engineering certification.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:35 min

Serving production machine learning workloads with KubeFlow and KServe

Aarno Aukia · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

3:33 min

Managing node capacity with default cluster autoscaling

Mario-Leander Reimer · World Congress 2023

5:48 min

Balancing delivery latency with stream reliability and scale

Phil Cluff · LIVE

4:13 min

Empowering developers with comprehensive AI software stacks

Markus Hacker Markus Hacker +1 · World Congress 2025

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · World Congress 2022

Videos

See all

Related articles

See all