Senior Software Engineer - ML/LLM Serving

ALLDUS INTERNATIONAL CONSULTING, INC.
San Jose, CA, United States
6 days ago
Apply on www.careerboard.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$180,000.0 - $220,000.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Profiling Computer Programming Data Security Distributed Systems Python (Programming Language) Machine Learning Recommender Systems Tensorflow Prometheus Software Engineering
+10 more
Google Cloud Large Language Models Grafana Caching Build Management Kubernetes Low Latency Optimization Algorithms Machine Learning Operations Docker

Job description

We are seeking a senior/staff Machine Learning Serving Software Engineer who thrives in a fast-paced, customer-focused environment and can build robust, flexible infrastructure to serve a diverse range of ML models including both LLMs and classical ML., As an ML Serving Engineer, you will design, implement, and optimize infrastructure that powers the deployment and inference of machine learning models across varied customer environments. You’ll work closely with product, research, and customer engineering teams to deliver low-latency, secure, and scalable ML serving solutions. Responsibilities

  • Design and build scalable, high-performance ML serving infrastructure capable of handling diverse model types (LLMs, recommendation systems, etc.).
  • Optimize inference pipelines for latency, throughput, and cost efficiency.
  • Integrate with a wide range of customer environments, adapting serving strategies to fit their infrastructure and compliance needs.
  • Deploy, monitor, and maintain ML models in production using modern deployment stacks.
  • Collaborate with ML researchers to operationalize new models and ensure seamless integration into customer workflows.
  • Ensure security and privacy best practices are applied to model deployment and inference, aligning with enterprise-grade data security requirements.
  • Stay up-to-date with the latest serving technologies and frameworks, evaluating and integrating them where relevant.

Requirements

Required

  • 5+ years of professional software engineering experience, with at least 3+ years focused on ML serving, inference infrastructure, or similar domains.
  • Proven experience deploying and optimizing large language models (LLMs) in production.
  • Hands-on expertise with multiple ML serving frameworks (eg, TensorFlow Serving, TorchServe, Triton Inference Server, BentoML, Ray Serve, vLLM, etc.).
  • Strong programming skills in Python, Go, or C+.
  • Experience with distributed systems and container orchestration tools (Kubernetes, Docker).
  • Familiarity with observability stacks (Prometheus, Grafana, OpenTelemetry) and performance profiling for inference workloads.
  • Solid understanding of secure data handling and privacy-preserving ML practices.
  • Knowledge of cloud platforms (AWS, GCP, Azure) and hybrid/on-prem deployment scenarios.

Preferred

  • Prior experience serving multiple model types beyond LLMs, eg, recommendation engines and classical ML models.
  • Exposure to model quantization, distillation, caching, and other optimization techniques for inference efficiency.
  • Experience working with enterprise customers or within compliance-heavy environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerboard.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:45 min

Introduction to serving large language models locally

Patrick Koss Patrick Koss · World Congress 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all