Senior Software Engineer - ML/LLM Serving
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
We are seeking a senior/staff Machine Learning Serving Software Engineer who thrives in a fast-paced, customer-focused environment and can build robust, flexible infrastructure to serve a diverse range of ML models including both LLMs and classical ML., As an ML Serving Engineer, you will design, implement, and optimize infrastructure that powers the deployment and inference of machine learning models across varied customer environments. You’ll work closely with product, research, and customer engineering teams to deliver low-latency, secure, and scalable ML serving solutions. Responsibilities
- Design and build scalable, high-performance ML serving infrastructure capable of handling diverse model types (LLMs, recommendation systems, etc.).
- Optimize inference pipelines for latency, throughput, and cost efficiency.
- Integrate with a wide range of customer environments, adapting serving strategies to fit their infrastructure and compliance needs.
- Deploy, monitor, and maintain ML models in production using modern deployment stacks.
- Collaborate with ML researchers to operationalize new models and ensure seamless integration into customer workflows.
- Ensure security and privacy best practices are applied to model deployment and inference, aligning with enterprise-grade data security requirements.
- Stay up-to-date with the latest serving technologies and frameworks, evaluating and integrating them where relevant.
Requirements
Required
- 5+ years of professional software engineering experience, with at least 3+ years focused on ML serving, inference infrastructure, or similar domains.
- Proven experience deploying and optimizing large language models (LLMs) in production.
- Hands-on expertise with multiple ML serving frameworks (eg, TensorFlow Serving, TorchServe, Triton Inference Server, BentoML, Ray Serve, vLLM, etc.).
- Strong programming skills in Python, Go, or C+.
- Experience with distributed systems and container orchestration tools (Kubernetes, Docker).
- Familiarity with observability stacks (Prometheus, Grafana, OpenTelemetry) and performance profiling for inference workloads.
- Solid understanding of secure data handling and privacy-preserving ML practices.
- Knowledge of cloud platforms (AWS, GCP, Azure) and hybrid/on-prem deployment scenarios.
Preferred
- Prior experience serving multiple model types beyond LLMs, eg, recommendation engines and classical ML models.
- Exposure to model quantization, distillation, caching, and other optimization techniques for inference efficiency.
- Experience working with enterprise customers or within compliance-heavy environments.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
The Best Large Language Models on The Market
Highest Paying Tech Companies for Developers