> Markdown version of [/jobs/ext/2536350-senior-software-engineer-ml-llm-serving](https://www.wearedevelopers.com/jobs/ext/2536350-senior-software-engineer-ml-llm-serving). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer - ML/LLM Serving - **Company:** ALLDUS INTERNATIONAL CONSULTING, INC. - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $180,000.0 - $220,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Profiling, Computer Programming, Data Security, Distributed Systems, Python (Programming Language), Machine Learning, Recommender Systems, Tensorflow, Prometheus, Software Engineering, Google Cloud, Large Language Models, Grafana, Caching, Build Management, Kubernetes, Low Latency, Optimization Algorithms, Machine Learning Operations, Docker - **Published:** August 31, 2026 - **Apply:** https://www.careerboard.com/us/en/find-jobs-in-United-States/-6464A6B27844DAD2E8/ ## About the Role Required * 5+ years of professional software engineering experience, with at least 3+ years focused on ML serving, inference infrastructure, or similar domains. * Proven experience deploying and optimizing large language models (LLMs) in production. * Hands-on expertise with multiple ML serving frameworks (eg, TensorFlow Serving, TorchServe, Triton Inference Server, BentoML, Ray Serve, vLLM, etc.). * Strong programming skills in Python, Go, or C+. * Experience with distributed systems and container orchestration tools (Kubernetes, Docker). * Familiarity with observability stacks (Prometheus, Grafana, OpenTelemetry) and performance profiling for inference workloads. * Solid understanding of secure data handling and privacy-preserving ML practices. * Knowledge of cloud platforms (AWS, GCP, Azure) and hybrid/on-prem deployment scenarios. Preferred * Prior experience serving multiple model types beyond LLMs, eg, recommendation engines and classical ML models. * Exposure to model quantization, distillation, caching, and other optimization techniques for inference efficiency. * Experience working with enterprise customers or within compliance-heavy environments. ## Description We are seeking a senior/staff Machine Learning Serving Software Engineer who thrives in a fast-paced, customer-focused environment and can build robust, flexible infrastructure to serve a diverse range of ML models including both LLMs and classical ML., As an ML Serving Engineer, you will design, implement, and optimize infrastructure that powers the deployment and inference of machine learning models across varied customer environments. You'll work closely with product, research, and customer engineering teams to deliver low-latency, secure, and scalable ML serving solutions. Responsibilities * Design and build scalable, high-performance ML serving infrastructure capable of handling diverse model types (LLMs, recommendation systems, etc.). * Optimize inference pipelines for latency, throughput, and cost efficiency. * Integrate with a wide range of customer environments, adapting serving strategies to fit their infrastructure and compliance needs. * Deploy, monitor, and maintain ML models in production using modern deployment stacks. * Collaborate with ML researchers to operationalize new models and ensure seamless integration into customer workflows. * Ensure security and privacy best practices are applied to model deployment and inference, aligning with enterprise-grade data security requirements. * Stay up-to-date with the latest serving technologies and frameworks, evaluating and integrating them where relevant. ## Related Videos - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Unveiling the Magic: Scaling Large Language Models to Serve Millions](https://www.wearedevelopers.com/videos/1619-unveiling-the-magic-scaling-large-language-models-to-serve-millions) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)