> Markdown version of [/jobs/ext/2835884-ml-infrastructure-engineer-model-inference](https://www.wearedevelopers.com/jobs/ext/2835884-ml-infrastructure-engineer-model-inference). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Infrastructure Engineer, Model Inference - **Company:** Abridge Partners, LLC - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Nvidia CUDA, Distributed Computing Environment, Distributed Systems, Machine Learning, Ansible, Tensorflow, Systems Architecture, Toolchain, Graphics Processing Unit (GPU), Pytorch, Large Language Models, Generative AI, Backend, Kubernetes, Machine Learning Operations, Api Design, Terraform - **Published:** September 10, 2026 - **Apply:** https://www.careerbuilder.com/job-details/machine-learning-infrastructure-engineer-model-inference-san-francisco-ca--cc001ec7-a345-4f12-bd19-1ebe36d1bc39 ## About the Role * 5+ years of experience in building and deploying machine learning models in production environments. * Deep understanding of container orchestration and distributed systems architecture * Expertise in Kubernetes administration, including custom resource definitions, operators, and cluster management * Experience developing APIs and managing distributed systems for both batch and real-time workloads * Excellent communication skills, with the ability to interface between research and product engineering Ideally, You Have * Expertise with model serving frameworks such as NVIDIA Triton Server, VLLM, TRT-LLM and so on. * Expertise with ML toolchains such as PyTorch, Tensorflow or distributed training and inference libraries. * Familiarity with GPU cluster management and CUDA optimization * Knowledge of infrastructure as code (Terraform, Ansible) and GitOps practices * Experience with container registries, image optimization, and multi-stage builds for ML workloads * Experience orchestrating across ASR models or LLM models for building various GenAI applications, Ansible, Application Programming Interface (API), Artificial Intelligence (AI), CUDA (Compute Unified Device Architecture), Clinical Study Publications, Coaching, Communication Skills, Distributed Computing, Establish Priorities, Financial Support, Fitness, GPU (Graphics Processing Unit), Health Plan, Healthcare, Industry Standards, Leadership, Machine Learning, Performance Management, Product Engineering, Production Systems, Psychiatry and Mental Health, Startup, System Architecture, Systems Administration/Management ## Description As an ML Infrastructure Engineer, Model Inference at Abridge, you'll play a pivotal role in building and optimizing the core inference infrastructure that powers our machine learning models. Your work will be instrumental in enhancing the scalability, efficiency, and performance of our AI-driven solutions. You will work with our Infrastructure and Research teams to build, deploy, optimize and orchestrate across our AI models. What You'll Do * Design, deploy and maintain scalable Kubernetes clusters for AI model inference and training * Develop, optimize, and maintain ML model serving infrastructure, ensuring high-performance and low-latency. * Collaborate with ML and product teams to scale backend infrastructure for AI-driven products, focusing on model deployment, throughput optimization, and compute efficiency. * Optimize compute-heavy workflows and enhance GPU utilization for ML workloads. * Build a robust model API orchestration system * Collaborate with leadership to define and implement strategies for scaling infrastructure as the company grows, ensuring long-term efficiency and performance. ## Related Videos - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)