> Markdown version of [/jobs/ext/2105881-ai-ml-developer-llmops-model-deployment](https://www.wearedevelopers.com/jobs/ext/2105881-ai-ml-developer-llmops-model-deployment). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI/ML Developer (LLMOps & Model Deployment) - **Company:** Shimento, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Temporary contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Microsoft Azure, Cloud Computing, Program Optimization, Nvidia CUDA, Continuous Integration, Monitoring of Systems, Python (Programming Language), Machine Learning, Open Source Technology, Performance Tuning, Tensorflow, Azure Machine Learning, Software Deployment, Software Engineering, Data Logging, Graphics Processing Unit (GPU), Pytorch, Delivery Pipeline, Large Language Models, Backend, Fastapi, Containerization, Kubernetes, Low Latency, Optimization Algorithms, HuggingFace, Azure AKS, Machine Learning Operations, TensorRT, Api Design, Docker, Key Vault - **Published:** August 18, 2026 - **Apply:** https://www.dice.com/job-detail/53d5addb-7c4a-46c3-bc9d-67f70207de3e ## About the Role We are seeking a highly skilled AI/ML Engineer to lead the deployment, optimization, and scaling of open-weights foundation models within our Azure cloud ecosystem. In this role, you will bridge the gap between machine learning and core software engineering, turning raw models into highly available, low-latency APIs that power our cloud applications. The ideal candidate has a deep understanding of LLMOps, hands-on experience optimizing model inference (including knowledge distillation), and a proven track record of architecting production-grade infrastructure on Azure., Experience: 4+ years of professional experience as an ML Engineer, Data Scientist, or Backend Engineer with a heavy focus on AI deployment. Cloud Proficiency: Strong hands-on experience with Microsoft Azure (Azure ML, AKS, Azure Container Apps, Key Vault). AI/ML Frameworks: Deep proficiency with PyTorch, Hugging Face ecosystem (Transformers, Accelerate), and LangChain or LlamaIndex. API & Backend: Strong Python programming skills and experience with containerization (Docker, Kubernetes) and API development (FastAPI). Model Optimization: Demonstrated experience with model distillation, pruning, quantization, and utilizing inference engines like vLLM or TGI. Soft Skills & Culture Fit Problem Solver: Ability to triage infrastructure bottlenecks, memory constraints (OOM errors), and CUDA-related issues independently. Collaborator: Comfortable working cross-functionally with backend engineers, product managers, and security teams. Cost-Conscious Mindset: A sharp focus on balancing model accuracy with cloud spend and compute efficiency. Preferred Qualifications Certifications such as Azure AI Engineer Associate or Azure Solutions Architect. Experience implementing secure guardrails (e.g., NeMo Guardrails, Azure AI Content Safety). Contributions to open-source ML/LLMOps projects. ## Description Model Deployment & API Development Host & Scale Open-Weights Models: Deploy and maintain state-of-the-art open-weights models (e.g., Llama, Mistral, Phi) on Azure infrastructure. API Engineering: Design, build, and secure high-performance, low-latency REST/gRPC APIs (using frameworks like FastAPI) to serve models to downstream cloud applications. Inference Optimization: Implement advanced optimization techniques (e.g., vLLM, TensorRT-LLM, DeepSpeed) to maximize throughput and minimize time-to-first-token (TTFT). LLMOps & Infrastructure Pipeline Automation: Build and manage end-to-end LLMOps pipelines for continuous integration, deployment, and monitoring of models. Azure Cloud Architecture: Leverage Azure AI Studio, Azure Machine Learning (Azure ML), Azure Kubernetes Service (AKS), and managed compute (GPUs like A100/H100) efficiently. Monitoring & Observability: Implement comprehensive logging, tracing, and evaluations for LLM outputs (tracking drift, latency, costs, and hallucination rates). Model Efficiency & Distillation Knowledge Distillation: Train smaller, task-specific student models from larger, high-performing teacher models to reduce operational costs and latency without sacrificing accuracy. Quantization & Fine-Tuning: Apply quantization techniques (AWQ, GPTQ, GGUF) and parameter-efficient fine-tuning (PEFT/LoRA) to adapt models to specific business domains. ## Related Videos - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Intro to FastAPI](https://www.wearedevelopers.com/videos/462-intro-to-fastapi) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)