AI/ML Developer (LLMOps & Model Deployment)

Shimento, Inc.
United States
1 day ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Microsoft Azure Cloud Computing Program Optimization Nvidia CUDA Continuous Integration Monitoring of Systems Python (Programming Language) Machine Learning Open Source Technology Performance Tuning
+22 more
Tensorflow Azure Machine Learning Software Deployment Software Engineering Data Logging Graphics Processing Unit (GPU) Pytorch Delivery Pipeline Large Language Models Backend Fastapi Containerization Kubernetes Low Latency Optimization Algorithms HuggingFace Azure AKS Machine Learning Operations TensorRT Api Design Docker Key Vault

Job description

Model Deployment & API Development

Host & Scale Open-Weights Models: Deploy and maintain state-of-the-art open-weights models (e.g., Llama, Mistral, Phi) on Azure infrastructure.

API Engineering: Design, build, and secure high-performance, low-latency REST/gRPC APIs (using frameworks like FastAPI) to serve models to downstream cloud applications.

Inference Optimization: Implement advanced optimization techniques (e.g., vLLM, TensorRT-LLM, DeepSpeed) to maximize throughput and minimize time-to-first-token (TTFT).

LLMOps & Infrastructure

Pipeline Automation: Build and manage end-to-end LLMOps pipelines for continuous integration, deployment, and monitoring of models.

Azure Cloud Architecture: Leverage Azure AI Studio, Azure Machine Learning (Azure ML), Azure Kubernetes Service (AKS), and managed compute (GPUs like A100/H100) efficiently.

Monitoring & Observability: Implement comprehensive logging, tracing, and evaluations for LLM outputs (tracking drift, latency, costs, and hallucination rates).

Model Efficiency & Distillation

Knowledge Distillation: Train smaller, task-specific student models from larger, high-performing teacher models to reduce operational costs and latency without sacrificing accuracy.

Quantization & Fine-Tuning: Apply quantization techniques (AWQ, GPTQ, GGUF) and parameter-efficient fine-tuning (PEFT/LoRA) to adapt models to specific business domains.

Requirements

We are seeking a highly skilled AI/ML Engineer to lead the deployment, optimization, and scaling of open-weights foundation models within our Azure cloud ecosystem. In this role, you will bridge the gap between machine learning and core software engineering, turning raw models into highly available, low-latency APIs that power our cloud applications.

The ideal candidate has a deep understanding of LLMOps, hands-on experience optimizing model inference (including knowledge distillation), and a proven track record of architecting production-grade infrastructure on Azure., Experience: 4+ years of professional experience as an ML Engineer, Data Scientist, or Backend Engineer with a heavy focus on AI deployment.

Cloud Proficiency: Strong hands-on experience with Microsoft Azure (Azure ML, AKS, Azure Container Apps, Key Vault).

AI/ML Frameworks: Deep proficiency with PyTorch, Hugging Face ecosystem (Transformers, Accelerate), and LangChain or LlamaIndex.

API & Backend: Strong Python programming skills and experience with containerization (Docker, Kubernetes) and API development (FastAPI).

Model Optimization: Demonstrated experience with model distillation, pruning, quantization, and utilizing inference engines like vLLM or TGI.

Soft Skills & Culture Fit

Problem Solver: Ability to triage infrastructure bottlenecks, memory constraints (OOM errors), and CUDA-related issues independently.

Collaborator: Comfortable working cross-functionally with backend engineers, product managers, and security teams.

Cost-Conscious Mindset: A sharp focus on balancing model accuracy with cloud spend and compute efficiency.

Preferred Qualifications

Certifications such as Azure AI Engineer Associate or Azure Solutions Architect.

Experience implementing secure guardrails (e.g., NeMo Guardrails, Azure AI Content Safety).

Contributions to open-source ML/LLMOps projects.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:33 min

Connecting frontends via a FastAPI proxy backend layer

Saoussen Chaabnia Saoussen Chaabnia · Europe 2026 Virtual

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

4:57 min

Centralizing LLMOps workflows within Azure AI Foundry

Maxim Salnikov Maxim Salnikov · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

Videos

See all

Related articles

See all