World Congress 2025 Aug 20, 2025 Session details

LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices

Anshul Jindal

Struggling to move generative AI from experimental notebooks to scalable production? Discover how NVIDIA NIM and NeMo Microservices unify fine-tuning, evaluation, and inference into an automated LLMOps pipeline.

Pause
Mute Enter Fullscreen
#1 about 5 min

The generative AI application lifecycle

How building data curation and model training tasks into loops enables continuous generative AI application updates.

#2 about 4 min

Solving application deployment complexities using LLMOps pipelines

Overcoming disconnected deployment stages involves operationalizing the machine learning pipeline from proof-of-concept to production.

#3 about 2 min

Mapping the developer pipeline with NVIDIA microservices

Integrating specialized microservice containers directly addresses the complexity of application customization and evaluation tasks.

#4 about 3 min

Integrating infrastructure enablers for LLMOps operations ecosystems

Combining data storage and operational workflow tools establishes a resilient automation foundation for language models.

#5 about 5 min

Executing fine-tuning and LLM evaluation API workflows

Pulling foundation models and datasets via API requests enables automated container training and adapter benchmarking.

#6 about 4 min

Scaling customized inference models with NVIDIA NIM

Using specialized Kubernetes operators to deploy optimized containers scales customized adapter loading into hardware cache.

#7 about 3 min

Automating discrete component pipelines using Argo Workflows

Writing deterministic component templates stitches discrete deployment operations into reproducible pre-production software workflows.

#8 about 4 min

Managing cluster platform infrastructure via GitOps principles

Establishing version control as the core infrastructure truth aligns application development with Kubernetes cluster operations.

#9 about 3 min

Demonstrating an integrated LLMOps cluster deployment environment

Orchestrating cluster synchronization automatically triggers pipeline sequences and centralizes model training metric visualization.

Matching moments

1:27 min

Automating model lifecycle management with the NIM operator

Kevin Klues Kevin Klues

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer · WWC 2023

2:14 min

Building and fine-tuning models with the NeMo framework

Anshul Jindal Anshul Jindal +1 · WWC Europe 2026

4:19 min

Introduction to DevOps for AI and MLOps

Aarno Aukia · LIVE

4:57 min

Centralizing LLMOps workflows within Azure AI Foundry

Maxim Salnikov Maxim Salnikov · LIVE

2:04 min

Optimizing and deploying containerized AI inference workloads

Ankit Patel Ankit Patel · WWC 2024

Upcoming sessions on this topic

Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

It’s Alive! Taming the MLOps Franken-Stack: Write, Run, and Serve with Michelangelo

Eric Wang, Paul Zimmerman

Eric Wang
Paul Zimmerman
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

Trust, But Verify: Continuous GPU Validation at Scale

Kyle Bell

VP of AI @ TensorWave

Kyle Bell
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Cedric Clyburn, Legare Kerrison

Cedric Clyburn
Legare Kerrison