World Congress 2024 Aug 20, 2024 Session details

Semi-Supervised Learning. How to overcome the lack of labels

Alex Timashov

Data labeling is an expensive bottleneck. What if you could hit 95% accuracy with just 100 labels? Discover how semi-supervised learning makes this a reality.

Pause
Mute Enter Fullscreen
#1 about 6 min

Financial and temporal costs of data labeling

High financial costs and extensive time requirements drive the need for alternatives to fully supervised machine learning.

#2 about 5 min

Core intuition connecting supervised and unsupervised learning

Combining small labeled sets with abundant unlabeled examples bridges the gap between purely supervised and unsupervised methods.

#3 about 3 min

Applying entropy minimization and pseudo labeling

Minimizing classification entropy on unlabeled data helps establish confident decision boundaries in semi-supervised training.

#4 about 4 min

Maintaining class consistency with stochastic image augmentations

Applying random transformations like rotation and cropping to unlabeled images preserves original class identities for consistent training.

#5 about 5 min

Techniques for text augmentation and virtual adversarial attacks

Advanced augmentation strategies like back-translation and virtual adversarial training generate reliable variations for complex text and image datasets.

#6 about 3 min

Leveraging generative models to learn semantic data structures

Generative modeling extracts underlying structural patterns from unlabeled data to substantially improve basic classification boundaries.

#7 about 4 min

Applying variational autoencoders to semi-supervised classification

Integrating variational autoencoders with standard classification models significantly boosts model accuracy despite extremely small labeled sample sizes.

#8 about 1 min

Summarizing the mechanics of semi-supervised frameworks

Effectively merging supervised constraints with expansive unlabelled distributions is the foundation of modern semi-supervised loss functions.

Matching moments

2:03 min

Structuring the machine learning solution approach architecture

Lukas Kölbl · LIVE

2:33 min

Generating synthetic training data to resolve labeling challenges

Teresa Conceicao · World Congress 2022

3:00 min

Cleaning missing values and expanding datasets with data augmentation

Lutske van der Meer Lutske van der Meer · World Congress 2024

4:38 min

Preparing data and labeling for supervised learning

Antonia Hahn · World Congress 2023

1:45 min

Supervised, unsupervised, and reinforcement learning paradigms explained

Alexandra Waldherr · LIVE

1:22 min

Increasing model accuracy through synthetic data generation

Ekaterina Sirazitdinova · World Congress 2023

Upcoming sessions on this topic

Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

Building Stuff with GenAI - The Open Minded Workshop beyond OpenAI

Andreas Erben

CTO for Applied AI and Metaverse at daenet

Andreas Erben
Open session

World Congress 2026 North America

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Legare Kerrison, Cedric Clyburn

Legare Kerrison
Cedric Clyburn
Open session

World Congress 2026 North America

Clean Rooms Demystified: Architecture and Patterns for Privacy-Safe Data Collaboration

Anurag Malik

Data Tech Lead @ Intuit

Anurag Malik
Open session

World Congress 2026 North America

Docker sandboxes: protect your secrets, tokens, and personal data from AI agent mistakes

Kristiyan Velkov

Front-End Advocate | Speaker | AI & DevOps | Docker Captain | Cursor Ambassador | DevReal | Tech Blogger | Book Author

Kristiyan Velkov
Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong