World Congress 2025 • Aug 20, 2025 • Session details

AI'll Be Back: Generative AI in Image, Video, and Audio Production

Anonymized By Request , Martin Foertsch , Thomas Endres

How do generative diffusion models map attention across video frames without hallucinating broken physics? Unpack the underlying architecture and the severe ethical battles currently reshaping multimedia workflows.

Pause
Mute Enter Fullscreen
#1 about 4 min

Overview of generative AI capabilities and current hype

Integrating generative artificial intelligence tools accelerates creative production across multiple media formats simultaneously.

#2 about 3 min

Fundamental mechanics of large language models and tokenization

Language models process sequential text data by dividing raw words into numerical tokens efficiently.

#3 about 2 min

Establishing semantic meaning through embedding vector spaces

Grouping vocabulary into multi-dimensional embeddings allows language models to evaluate conceptual relationships computationally.

#4 about 3 min

Transformer architecture and implementing the attention mechanism

Transformer algorithms utilize mathematical attention layers to correlate contextual relevance across separate inputs perfectly.

#5 about 3 min

Evolution of text-to-image models and semantic correlation

Contrastive pre-training aligns text prompts deeply with corresponding visual properties to generate accurate images.

#6 about 6 min

Reversing noise through the image diffusion process

Diffusion networks gradually denoise randomized static arrays to reconstruct cohesive final visual structures step-by-step.

#7 about 1 min

Optimizing processing time with latent diffusion models

Autoencoders compress pixel geometry into compact latent traits to accelerate visual calculations substantially.

#8 about 5 min

Adapting image generation techniques for coherent video production

Diffusion transformers slice visual data into temporal patches to maintain object consistency across consecutive frames.

#9 about 2 min

Advanced techniques for editing and extending generated videos

Modified generative prompts steer image interpolations and background alterations seamlessly across continuous video footage.

#10 about 2 min

Addressing physics limitations within generative video outputs

Current generative visual models lack inherent physics tracking, causing random geometric interactions and missing object permanence.

#11 about 3 min

Ethical concerns regarding AI training algorithms and data

Harvesting public media for model training introduces unresolved compliance issues around data consent and copyright liability.

Matching moments

59 sec

Creating automated generative media and artificial intelligence products

Nico Axtmann · WWC 2022

8:37 min

AI video generation tools and content creation challenges

Chris Heilmann +2 · LIVE

1:34 min

Mitigating the inherent challenges of generative AI tools

Mary Grygleski Mary Grygleski · LIVE

2:01 min

Evolution of generative artificial intelligence in modern technology

Romero Romero · WWC 2024

1:37 min

Exploring common use cases for modern generative AI

Sandra Ahlgrimm Sandra Ahlgrimm +1 · WWC 2024

3:47 min

Examples of generative AI applications in media and code

Cheuk Ho · WWC 2023

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Building Stuff with GenAI - The Open Minded Workshop beyond OpenAI

Andreas Erben

CTO for Applied AI and Metaverse at daenet

Andreas Erben
Open session

World Congress 2026 North America

The Broken Rung: How AI is Rebuilding Software Development from the Ground Up

Tomislav Tipurić

Chief Technology Officer, Nephos

Tomislav Tipurić
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

Engineering the Pivot: How Creative Strategy Solves the Hard Problems of AI Accuracy and Scale

Shruti Tiwari

AI/ML product manager, Dell

Shruti Tiwari
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

Who Tests the AI? Building Trustworthy AI Systems at Enterprise Scale

Him Raj Singh

PayPal, Manager, Software Engineer

Him Raj Singh