World Congress 2025 Jul 31, 2025 Session details

How AI Models Get Smarter

Ankit Patel

Standardized benchmarks won't guarantee production success. Learn how test-time reasoning and agentic workflows are turning structured English into the modern developer's most powerful programming language.

Pause
Mute Enter Fullscreen
#1 about 2 min

Leveraging chip sets and software for artificial intelligence efficiency

Hardware and software advancements enable more efficient large scale mathematical research.

#2 about 3 min

Measuring artificial intelligence capability using the public MMLU benchmark

Standardized multiple choice exams map conceptual boundaries by scoring performance against human expert averages.

#3 about 2 min

Overcoming human limitations in the pre-transformer computer vision era

Reliance on manually labeled data created significant bottlenecks for legacy architecture scaling.

#4 about 4 min

Utilizing the transformer architecture for unstructured raw text datasets

Predicting sequential vectors allows algorithms to absorb vast amounts of unstructured internet knowledge.

#5 about 3 min

Adapting pre-trained models using supervised fine tuning instructional techniques

Converting general linguistic knowledge into functional chat interfaces requires structured conversational demonstrations.

#6 about 5 min

Guiding model behavior utilizing reinforcement learning from human feedback

Deploying distinct reward algorithms automates the alignment of conversational responses with human preferences.

#7 about 2 min

Optimizing efficiency and reducing power consumption in computational hardware

System innovations drastically decrease the overall energy required to process massive architectural workloads.

#8 about 3 min

Enhancing accuracy through reasoning models and test time scaling

Programs that recursively prompt themselves to verify outputs significantly reduce factual errors.

#9 about 1 min

Dropping inference energy per token with high speed architecture

Hardware optimizations minimize the operational cost of generating ongoing programmatic output.

#10 about 2 min

Building modern applications utilizing structured natural language engineering prompts

Developers can orchestrate deep modular workflows simply by writing precise instructional sentences.

#11 about 2 min

Selecting between open source endpoints and proprietary reasoning services

Strategically choosing endpoints ensures the best balance of response quality and continuous functionality.

#12 about 1 min

Evaluating practical output quality independent of standardized quantitative benchmarks

Actual contextual testing inside target applications remains the only method for guaranteeing reliability.

#13 about 3 min

Preventing social engineering exploits with strict operational system guardrails

Applying rigid prompt boundaries keeps interfaces safe against manipulative gaslighting and forbidden inquiries.

#14 about 2 min

Designing modular application agents with integrated python tool chaining

Subdividing major features into smaller tools enables analytical programs to self direct dataset reviews.

#15 about 2 min

Exploring deep learning courses and free online organizational communities

Accessible training programs provide developers with essential templates for deploying real-world predictive utilities.

#16 about 5 min

Resolving core questions about probability and synthetic information distillation

Probabilistic mechanisms naturally create output variance while synthetic questions easily train smaller, efficient models.

Matching moments

2:37 min

Understanding core parameters and mechanics of large language models

Julián Duque Julián Duque · WWC 2025

5:01 min

Leveraging large language models for code optimization and development

Stephan Gillich Stephan Gillich +3 · WWC 2024

4:40 min

The increasing complexity and impact of modern AI

Jaap Kersten Jaap Kersten +1 · WWC 2024

2:18 min

Major breakthroughs shaping the artificial intelligence landscape

Nico Axtmann · WWC 2022

3:22 min

Evaluating advanced artificial intelligence platforms for daily recruitment

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

4:29 min

Designing AI applications defensively for inevitable failures

Krzysztof Cieślak Krzysztof Cieślak · WWC Europe 2026

Upcoming sessions on this topic

Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

Reinventing Testing Practices in the AI Era

Eric Deandrea

Java Champion & Senior Principal Software Engineer, IBM

Eric Deandrea
Open session

World Congress 2026 North America

Who Tests the AI? Building Trustworthy AI Systems at Enterprise Scale

Him Raj Singh

PayPal, Manager, Software Engineer

Him Raj Singh
Open session

World Congress 2026 North America

Engineering the Pivot: How Creative Strategy Solves the Hard Problems of AI Accuracy and Scale

Shruti Tiwari

AI/ML product manager, Dell

Shruti Tiwari
Open session

World Congress 2026 North America

Closing the Visibility Gap: Lessons from Safety Critical Agentic Systems

Vivek Pandit

Principal Engineer at Cadence

Vivek Pandit