WeAreDevelopers LIVE • Oct 12, 2020

What do language models really learn

Tanmay Bakshi

Are large language models just sophisticated readers disguised as writers? Discover why autoregressive models lack underlying thought and how blank-filling architectures will finally teach neural networks to plan.

Pause
Mute Enter Fullscreen
#1 about 7 min

Introduction to language models and human-computer communication

How generating natural language resolves point-and-click friction by enabling intuitive human-computer interfaces.

#2 about 3 min

Limitations of syntax in natural language processing

How rigid syntax rules fail to capture true meaning without human cognitive interpretation.

#3 about 4 min

How neural networks warp space for linear separability

Deep learning models transform data embeddings into linearly separable clusters to classify complex information.

#4 about 3 min

Defining model behavior through specific training objectives

How mathematical incentives and loss functions strictly dictate the internal representations learned by networks.

#5 about 5 min

Word embeddings and recurrent neural network limitations

Why sequence-based architectures struggle with processing speed and contextualizing identical words with different meanings.

#6 about 4 min

Transformer architectures and bidirectional context in BERT

How the self-attention mechanism solves slow processing by analyzing entire sequences simultaneously for deep context.

#7 about 5 min

Building a character-level transformer for masked prediction

Applying self-attention and positional embeddings to a scaled-down transformer to predict masked characters in text.

#8 about 5 min

Analyzing learned clusters and syntax trees in BERT

How mathematically optimized networks independently deduce phonetic similarities and linguistic syntax structures during pre-training.

#9 about 5 min

The fallacy of autoregressive language generation models

Why relying on greedy sampling and next-word prediction fails to replicate authentic human cognitive writing.

#10 about 4 min

Addressing the hype surrounding creative language generation

Examining the logical inconsistencies in large generative models to separate artificial creativity hype from reality.

#11 about 4 min

Implementing blank language models for non-linear generation

Exploring an alternative training objective that samples random trajectories by iteratively filling dynamically inserted blanks.

#12 about 3 min

Accelerating machine learning research with optimized compilers

How compiled frameworks reduce implementation complexity to help researchers iterate on cutting-edge architectures faster.

Matching moments

2:47 min

The fundamental mechanics of large language models

Tomislav Tipurić Tomislav Tipurić · World Congress 2025

2:26 min

Capabilities and applications of large language models

Aditi Godbole · LIVE

2:25 min

Understanding the evolution and nature of large language models

Krzysztof Cieślak Krzysztof Cieślak · World Congress 2024

3:52 min

Evolution of large language models and natural language processing

Natalie Pistunovich · LIVE

2:53 min

Understanding capabilities and limitations of large language models

Markus Walker Markus Walker · World Congress 2023

1:46 min

Evolution and impact of large language models

Vijay Krishan Gupta +1 · LIVE