> Markdown version of [/videos/42-what-do-language-models-really-learn](https://www.wearedevelopers.com/videos/42-what-do-language-models-really-learn). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # What do language models really learn Are large language models just sophisticated readers disguised as writers? Discover why autoregressive models lack underlying thought and how blank-filling architectures will finally teach neural networks to plan. - **Speakers:** Tanmay Bakshi - **Event:** WeAreDevelopers LIVE - **Published:** October 12, 2020 - **Duration:** 47:28 - **URL:** https://www.wearedevelopers.com/videos/42-what-do-language-models-really-learn ## Summary The quest for natural, intuitive human-computer interaction has driven massive hype around large language models. However, to truly understand what these models learn, we must look past the illusion of creativity. Neural networks fundamentally act as mathematical optimizers that warp data spaces to achieve linear separability, and their behavior is heavily dictated by training objectives—incentives that directly shape their internal embeddings. The evolution from Word2Vec and slow, context-limited recurrent neural networks to the modern transformer architecture completely transformed natural language processing. By leveraging self-attention mechanisms, models like BERT can look at entire sequences simultaneously. Trained on a simple masked language modeling objective, BERT astonishingly deduces linguistic syntax trees and deep semantic relationships from scratch, proving that the right training incentive organically builds structural understanding. Despite these leaps, modern autoregressive text generation models suffer from a fundamental logical inconsistency: they predict the next word without formulating an underlying thought, acting more as sophisticated readers disguised as writers. Moving forward, architectures like blank language models—which dynamically insert and fill blanks—offer a path toward models that actually plan their output. Unlocking this future relies heavily on modern infrastructure, such as Swift for TensorFlow and advanced LLVM compiler optimizations, enabling developers to prototype complex non-linear machine learning tasks efficiently. **Keywords:** large language models, neural network embeddings, transformer architecture, self-attention mechanisms, natural language processing, masked language modeling, bidirectional encoder representations, autoregressive text generation, blank language models, swift for tensorflow, lstm limitations, syntax tree parsing, word2vec semantic relationships, linear separability in deep learning, machine learning training objectives ## Chapters 1. **Introduction to language models and human-computer communication** (00:16) — How generating natural language resolves point-and-click friction by enabling intuitive human-computer interfaces. 1. **Limitations of syntax in natural language processing** (06:42) — How rigid syntax rules fail to capture true meaning without human cognitive interpretation. 1. **How neural networks warp space for linear separability** (09:31) — Deep learning models transform data embeddings into linearly separable clusters to classify complex information. 1. **Defining model behavior through specific training objectives** (13:05) — How mathematical incentives and loss functions strictly dictate the internal representations learned by networks. 1. **Word embeddings and recurrent neural network limitations** (16:03) — Why sequence-based architectures struggle with processing speed and contextualizing identical words with different meanings. 1. **Transformer architectures and bidirectional context in BERT** (20:33) — How the self-attention mechanism solves slow processing by analyzing entire sequences simultaneously for deep context. 1. **Building a character-level transformer for masked prediction** (24:28) — Applying self-attention and positional embeddings to a scaled-down transformer to predict masked characters in text. 1. **Analyzing learned clusters and syntax trees in BERT** (28:30) — How mathematically optimized networks independently deduce phonetic similarities and linguistic syntax structures during pre-training. 1. **The fallacy of autoregressive language generation models** (33:16) — Why relying on greedy sampling and next-word prediction fails to replicate authentic human cognitive writing. 1. **Addressing the hype surrounding creative language generation** (37:26) — Examining the logical inconsistencies in large generative models to separate artificial creativity hype from reality. 1. **Implementing blank language models for non-linear generation** (40:50) — Exploring an alternative training objective that samples random trajectories by iteratively filling dynamically inserted blanks. 1. **Accelerating machine learning research with optimized compilers** (44:27) — How compiled frameworks reduce implementation complexity to help researchers iterate on cutting-edge architectures faster. ## Related Moments - [The fundamental mechanics of large language models](https://www.wearedevelopers.com/videos/1457-exploring-llms-across-clouds) (from "Exploring LLMs across clouds") - [Capabilities and applications of large language models](https://www.wearedevelopers.com/videos/1218-data-privacy-in-llms-challenges-and-best-practices) (from "Data Privacy in LLMs: Challenges and Best Practices") - [Understanding the evolution and nature of large language models](https://www.wearedevelopers.com/videos/1010-bringing-the-power-of-ai-to-your-application) (from "Bringing the power of AI to your application.") - [Evolution of large language models and natural language processing](https://www.wearedevelopers.com/videos/347-mlops-and-ai-driven-development) (from "MLOps and AI Driven Development") - [Understanding capabilities and limitations of large language models](https://www.wearedevelopers.com/videos/631-chatgpt-create-a-presentation) (from "ChatGPT: Create a Presentation!") - [Evolution and impact of large language models](https://www.wearedevelopers.com/videos/899-creating-industry-ready-solutions-with-llm-models) (from "Creating Industry ready solutions with LLM Models") ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [WWC24 Talk - Scott Hanselman - AI: Superhero or Supervillain?](https://www.wearedevelopers.com/magazine/469-wwc24-talk-scott-hanselman-ai-superhero-or-supervillain) ## Related Jobs - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/2141919-machine-learning-engineer) at **Twilio** - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/1355348-machine-learning-engineer) at **TWILIO** - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/2023929-machine-learning-engineer) at **TWILIO** - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/2085365-machine-learning-engineer) at **TWILIO** - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/2145554-machine-learning-engineer) at **TWILIO** - [Staff, Machine Learning Engineer (L4)](https://www.wearedevelopers.com/jobs/ext/1202639-staff-machine-learning-engineer-l4) at **Twilio**