> Markdown version of [/events/world-congress-2026-north-america/sessions/1669-understanding-llm](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1669-understanding-llm). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Understanding LLM Architectures: Inside the Design of Modern Models - **Event:** World Congress 2026 North America ## Description Large Language Models are often described as if each generation introduces an entirely new architecture. In practice, most modern LLMs still retain the transformer core, but their real progress comes from a series of targeted design changes around attention, positional handling, feed-forward computation, routing, and memory efficiency. This talk explains LLM architectures through the engineering tradeoffs that shaped modern models: why some attention mechanisms evolved for lower inference cost, how architectural choices influence long-context behavior, why sparse activation changed the economics of scale, and how these shifts affect real-world deployment. Rather than treating LLM architecture as a static diagram, this session presents it as a set of design decisions made in response to practical constraints in latency, memory, context length, and system efficiency. Attendees will leave with a clearer mental model of what remained stable, what changed, and why those changes matter. ## Speaker ### [Jofia Jose Prakash](https://www.wearedevelopers.com/@jofia-jose-prakash) Enterprise AI Architect at American Chemical Society ## Related talks at this congress - [Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1415-fast-cheap-and) — Legare Kerrison, Cedric Clyburn - [No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1671-no-single-model-to) — Emmanuel Acheampong - [You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1677-you-can-t-re-run) — An Phan - [Headroom: A Context Optimization Layer for LLM Applications](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1419-headroom-a-context) — Tejas Chopra ## Watch remotely Can’t make it to San José? Watch this session live with Pro. You also get: - All full videos, bookmarks, and playlists - World Congress livestreams [See pricing](https://www.wearedevelopers.com/pricing) ## Links - [Get tickets](https://www.wearedevelopers.com/world-congress-north-america/tickets)