
E23 How LLMs Tick
Ryan and Luca pull back the curtain on large language models, explaining what's really happening when you chat with an AI. They break down tokenization, context windows, and the stateless nature of LLMs—revealing why these tools aren't actually thinking or remembering, just generating the most likely next token based on massive matrices of weights. The conversation covers practical implications like why long sessions deteriorate, how caching affects costs, and why that apologetic "I won't do it again" from your LLM is meaningless. They also tackle the misleading anthropomorphization in AI marketing and share frustrations with features like Claude's opaque memory system. If you've ever wondered why your LLM seems to forget instructions or why it confidently states nonsense, this episode explains the mathematical reality behind the conversational illusion. Key Topics: [02:30] Tokenization: How LLMs break language into mathematical units [08:45] LLMs as stochastic algorithms: Predicting the next most likely token [12:20] Training models: Billions of parameters and matrix multiplication [15:10] Quantization: Trading precision for memory efficiency [18:00] Context windows and token limits: Why size matters and costs money [22:15] Stateless processing: LLMs don't remember, they reprocess everything [26:40] Caching and timeouts: The hidden costs of pausing your session [30:00] Hallucinations aren't bugs—they're the default mode of operation [35:20] Attention mechanisms and why large context windows cause degradation [40:15] Claude's memory system: Well-intentioned but problematic in practice [44:30] Mixture of experts: Sparse vs. dense models and routing tokens Notable Quotes: "The LLM doesn't even know what truth is, so it can't lie to you. It's not lying to the truth either. It's bullshitting—it doesn't care which way is true or false." — Ryan "All LLMs know how to do is hallucinate. They just generate chains of tokens. If you're lucky, those chains have some connection to the real world and are actually helpful." — Luca "I wish it wasn't programmed to just lie to me. If the LLM says 'I won't do it again,' yes it will, because it has no memory of this incident. It will behave the same way tomorrow." — Luca Resources Mentioned: Luca's Training and Consulting - Luca's website with links to AI and embedded systems training courses Tulip Tree Tech Emulator - Ryan's commercial emulator for embedded systems development (Raspberry Pi, STM, ESP, Microchip) Google's Attention Paper - The foundational 2013 paper introducing the attention mechanism that enabled modern LLMs On Bullshit by Harry Frankfurt - Philosophical work distinguishing bullshitting from lying—relevant to understanding LLM outputs Agile Embedded Podcast Slack - Community discussion channel where you can reach Ryan and Luca, with a dedicated sub-channel for Embedded AI topics
