
Cracking LLMs Open: Emmanuel Ameisen Live with Tim O’Reilly
Anthropic's interpretability team studies Claude's internals the way neuroscientists study a brain, but the science is younger, messier, and far less mapped. Emmanuel Ameisen has been part of that effort for the last two years, and at O'Reilly's recent Foo Camp, he shared some of what’s happening inside of Claude when it processes information. Emmanuel joined Tim to reprise that talk, then take questions from the audience. Token prediction is often dismissed as simply “pattern matching,” Emmanuel explained, but that undersells what’s actually happening. Ask Claude to finish a sentence about a short bike ride across the “GG bridge” and it correctly infers that the bridge must be the Golden Gate, placing the user in San Francisco. From there, the model can build on this understanding to answer questions, for instance how long it will take to get to “world-class skiing” at Lake Tahoe. Emmanuel and his team study the patterns that show up when the model encounters particular ideas, and they can even intervene in those activations and see how the model’s behavior changes. Models are building a complex understanding of the world, and as Emmanuel pointed out, they “use their world model at every token.” Emmanuel and Tim discussed what Anthropic's interpretability team has learned by "pushing or pulling" on a model’s activations, what AI-generated poetry tells us about planning and world models, how models do math by placing numbers on “particularly shaped curves,” and why we should think about LLMs as actors playing a role down to the emotions they represent. They also explored what humans can do that models still can’t and the bigger question of what studying LLMs might teach us about how our own minds work.