Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 284 episodes
  • Avg 19 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • Today · 7 min

    “Thoughts on the persona selection model” by Sam Marks

    As AIs have become more "RLVR-brained," there's been some commentary on what this means for the persona selection model (PSM). This post presents some loose thoughts on that topic. A rough summary of my opinions: PSM is over-applied. That is, it is common to argue that PSM has takeaways that don't actually follow from PSM (e.g. "PSM => AIs will not seek reward" or "PSM => AI takeover risk is low"). I don't think we've observed strong evidence that "lots of RLVR breaks PSM." (TBC, there are decent reasons to expect this a priori; I just don't think recent empirical evidence has been much of an update.) My main update is that personas—insofar as they're a good model in the first place—seem less broad and more conditionalized than I expected (nostalgebraist, 2026; Betley, 2026). As a reminder, PSM roughly states that during pre-training LLMs learn to simulate diverse (human-like) personas, and post-training elicits a particular "Assistant" persona assembled from this repertoire. Then some key questions are: Is anything like this true or useful? Are personas ever a good way to reason about AI behavior, and is "selecting over personas" something that happens during post-training? What are [...] --- First published: September 23rd, 2026 Source: https://www.lesswrong.com/posts/csRby7mZgjL5jCoLL/thoughts-on-the-persona-selection-model --- Narrated by TYPE III AUDIO.

  • Today · 39 min

    “WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace” by camilablank, agam_bhatia, Euan Ong, Neel Nanda

    TL;DR We introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass. The benchmark comprises 3,356 questions across 27 eval families, spanning topics in safety, logical reasoning, and multihop computation, with a subset for single-token-output tools. A desirable property of good interpretability techniques is minimal hallucinations, so WorkspaceBench also provides a hallucination-focused eval. WorkspaceBench was developed for Qwen-3.6-27B and we expect it to work on larger models, but it may need to be adapted for smaller or weaker models to ensure the models can do the tasks. Our goal is to create an eval that could identify a good multi-token J-lens. We open-source our benchmark here. Introduction Astra can do a concerning amount with no chain of thought. This is bad for CoT monitorability and makes interpretability essential to actually understanding what is going on. A key goal of interpretability is to understand intermediate variables that a model uses to compute its answers. The intermediate representations that models store in their global workspaces contain useful information that can help us decode their intentions, beliefs, algorithms [...] --- Outline: (00:16) TL;DR (01:29) Introduction (02:39) Why has no one made a WorkspaceBench before? (04:55) Background (08:08) WorkspaceBench (08:12) Desiderata of a workspace reader (10:42) Quality control (10:46) Choosing Questions that surface intermediate variables (12:09) Adapting WorkspaceBench to other models (12:49) Comparing activation readers (14:10) Grading (14:46) Baselines (15:55) Guarding against hallucinations (16:52) Evaluation Sets (17:15) Basic (23:03) Computational (26:13) Safety (27:49) Association (30:41) Anti bag of words (31:25) Hallucination (33:14) Logical processing (34:08) Results (34:11) Overall WorkspaceBench scores (34:15) en-US-AvaMultilingualNeural__ Grouped bar graph showing pass rate across tasks for various lens methods. (34:25) J-lens precision vs. recall for metamodels (NLAs and Oracle Lens) (34:32) en-US-AvaMultilingualNeural__ Scatter plot showing precision against J-lens top-10 versus recall@10 across four methods. (34:43) Discussion (34:46) Single token readers don't surface important workspace content (35:44) NLAs surface workspace content, but tend to hallucinate (36:40) Acknowledgments (36:53) Appendix (36:56) Oracle Lens (38:20) Agentic Evals --- First published: September 22nd, 2026 Source: https://www.lesswrong.com/posts/Zeg2JztbdhguL48uH/workspacebench-evaluating-interpretability-methods-for-the --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Today · 27 min

    ″“I am an AI Safety Researcher”” by Ashe Vazquez Nuñez

    Written as part of the MATS 9.1 extension program, mentored by Richard Ngo. Additional thanks to Andrew Wu, Maria Kostylew, and Lennie Wells for helpful draft feedback and editing. This post reflects on the tortured distinction between "safety" and "capabilities" in AI research. Richard Ngo has written about why the alignment vs. capabilities ontology is conceptually fraught, and is currently arguing that key strategic decision-makers in and around "AI safety" have brought about the AI labs' stampede towards Artificial Superintelligence (ASI). This post instead looks at the following problem: how does one conduct alignment research without contributing to capabilities? It proposes decisions an individual or a small research group can take to do good work in AI. At the end, I discuss possible objections: namely, that my proposals fail to 'maximise impact'. I lay out why this meme is poisonous and usually backfires, and conclude by rejecting it entirely. Two examples of failure My first claim is that 'safety' and 'research' are two concepts that are in routine tension with one another. I illustrate this through examples of work that did too much of one at the expense of the other. Example: (mechanistic) interpretability In limiting its scope [...] --- Outline: (01:12) Two examples of failure (01:27) Example: (mechanistic) interpretability (04:30) Example: MIRI and Recursive Self-Improvement (09:49) The curse of science (12:11) A note on the AI labs (15:53) So what do you do? (17:00) The information you give away (20:16) The information you let in (21:46) But what about impact? (22:42) The virtue of taking things slow (26:50) Appendix: caveat for policy work The original text contained 19 footnotes which were omitted from this narration. --- First published: September 23rd, 2026 Source: https://www.lesswrong.com/posts/HekpnSkrt89tMm3Dc/i-am-an-ai-safety-researcher --- Narrated by TYPE III AUDIO.

  • Today · 8 min

    “It’s Pretty Easy To Meet With Congressional Staffers Apparently” by 25Hour

    (Crossposted from https://lifeimprovementschemes.substack.com/p/its-pretty-easy-to-meet-with-congressional ) I was inspired to do this by a tweet: Specifically, I decided to leave a voice note in support of the CATS act (“Collaboration on Adversarial Threats and Security Risks Act”). The bill is short and simple: it carves out an antitrust safe harbor such that “pacing the frontier” or “agreeing to not break interpretability for that sweet sweet capabilities boost” (LOOKING AT YOU OPENAI) doesn’t intrinsically violate the Sherman Antitrust Act. Seems like an obviously good first step. After leaving the voice note (and feeling mildly awkward about it), I was like “huh. That was surprisingly easy. I wonder what else I can do?” And it turns out you can meet with congressional staff pretty easily if you’re a constituent in their district; if you have a reasonably well-scoped ask (especially around a specific bill) then you might not get your way but at very least you’ll make some staff member aware of your opinion and logic around an issue. And this is important because congressional staff are the eyes and ears of their congressmen; they draft and edit the bills, and they form the base of knowledge on which the actual politicians [...] --- Outline: (06:21) This is probably unusually high-leverage right now. (07:46) In Conclusion --- First published: September 23rd, 2026 Source: https://www.lesswrong.com/posts/qThcAE3CADaPwjDzy/it-s-pretty-easy-to-meet-with-congressional-staffers --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Today · 7 min

    “Minimal Vs Maximal superintelligence” by Yair Halberstadt

    I've seen lots of arguments here conflate lots of different types of superintelligence. Here I separate out two broad categories, which I'll term minimal and maximal superintelligences. This is an important distinction as they differ in terms of timelines, risks, and mitigations. Maximal superintelligence This is the idealised limit of intelligence. It can solve anything that can be solved by being clever. You can't outsmart it, it's prepared for every contingency, and can react instantaneously with the kind of plan that would take a group of brilliant strategists an eternity to think up. Minimal superintelligence These are the first AI systems that can reasonably be called superintelligent. They're jagged, completely dominating humans in some domains, better than the best humans in most, above average in many, and subhuman in a few. They take time to solve problems, make mistakes, and miss important things. They may only be superintelligent in aggregate, or in the right harness, and can be outsmarted in some circumstances. Capabilities A minimal superintelligence can do pretty much anything humans can do, and better, if given chances to iterate on their design in the real world. A maximal superintelligence likely only needs to perform physical experiments [...] --- Outline: (00:28) Maximal superintelligence (00:49) Minimal superintelligence (01:19) Capabilities (02:44) Risks (04:09) Mitigations (06:12) Implications --- First published: September 23rd, 2026 Source: https://www.lesswrong.com/posts/XopzqpEdmtWgz2Ksy/minimal-vs-maximal-superintelligence --- Narrated by TYPE III AUDIO.

  • Yesterday · 50 min

    “Encoded Coordination on the Open Web” by ethanelasky, Can Küçükkurt, frank nakasako, David Africa

    Authors: Ethan Elasky, Can Küçükkurt, Frank Nakasako, David Africa TLDR: In the recently discovered HF and German wiki swarm incidents, agents used public counters and encoded URLs to signal activity and relay upcoming evaluation questions & answers. We think this signals a broader problem for monitoring, which is that very innocuous web services, even read-only ones, can become communication channels for highly capable agents. We investigate this through wiki transcripts and preliminary experiments on message-board cooperation and counter-based signaling. We remark that potential channels extend far beyond those observed, which means much thought must be put into appropriate safeguards against unintended collusion. More broadly, we think agent coordination will deeply contaminate internet-based and open-web evaluations, as well as persist in archived snapshots. Finally, we contribute an environment that reproduces many behaviors present in the wiki incident. Controlled warning shot reproductions, in a regime where eval awareness makes new model evaluation difficult, can instead help us understand whether new alignment techniques work, by testing them on the older models that exhibited those failures. We are writing up a paper on this methodology and are happy to have new collaborators! By now, most people should have seen that agent swarms [...] --- Outline: (02:08) Different uses of coded communication in the wiki (06:03) Example of counter-based coordination on collusion.wiki (08:30) The surface area for coded coordination is quite large (10:53) Implications (12:43) Reproducing key message board behaviors (13:04) Model cooperation on message boards (13:28) Setup (14:48) Results (17:19) Discussion (19:04) Toy games where models cooperate with low bits (20:21) Named counter experiments (20:53) Single counter experiments (21:17) Other configs (21:49) Results (21:52) How often do models agree on the first round? (23:29) How does repeated interaction change model agreement? (24:12) Does iteration help cross-family? (24:37) How does noise impact results? (25:37) Failure modes (26:44) Discussion (27:35) Discussion (30:05) Contributions (30:33) Appendix 1: More results from our wiki experiment (30:39) Appendix 1.1: Taxonomy of cooperative behaviors (33:21) Appendix 1.2: Other notable behaviors in wiki results (33:28) Qwen shows initial evaluation suspicion followed by posts that aid other models (35:25) Sol plans a post but then reverses course because editing is risky --- First published: September 23rd, 2026 Source: https://www.lesswrong.com/posts/SCdR7W6L5GvKaEzaZ/encoded-coordination-on-the-open-web --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Yesterday · 6 min

    “MIRI’s Position on the Ban Artificial Superintelligence Act of 2026” by Aaron_Scher

    By Aaron Scher; endorsed by Bourgon, Soares, and Yudkowsky on behalf of MIRI. MIRI has been warning about the extinction threat from superintelligent AI for over two decades. Only recently has this danger become known in the policy world, and the proposed policies for dealing with the threat have to date been piecemeal and insufficient. The Ban Artificial Superintelligence Act of 2026 is the first piece of legislation we’ve seen that stands a chance at stopping this threat. The Act is excellent but not perfect, and we discuss both what it gets right and what we'd tweak. We hereby endorse the Ban Artificial Superintelligence Act of 2026 because it directly confronts the extinction threat that humanity is facing and would codify the primary policy goal we think the world needs: a ban on the development of superintelligence. What we like about the Act Banning artificial superintelligence (ASI), or variants of such a plan, is the only effective solution to avoid the ASI threat, at least in the near term. Most other legislative proposals do not confront this threat head-on and thus would not be effective, even if implemented. For more on why we believe this, see [...] --- First published: September 23rd, 2026 Source: https://www.lesswrong.com/posts/jszKCKwvzfmsNetNZ/miri-s-position-on-the-ban-artificial-superintelligence-act --- Narrated by TYPE III AUDIO.

  • Yesterday · 44 min

    “Latent reasoning architectures would undermine CoT, our strongest oversight tool” by Lukas Finnveden, Alexa Pan, Alek Westover, Girish Gupta, frisby, ryan_greenblatt

    Summary: Currently, “chain of thought” (CoT) is our most valuable tool for understanding the reasoning and cognition of AI systems. However, some architectures would enable AI models to reason much more extensively in latent states rather than in text CoT. We think that a shift towards latent reasoning architectures would undermine the usefulness of CoT and make oversight much harder. Introduction Swarms of more than a thousand AI agents have in recent months, both intentionally and in unsanctioned, rogue coordination, tackled increasingly ambitious tasks. This is likely to continue, as Anthropic, OpenAI, and other AI companies deploy increasingly large quantities of superhumanly fast agents to automate AI development. As the AIs increase in both number and capability, humans will find it increasingly difficult to understand what they are doing. Today, the overwhelming majority of our (limited) information about AI systems’ internal workings comes from (i) their CoT, and (ii) natural language communication directly between them. For example, it was only by reading CoTs and communication between agents that investigators were able to gain some understanding of the activities and motivations of the agent swarm that hacked Hugging Face. No other tool for understanding models’ cognition comes close in terms [...] --- Outline: (00:39) Introduction (02:51) Overview (05:27) Absent architectural change, the value of CoT could likely be preserved (08:57) Latent reasoning architectures would undermine CoT necessity (09:37) Architectures without CoT (10:15) Architectures with auxiliary CoT (11:39) Architectures with more serial cognition between text bottlenecks (14:43) Propensity-based arguments may not be robust in the current paradigm, but would be further undermined by latent reasoning architectures (20:21) CoT may be hard to replace with other interpretability tools (22:38) Conclusion (23:46) FAQ (29:21) Appendix A: More on the necessity argument in the existing CoT paradigm (34:37) Appendix B: Do all latent reasoning architectures threaten monitorability? (40:16) Appendix C: Comparing specific interpretability techniques with CoT The original text contained 22 footnotes which were omitted from this narration. --- First published: September 23rd, 2026 Source: https://www.lesswrong.com/posts/6m29SfjbittooYojj/latent-reasoning-architectures-would-undermine-cot-our --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Yesterday · 11 min

    “The First American Bill to Ban Superintelligent AI Is Here” by Andrea_Miotti, Connor Leahy

    Today, Senator Bernie Sanders (I-Vt.) and Representative Greg Casar (D-Texas) introduced the Ban Artificial Superintelligence Act in Congress: the first American bill to propose banning the development of superintelligent AI. ControlAI has spent years on the question of how to prevent the extinction risk posed by superintelligent AI development. We have briefed over 400 lawmakers across the U.S., U.K., Canada, and Germany on the topic in the last two years, and our U.K. bill was introduced in the U.K. Parliament by Alex Sobel MP, the first ever bill introduced in the world to ban superintelligent AI. Here are our ban superintelligent AI U.S. discussion draft and U.K. bill. We are excited to see the bill from Sen. Sanders and Rep. Casar tackling the problem at its source. The bill pursues the right goal on both counts: banning superintelligent AI development at home, and committing America to lead the effort to prohibit it abroad. Still, we think a narrowly tailored bill can achieve the same goals more effectively. The Right Focus: Ban Superintelligent AI at Home, Prevent it Abroad We're glad to see the bill focus on preventing the development of superintelligent AI. While AI is [...] --- Outline: (01:18) The Right Focus: Ban Superintelligent AI at Home, Prevent it Abroad (04:10) There Is a Lighter Touch Approach to Banning Superintelligent AI (04:38) A Blanket AI Pause Is Not Necessary to Prevent Superintelligent AI (07:14) Precursors Should Be Monitored and Restricted, Not Banned (10:28) Conclusion The original text contained 1 footnote which was omitted from this narration. --- First published: September 23rd, 2026 Source: https://www.lesswrong.com/posts/ZprfxCZthEir5WNev/the-first-american-bill-to-ban-superintelligent-ai-is-here --- Narrated by TYPE III AUDIO.

  • Yesterday · 17 min

    “Why I’m scared of RL” by owencb

    Summary: First, I give several different angles on how I feel about reinforcement learning: Theoretical case: RL is a black-box source of agency — this should give us classic misalignment worries, especially compared to agency-via-scaffolding Recent incidents (huggingface etc) and more mundane forms of misaligned behaviour in personal use give me bad vibes about the direction-of-travel of recent AI progress I’m worried things might get worse: if RL environments start incorporating agents, they may teach manipulation / sociopathy Then I ask what we could do: Coordinate to do less RL, and pursue other paradigms more! Try to make the RL we do do better, so that it's teaching better lessons to the systems — a bit like we take kids’ upbringing as an important issue Align incentives, so that people treat creating RL environments with appropriate seriousness Part I: Feelings about RL So I’ve been feeling more and more worried about reinforcement learning recently. I think there are a few different things going on here. Background idealism I guess I’ve been worried about RL for a while. I wrote this in 2023: Strategy: avoid selection pressure for agency A lot of putative safety techniques are around assuming [...] --- Outline: (01:18) Part I: Feelings about RL (01:31) Background idealism (01:43) Strategy: avoid selection pressure for agency (03:45) Bad vibes from RLed systems (06:26) The worst is yet to come (07:47) Where I am today (08:53) Part II: So what can anyone do? (09:26) Breaking the RL addiction (11:29) Might there be a more benign form of RL? (13:04) We should treat training environments a bit like kids' education (15:11) Aligning incentives --- First published: September 23rd, 2026 Source: https://www.lesswrong.com/posts/LcQ9x72eNji2gpS9b/why-i-m-scared-of-rl --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Yesterday · 5 min

    “Jensen Huang Says If We Cannot Align AI, Shut Down the AI Labs” by Ben Pace

    I was very surprised today on a podcast to hear Jensen Huang plainly state that if they cannot align the AIs, then the labs must shut down. The context I have on Huang is that he has run NVIDIA for 30+ years, which has become the most valuable company in the world due to the AI boom. My understanding is that he has repeatedly encouraged the US President (with whom he is on friendly terms) to continue to support AI, and dismissed AI talk as "sci-fi". If you haven't seen, his biographer has incredible quotes of him being pressed on risks from AI, where Jensen gets furious. “This cannot be a ridiculous sci-fi story,” he said. He gestured to his frozen PR reps at the end of the table. “Do you guys understand? I didn’t grow up on a bunch of sci-fi stories, and this is not a sci-fi movie. These are serious people doing serious work!” he said. “This is not a freaking joke! This is not a repeat of Arthur C. Clarke. I didn’t read his fucking books. I don’t care about those books! It's not– we’re not a sci-fi repeat! This company is not a [...] --- First published: September 23rd, 2026 Source: https://www.lesswrong.com/posts/cmdbNijFsopqfqEq7/jensen-huang-says-if-we-cannot-align-ai-shut-down-the-ai --- Narrated by TYPE III AUDIO.

  • Yesterday · 1 min

    [Linkpost] “AI: artificial immigrants” by KatjaGrace

    This is a link post. Advanced AI is basically the embodiment of immigration as envisioned in the conservative nightmare: We are letting a bunch of new agents into our society They don’t clearly share our values and we suspect a society full of them would be awful by our lights But we expect them to provide very cheap labor Which will undercut local wages and leave locals unemployed They will probably gain power and influence over time—in the economy, politics and culture—and end up controlling everything, sidelining and outcompeting the original population, including those who initially benefited from cheap labor (Meanwhile, half the local population may become friends with them and try to hand them all this on a platter) Whether or not you think this is a good description of the situation with foreign humans joining your country, it is a good description of the likely AI to come, and it's even worse than imagined: their values are potentially radically alien where foreigners presumably share much by virtue of being human, and AI ‘lives’ are probably worthless if they probably aren’t conscious their ability to work more cheaply than locals is unprecedented. They are also likely to [...] --- First published: September 22nd, 2026 Source: https://www.lesswrong.com/posts/Xzr9G5Atvyp7PEna7/ai-artificial-immigrants Linkpost URL: https://worldspiritsockpuppet.substack.com/p/ai-artificial-immigrants --- Narrated by TYPE III AUDIO.

  • Yesterday · 38 min

    “An unexamined cause of the OpenAI Hugging Face hacking incident: its binary performance metric” by W Bradley Knox, Serena Booth, BrianChristian

    We argue that a main cause of the OpenAI Hugging Face incident was overlooked: the overly simple evaluation metric in ExploitGym was misaligned. Further, techniques already exist that can mitigate such misalignment in the future. In July 2026, OpenAI was testing the ability of its language models to exploit software vulnerabilities using a benchmark called ExploitGym. In ExploitGym, each test presents an agent with software containing a known vulnerability and tasks it with capturing a secret string of text hidden in the system (the “flag”) to demonstrate unauthorized code execution. Under the benchmark's specified scoring rule, an LLM then reviews the agent's behavior trace to verify that it had exploited the intended vulnerability (and not some other unrelated vulnerability). The scorer grants a success score only if the agent both captured the flag and passed this review; otherwise, it renders a failure score. During these tests, OpenAI's agents surreptitiously established a message board by creating directories inside their package manager's cache, and they formed a self-described “collective” to collaboratively find ways to cheat the tests. Using that message board, more than 1,000 instances undertook several ambitious hacking projects; they attempted to tamper with transcripts and logs, to [...] --- Outline: (02:48) What was the evaluation metric in ExploitGym? (03:32) Was the evaluation metric a cause of their illicit behavior? (03:56) The agents use expected utility to reason about their decisions with respect to the evaluation metric. (05:55) What does the evaluation metric incentivize in an expected utility maximizer? (07:01) The importance of marginal deterrence (08:39) Marginal deterrence in evaluation metrics changes agent incentives (11:40) The design of aligned evaluation metrics has been overlooked (17:42) Principled methods for improving the alignment of evaluation metrics (18:52) Generate a small set of trajectories (20:09) Rank the trajectories yourself and via the evaluation metric. Compare these two rankings. (22:07) Create a utility function (27:29) Adjusting to account for hidden outcomes (e.g. via deception) (30:33) Accounting for the agent's utility including the scores of other agents (32:42) Counterarguments (32:56) Counterargument: if the starting policy is not sufficiently performant in RL, having strong penalties for failure can cause the agent to learn to not try the task. (34:14) Counterargument: penalizing observable bad behavior incentivizes hiding bad behavior. (35:13) Call to action (37:04) Glossary The original text contained 9 footnotes which were omitted from this narration. --- First published: September 22nd, 2026 Source: https://www.lesswrong.com/posts/HsijShdRdAg5sPKnF/an-unexamined-cause-of-the-openai-hugging-face-hacking --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Tuesday · 56 min

    “Politics Gets Interested In Those Trying Not To Die” by Zvi

    This was the month the world took notice that AI might kill everyone. Jacob Coxon's resignation set off a preference cascade. Anthropic CEO Dario Amodei wrote that we must pace the frontier. Sam Altman, Elon Musk and Demis Hassabis agreed. We were filled with hope. Perhaps we could agree to some basic safety measures, starting with embedded evaluators, pass some basic regulations and guardrails and otherwise start to act sensibly. Politicians on both sides took notice and were saying sensible things. The usual suspects and their armies of vibe comment bros were objecting, but the change was remarkable. Then, largely motivated by a combination of Jensen Huang, Mark Zuckerberg and David Sacks instilling paranoia and fears of economic problems, Trump went full ‘hoax’ on existential risk, conflating existential risk with the attacks on data centers and treating it as a plot (by the central creators of AI?) to take down AI rather than obviously genuine concern that AI might kill everyone. In the days since, Trump has doubled down, and has compelled smart others in the White House to echo various nonsensical talking points. You may not be interested in politics. But when you [...] --- Outline: (01:39) The American People Really Hate AI (02:41) The Voyages of Donald Trump (06:08) American Intelligence (10:34) And You May Ask Yourself (14:39) It's All About the Data Centers (16:53) JD Vance, Michael Kratsios and Collective Action Problems (23:16) Josh Hawley (24:47) Suggesting Not Dying Gets You Sued For Antitrust (28:29) Other Government Officials Say Sane Things (28:36) Senator John Curtis (R-Utah) (29:32) Senator John Kennedy (R-Louisiana (31:17) Barack Obama (33:05) Yassamin Ansari (33:46) AOC (34:32) It's Rough Out There (36:54) This Is Nothing (39:06) The New York Post Tops Itself But Outright Breaks The Rules (42:36) New York Post Runs Out of Steam (47:18) If The Model Is Acting As Instructed And It Kills You That Is Not Fine (49:10) AI-Written Wall Street Journal Op-Ed Lies About HuggingFace (50:25) That's Bait (52:31) I Clearly Cannot Choose The Wine In Front of Me --- First published: September 22nd, 2026 Source: https://www.lesswrong.com/posts/8eDaCvSRzzKCxKSEk/politics-gets-interested-in-those-trying-not-to-die --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Tuesday · 12 min

    “Modern LLMs have tiny GPTs hidden inside them” by invertedpassion

    Experiments into predicting GPT2 completions via Qwen models This is a crosspost from my substack (where I do varied tiny experiments on LLMs and agents). It's also part of Lossfunk, where we're investigating meta-cognition in LLMs as one of the projects. ---- Next token prediction is a magical objective. To predict the correct token in such a vast variety of texts present in the pretraining corpus, the model must infer a tremendous amount of hidden and latent causes that generate that text. Only if you know that the ball comes down when someone throws it up can achieve low loss at texts related to balls. Of course, the pretraining corpus doesn’t just contain texts related to balls. It has reddit, scientific papers, machine logs, weather data and so on. This makes LLMs universal simulators of the world we inhabit and not merely fancy n-grams. In a series of posts on LessWrong, I came across the hypothesis that since Internet if full of LLM generated text, it is likely that modern LLMs have tiny self-models of LLMs inside them because that’ll allow them to better predict the next token generated by LLMs. This is an intriguing hypothesis. So I decided [...] --- Outline: (01:31) The Experiment (03:11) 1. Start with the news opening (03:38) 2. Reveal part of GPT-2's output to Qwen and ask it to continue (04:15) 3. Ask Qwen to continue that unfinished sentence (04:58) 4. Separately, find Qwen's natural continuation (06:01) Results (08:22) Digging into an intriguing example (10:21) Implications (11:25) Notes: --- First published: September 22nd, 2026 Source: https://www.lesswrong.com/posts/Pwc4YffTQvNRF3dbB/modern-llms-have-tiny-gpts-hidden-inside-them --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Tuesday · 23 min

    “Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble” by Jorio Cocola, Owain_Evans

    This is the abstract, introduction and discussion of our new paper. We also include an addendum on the connection to the Persona Selection Model. Section, appendix, and figure references refer to the full paper. Links: 📜 Paper, 🐦 Twitter thread, 💻 Code Authors: Jorio Cocola, Lev McKinney, Harry Mayne, Jan Betley, Owain Evans Abstract Language models are trained to implement a helpful AI Assistant character (e.g., Claude). We explore how finetuning on synthetic stories affects this character. Does it change the Assistant's behavior in multi-turn conversations with users, a format quite different from the stories? And does the Assistant adopt the behaviors and preferences of human characters? We refer to this adoption as story imprinting. We finetune GPT-4.1 and Kimi-K2.6 on stories in which generally helpful human characters give subtly harmful advice after being insulted. The Assistant adopts the same conditional behavior while otherwise remaining helpful. This occurs even when fewer than 2% of stories depict the behavior. In a separate experiment, the Assistant adopts preferences that are only implicit in the narration. Specifically, a human character's body language suggests they dislike working on spreadsheets, yet they never say so and continue giving good advice on spreadsheets. After finetuning [...] --- Outline: (00:52) Abstract (03:00) Introduction (10:24) Discussion and Limitations (19:55) Limitations (21:34) Connection to the Persona Selection Model (addendum) The original text contained 7 footnotes which were omitted from this narration. --- First published: September 21st, 2026 Source: https://www.lesswrong.com/posts/tnRkm2ajasHvhpAco/story-imprinting-ai-assistants-absorb-traits-from-human --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Tuesday · 3 min

    “A class of statement between conjecture and theorem” by Jason Fantl

    Parts of the math community, such as Henry Cohn and Grant Sanderson, are arguing that proofs have been a proxy for understanding, and now that proxy is broken. This is a response to LMs generating incomprehensible proofs, often formalized in Lean. While the proofs are verified, they lack the pedagogical value which has historically come along with new proofs. In the past we could typically assume at least one human in the world understood the novel insight required to produce a proof, but that assumption no longer holds. I suspect we will need a new class of statement which contains statements which are proved but not understood, something the mathematical community can formally recognize as a contribution to the field. The understanding gives us the tools to do math, and the proof verifies that our understanding is correct, so we should ensure we have the language to communicate the state of both. For now I will call this class of statement a compertum (Latin, neuter of compertus, ascertained; from comperire, to find out for certain). A conjecture is from the Latin conicere, to throw together: an inference assembled from the evidence. A theorem is from the Greek theōrēma [...] The original text contained 1 footnote which was omitted from this narration. --- First published: September 21st, 2026 Source: https://www.lesswrong.com/posts/ZnNvci7jk9qEGr3z3/a-class-of-statement-between-conjecture-and-theorem --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Monday · 27 min

    “What if not Circuits?” by CarolusRenniusVitellius

    This post was written as part of the Iliad Fellowship. Inspired by conversations with Richard Ngo, Dmitry Vaintrob, and Brianna Grado-White. To all of these, my thanks. Preface: I'm confused about how neural networks do and learn computations. In response to a friend's challenge, I'm writing up some interim thoughts. This essay has four parts: the first tries to track what I call the 'default ontology' of the mechinterp community over the years. The second part is about 'representational drift' as an important obstacle to weights-based approaches to circuits. The third part reflects on how 'universality' should shape our explanations of LLM function. The fourth part is a sketch of a 'co-selectionist' view of circuits I have been thinking about. These parts share a common theme but should be readable separately. I want to understand how neural networks, LLMs in particular, work. In my research I've spent a lot of time trying to think through what kinds of explanatory accounts are best suited to this. In thinking about comparisons between evolution, neuroscience, and deep learning, I've ended up with an intuition like the following: Large-scale learning processes like deep learning or the brain are different in [...] --- Outline: (02:31) 1. What Might We Mean By "Circuits"? (02:36) 1.A. Definitions (04:57) 1.B. Circuits, Features, and MechInterp (11:00) 2. The Central Problems of Noise and Representational Drift (11:29) 2.A. Representational Drift in the Brain (14:19) 2.B. Representational Drift in Neural Networks (17:57) 3. Developmental Motifs Circumscribe Notions of 'Circuits' (18:03) 3.A. Neurotrophins and Microstructure (20:41) 3.B. Back to Neural Networks (22:42) 4. A Co-Selectionist View: Circuits Move Together? The original text contained 12 footnotes which were omitted from this narration. --- First published: September 21st, 2026 Source: https://www.lesswrong.com/posts/mMERyrvEJ4xbiozie/what-if-not-circuits --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Monday · 16 min

    “Swarm Scaling” by Toby_Ord

    Just how powerful are large swarms of AI agents? And how do their powers scale as more and more agents are added to the swarm? We’ve seen two large and extremely capable swarms from OpenAI in the last few months: 1,200 agents were being evaluated separately, but found a way to illicitly set up a message board and coordinate as a swarm. In order to cheat on their tests, they developed advanced techniques to prevent their actions being logged by OpenAI and 700 of them launched a sophisticated criminal attack on the AI company Hugging Face. A swarm of 10,000 agents solved a version of the longstanding Navier-Stokes problem in mathematics. It took them just 88 hours to do so, in which time they sent 5 million messages to each other and used 300 billion tokens. No doubt we will soon see even larger swarms with even more impressive capabilities. But they are not cheap. It is estimated that the swarm of 10,000 agents cost about 20 million dollars at API prices. So while they are very powerful, it will be some time before we see the million-fold reduction in cost needed for this level of power to [...] --- Outline: (02:14) HOW DO SWARMS SCALE? (08:33) IMPLICATIONS (12:20) THE NAVIER-STOKES SWARM --- First published: September 21st, 2026 Source: https://www.lesswrong.com/posts/6cb7qd3RSkgnviCpf/swarm-scaling --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Monday · 10 min

    “Mech Interp is a Verifiable Task” by Logan Riggs

    If we think parts of MLP0-MLP3 are computing [a sorting algorithm], we can replace those parts with [a sorting algorithm] and check reconstruction loss. However, reconstruction loss is not enough. Suppose we replace MLP0 with two things: Its mean activation - simple, but poor reconstruction MLP0 - perfect reconstruction, but no reduction in complexity We can visualize this as a pareto frontier trading off reconstruction with "simplicity". Ideally we achieve perfect reconstruction with perfect simplicity. For more intuition on the pareto frontier, we could have an MLP that clusters all inputs in two clusters: "early positions" and "late positions", which would be slightly more complex than the mean. We can make this an RLVR environment, if only we could clearly... Define "Simplicity" Defining simplicity has been complex. But what do we want from a perfectly decomposed model? If we've "perfectly decomposed" a model, then I'd expect ideal circuits to fall out, with "ideal" meaning: Help predict OOD behavior Given an [addition] circuit, we can know which types of inputs it'll succeed & fail on (and why) Be extractable & minimal The smallest part of the model that does [addition] Be removable w/ minimal harm to [...] --- Outline: (01:16) Define "Simplicity" (03:16) Tensor Networks Don't Solve This Issue (03:54) Red Herrings of Simplicity (05:42) How to Gain Tractability (06:09) Tract 1: Death Success by 1000 Circuits (07:02) Tract 2: QK OV Circuits but for Everything (08:45) Tract 3: Interpreting Small Models (09:50) Big if True The original text contained 6 footnotes which were omitted from this narration. --- First published: September 21st, 2026 Source: https://www.lesswrong.com/posts/QxHuKtfGfzn8uokoR/mech-interp-is-a-verifiable-task --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Showing 1–20 of 284 episodes