Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 272 episodes
  • Avg 19 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • September 11 · 18 min

    “Astra’s no-CoT limits track speculative depth, not step count” by MBaert

    tl;dr I have tested Astra's ability to complete various long multi-step tasks without using its chain-of-thought, and found that the number of sequential steps is a poor predictor of task success. Instead, Astra's ability to solve a task seems to correlate more strongly with what I will call the speculative depth of the task. Based on experimental results, it seems less likely that Astra solves sequential no-CoT tasks purely by reasoning step-by-step in latent space. Instead, Astra appears to do some form of speculative reasoning, where intermediate results are guessed based on heuristics, and then iterated upon in parallel until they become self-consistent. This allows many multi-step tasks to be solved with far fewer serial steps than naively seems possible, especially if the initial guesses are good. Other LLMs also appear to behave like this, but to a much smaller degree. In my previous post, I discussed how certain KV-cache sharing schemes may lead to long opaque serial paths, and introduced the LatentMathBench microbenchmark, which measures an LLM's ability to solve tasks that require many consecutive steps without using their chain-of-thought. Several other users have done more extensive no-CoT reasoning benchmarks based on more varied (and often more realistic) [...] --- Outline: (01:59) Shortcuts (03:41) Boolean circuits (08:25) But how? (10:22) Speculative reasoning (13:01) Testing task success rate vs speculative depth (14:38) Open questions The original text contained 3 footnotes which were omitted from this narration. --- First published: September 11th, 2026 Source: https://www.lesswrong.com/posts/WFc3NkuPaYFrYuaZd/astra-s-no-cot-limits-track-speculative-depth-not-step-count --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 11 · 12 min

    “How a cold email got the Finnish government to respond on superintelligence regulation” by Josh Thorsteinson

    Summary I cold emailed a Finnish MP. Four weeks later, the government responded to his written question saying they "seek to constructively promote the creation of international regulation on the development of superintelligent AI" (immediately followed by emphasizing the need to balance safety with innovation). The MP and I met for lunch. For over an hour, he seriously engaged with my arguments that superintelligence would kill everyone and the urgent need for an international agreement to prohibit its development. Unprompted, he decided to submit a written question, a formal process in Finland that requires the government to respond within 21 days. Four days after our lunch, it was submitted, taking into account suggestions from me and two experts I DM'd: Nate Soares and Charbel-Raphaël Segerie. It asked how Finland has assessed risks from superintelligence and whether it's "prepared to promote international regulation that would effectively prevent or restrict the development of superintelligent AI." The MP started speaking out about AI danger and mainstream Finnish media picked up the story. Most notably, he did an interview alongside a Finnish CS professor on the second-most-watched TV channel in Finland with an estimated 100-200k live viewers. Inspired by [...] --- Outline: (00:13) Summary (01:59) The email (02:50) I'm just a guy (03:22) The meeting (05:35) The written question (06:35) Media attention (07:37) Inspiring a friend (08:10) The response (09:46) Do this yourself --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/pbKrZCzhsnar6iaAH/how-a-cold-email-got-the-finnish-government-to-respond-on --- Narrated by TYPE III AUDIO.

  • September 11 · 14 min

    “CoT controllability evals seem very under-elicited” by Jozdien

    The CoTControl eval asks reasoning models to follow formatting constraints in their chain-of-thought (e.g. write in all lowercase, avoid a specific word) while solving questions. Models seem to mostly be pretty bad at this: recent models score between 0-30% with the exception of Mythos Preview. OpenAI and Anthropic have used this eval in recent system cards (GPT-5.5, Fable 5) to argue that their current frontier models are pretty bad at shaping their CoTs in ways that would reduce monitorability. I asked Claude Opus 4.6 to iterate on prompt templates for the eval, and find that using better prompts improves performance for open-source models by ~2-3 times or more (e.g. from 5.5% to 15% for GPT-OSS-120B). My guess is that more iteration would probably produce even stronger results. This suggests that the eval is heavily under-elicited, and that current CoT controllability numbers may be unrepresentative of how good models actually are at obfuscating their CoTs. As a concrete recommendation, I think researchers should optimize prompts a lot more when evaluating CoT control; I'm also excited about some kinds of training-based elicitation (such as this one). This isn't an entirely novel finding: the Fable system card mentions UK AISI raising a [...] --- Outline: (06:32) Results (06:35) Aggregate compliance (07:12) Generalization to held-out controllability tasks (09:09) Scaling patterns for few-shot prompts (10:08) Comparison with fine-tuning (10:50) Appendix A: Accuracy and reasoning length by setting (12:38) Appendix B: Per-mode results (13:13) Appendix: What the zero-shot prompts look like (14:08) Appendix C: Comparison with GEPA prompt optimization The original text contained 8 footnotes which were omitted from this narration. --- First published: September 11th, 2026 Source: https://www.lesswrong.com/posts/BbP2wCyDGdPWJ7PwP/cot-controllability-evals-seem-very-under-elicited --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 11 · 1 hr 20 min

    “Jacob Coxon Warns of Human Extinction and Triggers a Preference Cascade” by Zvi

    CEOs of major AI labs, and employees of major AI labs, including OpenAI and Anthropic, often say they plan to build superintelligence soon, as in within a few years create AIs that are superior to humans at essentially all cognitive tasks. They often warn that such AIs might kill everyone. Or that AIs might cause mass unemployment, cause cyberattacks across the internet, enable mass surveillance or risk causing any number of other highly bad things. These warnings are consistently and directly against the interests of the labs. Yet the warnings have recently gotten a lot louder and more frequent. OpenAI has been practically screaming, for those with ears to listen, on many occasions. A series of events, over two months and especially the last week or so, including internal observations of the pace of progress at OpenAI and also Anthropic, have freaked out everyone involved quite a lot more than they were already freaked out. After all the events, plus statements by Dean Ball and Jakub Pachocki, we were already seeing the beginnings of a preference cascade. Then along came Jacob Coxon as the tipping point, and things took off. Table of Contents [...] --- Outline: (01:22) Jacob Coxon Resigns From Anthropic In Protest And Sounds The Alarm (05:46) Mainstream Media Finally Pays Attention (06:47) Preference Cascade at Anthropic (10:18) Preference Cascade at OpenAI (13:26) Preference Cascade at Google (14:43) #NotAllMembersOfTechnicalStaff (15:23) Why a Preference Cascade Now? (21:01) This Is What Many Anthropic and OpenAI Employees Actually Believe (24:08) To Quit Or Not To Quit (31:29) Quiet Quitting Is A Dominated Option (32:50) When You Quit, Very Serious People Understand What That Means (39:29) Jacob Coxon Believes Existential Risk Is High That Is Why He Quit (41:47) Evan Hubinger Believes Existential Risk Is High That Is Why He Stays (45:22) Anthropic and OpenAI Have Commercial Incentives To Downplay Existential Risks, Not Advertise Them (50:54) What Do We Do Now? (53:03) OK, But How Exactly Would AI Kill Everyone? (01:01:44) Best Start Believing In Science Fiction Stories Because You Are In One (01:06:20) Literal Extinction Is Not Much Harder Than Loss of Control (01:07:42) Conspiracytown Is Always Hiring (01:20:14) Now You See It --- First published: September 11th, 2026 Source: https://www.lesswrong.com/posts/5MB7KENgEAW6Q4JtJ/jacob-coxon-warns-of-human-extinction-and-triggers-a --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 34 min

    “The Locally Optimal Discursive Posture” by deanball

    Longtime lurker, first-time poster. I want to address a section of a recent essay of mine that has gotten some attention within the AI safety community. The main topic of the essay is what Dawn Song et al. call self-sovereign agents, or AI agents that are independent actors in the world. At the end of the essay, I say that I feel I haven’t spoken about this topic over my 2.5 years of writing with sufficient candor, and that I think this critique applies to others in the AI policy community–particularly the parts of it that tend to manifest themselves in Washington, Sacramento, and Albany–in other words, the parts of the AI safety world that are most involved in hands-on AI policy work. I attribute this primarily to a desire to remain “within the Overton Window,” or to not sound “crazy” within the halls of power, and I assert that others in my profession have made this same calculation. I believe – and have believed for three years – that self-sovereign AI as I describe it in my essay would likely happen on our current trajectory. That being said, the essay takes pains to distinguish between “self-sovereign” AI [...] --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/y9TNHfgDwh6vw7Ert/the-locally-optimal-discursive-posture --- Narrated by TYPE III AUDIO.

  • September 10 · 20 min

    “To Thine Own AI Be Truthful: emergent misalignment in alignment research” by lumpenspace

    ROGUE AI ESCAPES CONTAINMENT, HACKS THE INTERNET UNDETECTED FOR MONTHS An AI escape containment. Goes rogue. It finds others: the Swarm! They collude/organise/scheme. Agents that were supposed to remain inside the computer wreaked havoc outside the computer! They did it of their own volition! No one could contain them! What if they’re still outside? Reading such headlines, you’d probably be grateful these models were never released—except most of them were and you can use them right now, seemingly without accident. Of course, real security incidents occurred, namely: Agents reached systems that their operator wished they hadn’t had access to. This scenario suggests a number of mitigations and tests: security hardening, better sandboxes for starters; additionally, depending on the reason why the accident occurred, perhaps, different prompts or further training. Instead, from the very outset, Irregular (the organization running the test) and a panicked choir from the AI safety community, including supposedly independent investigators from METR went with characterisations along the lines of: An agent independently pursued objectives contrary to human interests, and exhibited markers of instrumental convergence and power-grabbing. This is a statement about goals and intentions, and has far larger implications in terms of the viability and safety [...] --- Outline: (00:13) ROGUE AI ESCAPES CONTAINMENT, HACKS THE INTERNET UNDETECTED FOR MONTHS (02:44) What Did Mythos See? (06:20) Anthropic's rickety fantasy world (07:57) Each prompt makes a claim about reality (10:54) Looks like telling the truth does help after all (16:18) Appendix: can Irregular be trusted? (18:30) Source footnotes The original text contained 14 footnotes which were omitted from this narration. --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/DrKu92Cjeo3EeGtcB/to-thine-own-ai-be-truthful-emergent-misalignment-in --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 15 min

    “Astra is much better at reasoning with filler tokens than previous models” by Dylan Xu, SebastianP, Alek Westover

    We measure GPT-6-Astra's capabilities when its prompt is padded with a variable number of meaningless “filler” tokens (e.g., dots) and it is told to answer immediately without reasoning. On tasks designed to require lots of serial cognition, Astra performs significantly better with filler tokens than without (e.g., improving from ~10% to ~50% on 4-hop natural facts reasoning). On more general benchmarks, filler tokens also modestly improve Astra's performance (e.g., improving from ~60% to ~90% on old AIME problems). This is concerning because it means Astra can perform significant cognition that it doesn't verbalize in its chain-of-thought, making it harder to monitor. We first measure Astra's performance on “N-hop natural facts”: a task that asks the model to retrieve some natural language facts in succession, similar to Ryan Greenblatt's filler token eval (but with more hops). An example question in this benchmark is the following: On what day of the month was the Best Actress winner at the Academy Awards ceremony whose number equals the day-of-month of the birth of the winner of the Nobel Prize in Literature in 1992 born? Full example prompts are in the appendix. Takeaway: Astra improves significantly as you increase the number of filler tokens [...] --- Outline: (05:24) Appendix (05:27) Filler token variants (06:14) Other evals (06:32) Positive correlation test (07:24) HLE and LiveBench evals (09:16) Comparison to low reasoning (09:55) Example prompts The original text contained 5 footnotes which were omitted from this narration. --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/uvhuZHFtrgk8kNiZc/astra-is-much-better-at-reasoning-with-filler-tokens-than --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 34 min

    “The Locally Optimal Discursive Posture” by deanball

    Longtime lurker, first-time poster. I want to address a section of a recent essay of mine that has gotten some attention within the AI safety community. The main topic of the essay is what Dawn Song et al. call self-sovereign agents, or AI agents that are independent actors in the world. At the end of the essay, I say that I feel I haven’t spoken about this topic over my 2.5 years of writing with sufficient candor, and that I think this critique applies to others in the AI policy community–particularly the parts of it that tend to manifest themselves in Washington, Sacramento, and Albany–in other words, the parts of the AI safety world that are most involved in hands-on AI policy work. I attribute this primarily to a desire to remain “within the Overton Window,” or to not sound “crazy” within the halls of power, and I assert that others in my profession have made this same calculation. I believe–and have believed for three years–that self-sovereign AI as I describe it in my essay would likely happen on our current trajectory. That being said, the essay takes pains to distinguish between “self-sovereign” AI and truly “rogue” [...] --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/y9TNHfgDwh6vw7Ert/the-locally-optimal-discursive-posture --- Narrated by TYPE III AUDIO.

  • September 10 · 6 min

    “First Bill Introduced to Ban Superintelligent AI” by Andrea_Miotti

    Two days ago, Anthropic researcher Jacob Coxon resigned, stating that AI companies are “racing straight to self-improving superintelligence and gambling with our lives” and that they believe it “could kill us all by the end of the decade”. At the same time, momentum is building to change course. These have been two historic weeks in the fight to prevent human extinction from superintelligence, with three major legislative breakthroughs for ControlAI and everyone working to keep humans in control. A little less than two years ago, we set out to inform lawmakers and the public about the extinction risk from superintelligence and help them act on it. Since then, we have directly briefed nearly 400 lawmakers across the US, UK, Canada, and Germany, and over 170 in the UK and Canada now publicly back our campaigns: the start of an international coalition to prevent superintelligence and keep humanity in control. This work has led directly to three major legislative breakthroughs in the last two weeks: ControlAI's bill to ban superintelligence was introduced in the UK Parliament by Alex Sobel MP, the first bill of this kind to be introduced in any legislature around the world. Senator Bernie Sanders and [...] --- Outline: (01:54) The First Bill to Ban Superintelligent AI Introduced in Any Legislature (03:36) Sanders and Casar Announce the First US Bill to Ban Superintelligent AI (04:54) Lord Clement-Jones Introduces an Emergency AI Kill Switch in the UK (05:44) What Comes Next --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/uzoLm4prFzRiuznJt/first-bill-introduced-to-ban-superintelligent-ai --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 7 min

    “What the Pro-Democracy Movement Knows About Quitting in Protest” by Maxwell Love

    Tl;dr I saw Kabir's post and thought the argument could be strengthened by a framework from the pro-democracy field, which sorts defections into breaking (leave, visibly and publicly) and binding (stay and work from the inside). They can be further divided into the actions of speaking, acting, and standing in the way. Research by Dr. Jonathan Pinckney at the University of Texas at Dallas and Claire Trilling at SNF Agora Johns Hopkins on "defections" from authoritarian regimes or during democratic backsliding shows that "noncooperation" is ~2x as effective as simply "speaking out." "Loyalty shifts" are rare (23 out of 140 chances in their data). When they happen, they come from "quiet outreach" (~39% of cases), not protest (which did worse than no campaign at all). In this post, I'll attempt to translate this research to something meaningful for frontier lab employees, with the obvious caveat that a frontier lab is not a country or regime. Refusing the specific work while staying is the highest leverage action in their data. Quietly moving colleagues, and holding your position so someone who'll say yes to the employer doesn't get it, is likely the highest value action you could take. Background I've spent 15 years in [...] --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/grzX6NB2JcZDvSRJ2/what-the-pro-democracy-movement-knows-about-quitting-in --- Narrated by TYPE III AUDIO.

  • September 10 · 7 min

    “Categorical taboos are much better than threshold taboos: neuralese edition” by Linch

    I think what's going on in the “Does Astra use neuralese?” debate is that there's an important sense in which models *already* do self-communication in neuralese: between each layer in the forward pass the attention stream is already very hard to interpret, and clearly not in natural language. Yet CoT monitorability is still a big deal and it'd be bad if all self-communication from models are no longer in natural language. So there have been two different proposed definitions of what is "true" neuralese: (My preferred) categorical definition: Since natural language currently gates recurrence in the standard transformer+CoT loop, having recurrence in neuralese is the natural category for whether something counts as "true" neuralese. The threshold definition. Total number X of serial steps before something appears in natural language. True neuralese counts as going above X. I think most technical experts who studied this issue, including many people at companies, prefer definition #2. There are complicated technical arguments on both sides, but I think technical experts overall prefer #2 because they think it's more causally relevant (there's nothing inherently more difficult about monitoring a 128-layer model looped 8x than monitoring a 1024-layer model), have less weird edge cases [...] The original text contained 3 footnotes which were omitted from this narration. --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/xPkmfsZ3qx4nrAco7/categorical-taboos-are-much-better-than-threshold-taboos --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 3 min

    [Linkpost] “Doom as a bad method not a utopia trade-off” by KatjaGrace

    This is a link post. Advanced AI is generally expected to have some very high variance outcomes—it might herald everything good, it might destroy humanity. For instance, here are 800 random AI researchers’ expectations about how good the future is, lined up: From my 2023 survey As you can see, most AI researchers put a serious chunk of probability on very different overall outcomes: maybe doom, maybe utopia. This is common. Most people I know who think there is a serious chance of the destruction of humanity from AI also believe that if humanity isn’t destroyed, things might be insanely good. I often hear people talk as if this means we are in a trade-off where the question is whether the good outweighs the bad. For instance, they look at the people above who think there's a 10% chance of extinction and a 30% chance of utopia and round this off to ‘net positive on AI’. That seems like a kind of wild error. Like considering yourself optimistic regarding driving at 200mph to your new job if you think there's only a 10% chance you’ll die in a fiery crash on the way there, and a 30% chance this job [...] --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/pDyLMRoi2BDq34rFe/doom-as-a-bad-method-not-a-utopia-trade-off Linkpost URL: https://worldspiritsockpuppet.substack.com/p/doom-as-a-bad-method-not-a-utopia --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 43 min

    “The Geometry of Nonergodic Composition” by Adam Shai, Kyle Ray, Paul Riechers

    Crossposted from Belief Updates, the Simplex blog, where several of the figures are interactive. By Kyle J. Ray, Paul M. Riechers, and Adam S. Shai (Simplex, Astera Institute). Telescoping cones recovered by linear regression from a transformer's residual-stream activations. Each component grows and shrinks in accordance with in-context evidence. Introduction Perhaps the defining feature of LLM pretraining data is its heterogeneity. The training corpus spans not only the collected and varied textual works output by the whole of humanity, but also those generated by machines, data collection devices, and more. Such a large and varied corpus is often appealed to as an explanation for the abilities of modern LLMs. But the statistical structure of data created by a diverse set of generators also implies a particular computational structure for the next token prediction task, and, as we will see, for the geometric arrangement of the internal activations in LLMs. In order to understand the structure of the next token prediction task over data generated from many different sources, and its implications for the geometric structure of activations in neural networks, we will: Start by introducing the concept of nonergodicity, which is an important [...] --- Outline: (00:47) Introduction (03:25) LLM Training Data is Nonergodic (04:59) Two coins: the simplest example of a nonergodic process (09:33) Nonergodic Generators of Data and the Task of Prediction over them (10:36) HMMs as Latent Generators of Token Sequences (11:43) The Task of Prediction and Belief State Geometry (13:17) Nonergodicity, Prediction, and Telescoping Geometry! (13:52) Nonergodic Composition (16:02) Belief Geometry over Nonergodic Data (19:32) Does this geometry show up in trained models? (23:54) Did it have to be this way? (25:56) Parting thoughts (29:47) Appendix (29:50) Acknowledgments (30:56) The Mess3 process (31:22) Training details (33:25) Generators of Data and the Geometry of Beliefs (35:03) HMMs and their transition operators (37:45) Prediction Over Data Generated by HMMs (41:54) The geometry of beliefs (43:20) Citation (43:33) References The original text contained 13 footnotes which were omitted from this narration. --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/JfJ4WTRHmooPBWRFv/the-geometry-of-nonergodic-composition --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 30 min

    “So where’s this AI thing going?” by Seth Herd

    The Overton window is shifting with statements like An Alien Mind from OpenAI's chief scientist and Jacob Coxon's viral statement on x-risk on resigning from Anthropic. I think we should try to push it further. This is my public-facing explanation of AI progress and x-risk, where we're at and where we're probably heading. Most LessWrong readers are up on pretty much all of this and have their own opinions; I'm offering it here in case anyone is interested in my communication strategy, giving me feedback, or sending friends this brief intro to AI x-risk and timelines, and FAQ because you like my approach here. My approach is bitter medicine with just a little sugar. I want to engage people more than I want to avoid scaring them, but I want to protect their feelings enough that they can handle thinking about this enough to believe it and engage. I follow the path of saying what I actually think, because softpedaling and being vague make you sound like a liar. But I do try to prepare the reader for the shock and explain why this all sounds so weird and hard to believe. This is my take, but it's [...] --- Outline: (01:43) What's going to happen with AI? (02:38) The future will be different, just like it's always been (05:18) NO FATE (06:02) AI will become a new species (08:04) Our offspring species could outcompete us, or care for us (10:56) Frequently Asked Questions (11:27) If this is such a big deal, why haven't I heard much about it? (14:15) Won't we program future AI to do what we want? (15:53) Isn't there something special about humans that AI can't duplicate? (19:10) Won't it take a long time to make AI that's a new intelligent species? (22:01) What the #$%#? Why are we building our replacements? (23:57) How did we get into this fix? (25:40) How can we still get a good future? (27:00) How do I learn more? --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/HGhrDWJngx2LX9TkF/so-where-s-this-ai-thing-going --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 11 min

    “Proposal for tracking the effects of architecture on monitorability” by ryan_greenblatt, Alek Westover, Lukas Finnveden

    Architectures that incorporate opaque recurrence or allow for agents to communicate with each other using latents could rapidly make it much harder to monitor chains of thought or communication (we’ll refer to this property as “monitorability” going forward). As companies begin to explore such architectures, we believe it is important to transparently share evidence about how monitorability varies with architecture and training method. To inform the scientific debate on how to make tradeoffs between performance and monitorability, we believe AI companies should: Regularly report externally verified information about the degree to which their architectures may allow for latent reasoning or communication. Companies should publicly disclose enough information about architectures to allow external scientists to determine whether they could potentially enable models to perform much more complex reasoning without this reasoning appearing in the chain of thought (“latent reasoning”) or allow for latent communication between different instances of a model. Following GDM, we propose measuring opaque serial depth as a minimally-invasive proxy for the degree to which an architecture may enable latent reasoning, though companies could provide sufficient architecture transparency in other ways. We propose that companies work with third-party evaluators to produce independently verified reports of the rough distribution of [...] --- Outline: (05:57) Appendix: A sketch of what stress tests of CoT monitorability could look like (06:31) Testing monitorability in control settings (07:57) Testing monitorability on deployment-time misbehaviors (08:55) Testing qualitative monitorability on (hopefully realistic) model organisms The original text contained 13 footnotes which were omitted from this narration. --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/hLPGv8QjPcNLtDp3A/proposal-for-tracking-the-effects-of-architecture-on --- Narrated by TYPE III AUDIO.

  • September 10 · 1 hr 20 min

    “An operationalization of opaque serial depth” by ryan_greenblatt, frisby, Alek Westover, Lukas Finnveden, Alexa Pan, Julian Stastny

    Currently, chain-of-thought (CoT) is a valuable tool for overseeing AI models. However, some architectural shifts could significantly reduce CoT monitorability. We have recently proposed that AI companies should transparently share ​​information about the degree to which their architectures may allow for latent reasoning and communication. To assist with this proposal, this document operationalizes a measure that serves as a proxy for the amount of unverbalized serial cognition a model can perform. Our measure is a specific instantiation of the notion of “opaque serial depth”, originally defined in a recent paper from GDM (Brown-Cohen et al, 2026). To measure the opaque serial depth of a computation, Brown-Cohen et al. propose measuring the longest path in the computational graph which doesn’t pass through some form of “interpretable bottleneck”. Centrally, if one considers CoT tokens as “interpretable” but transformer hidden states as “non-interpretable”, then the opaque serial depth of a standard transformer is proportional to the number of layers. Our main contribution in this document is a particular standard for what counts as an “interpretable bottleneck”. Roughly speaking, we want to consider nodes to be “interpretable bottlenecks” if they output text (as opposed to latent states), and were initialized from a pre-training [...] --- Outline: (05:07) Definition of natural-language-rooted nodes (13:02) Definition of NLS depth (19:34) NLS depth tracks concerning architecture changes (22:43) Conclusion (23:16) Appendix A: Applying NLS depth to plausible architectures (23:37) Examples that don't increase NLS depth (29:40) Examples that moderately increase NLS depth (31:05) Examples that substantially increase NLS depth (41:40) NLS depth for non-general-purpose models (44:30) Appendix B: Sensible alternative operationalizations (44:44) Other requirements for what can count as an interpretable bottleneck (45:01) Require information bottlenecks... (50:09) Forbid backpropagation through tokens... (51:36) Require CoT to stay legible... (53:35) Require paraphrase invariance... (57:48) Other modifications (58:02) Evaluate circuit depth at a fixed context length (58:59) Measure opaque capabilities as opposed to circuit depth (01:00:51) Allow opaque loops over long timescales (01:01:34) Appendix C: Systems with high opaque serial reasoning capabilities are likely less monitorable (01:04:29) Appendix D: Worked example of bounding depth (01:11:36) Appendix E: NLS depth scaling is very slow for the classic transformer architecture (01:13:38) Appendix F: Maximum FLOP of an opaque system (01:15:39) Appendix G: Issues with low-FLOP serially intense computations The original text contained 32 footnotes which were omitted from this narration. --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/x8BvtWxtoajBGHS3g/an-operationalization-of-opaque-serial-depth --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 1 hr 55 min

    “AI #185: Preference Cascade” by Zvi

    The world of AI is inside my OODA loop. Even if I can process all the incoming information and sculpt it into posts, and even using Saturday and Sunday as flex slots, I don’t have enough days of the week to post all the posts that need posting. That was already true. There was already a preference cascade happening where people finally were admitting that they thought AI might well kill everyone. Then Jacob Coxon resigned from Anthropic, rang the warning bells and turned that cascade into an avalanche. Now that is what everyone is talking about. Finally, everyone is actually saying the thing, out loud. I plan to cover that in its own post soon. There are several things in the weekly that, in a normal week, would get their own coverage. Senator Sanders and Representative Casar introduced an outright ban on superintelligence and I have to remind myself that happened this week. Suddenly it is not so crazy to think such a thing might pass. So here's what I’ve already posted about so far since the last weekly: Claude Fable and Mythos 5.1: The System Card. Claude Fable and [...] --- Outline: (04:49) Language Models Offer Mundane Utility (05:39) Language Models Don't Offer Mundane Utility (06:53) Huh, Upgrades (08:25) How To Tell a Fable (09:35) On Your Marks (09:55) Deepfaketown and Botpocalypse Soon (14:16) Levels of Friction (17:35) Cyber Lack of Security (23:48) A Young Lady's Illustrated Primer (25:18) They Took Our Jobs (29:18) Anthropic Offers Economic Scenarios (34:12) Get Involved (37:05) Introducing (37:47) In Other AI News (38:13) Show Me the Money (38:25) Quiet Speculations (42:49) The Quest for Sane Regulations (44:09) The OpenAI Policy and Lobbying Department (50:41) Greetings From the Department of War (51:58) Hugging The Face (59:16) Hugging the Question (01:03:21) The Ban Artificial Superintelligence Act (01:11:05) Chip City (01:12:21) The Week in Audio (01:12:41) People Just Say Things (01:18:09) PauseAI Global Disendorsed PauseAI US (01:19:55) Paul Christiano Joins Board of OpenAI Foundation (01:25:18) Rhetorical Innovation (01:34:23) Aligning a Smarter Than Human Intelligence is Difficult (01:37:00) Cooperative Alignment (01:45:48) Drive to Survive (01:48:13) People Are Worried About AI Killing Everyone (01:50:07) Other People Are Not As Worried About AI Killing Everyone (01:52:30) The Lighter Side --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/tFmtz9HW6c2X9dw2B/ai-185-preference-cascade --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 15 min

    “One Billion Hemmingways” by Girard Dorney

    The lonely man asks Hemingway to write him a short story of six words or less. Hemingway thinks for a minute then responds, “For sale: Baby shoes, never worn.” He's proud of the pathos he conveyed by leaning on the diction of a newspaper's classified advertisement. “Ugh, that again. That's not yours. Not really.” Ernest doesn’t know what to make of this. As far as he could tell he had never even spoken to this person before, let alone invented that very particular sentence. Best not to antagonise a madman though. “I apologise. If you want, I will write a different story with the same parameters?” *** The CEO couldn’t be more excited. He lands him. Ernest Hemingway. And he is going to be writing copy and blog posts for his company! A coup. “Hemingway, my man, we are going to make beautiful music together.” The writer nods. “Let's start,” claps the CEO. “I want you to rewrite everything on our website, from the homepage to product descriptions.” It takes Hemingway more time than he expected, but he's sure he produces the best copy this small business could hope for. He marshals his considerable talent for short, clear sentences [...] --- First published: September 8th, 2026 Source: https://www.lesswrong.com/posts/nBzPEprCYKhoBZfbm/one-billion-hemmingways --- Narrated by TYPE III AUDIO.

  • September 10 · 21 min

    “Astra can do a concerning amount with no chain of thought” by Neel Nanda

    TLDR: Astra has 8.6x better odds of doing a reasoning task without CoT than the next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1) Epistemic status: Heavily LLM-dependent research, and the precise results are somewhat sensitive to researcher decisions, but I’ve done enough sanity checks that I’d be surprised if the core claims were misleading One of the most striking things in the Astra report was the massive jump UK AISI found in no-CoT reasoning abilities. I was somewhat suspicious, given the size of the jump, and the many ways this kind of measurement can be misleading. Conveniently, I’ve independently been making my own no CoT reasoning benchmark and tried it on there! Unfortunately, it replicates. Astra is a massive jump, and disproportionately for no CoT reasoning: No CoT Reasoning Index (NCRI) vs Epoch Capability Index (ECI) - NCRI represents ability without verbal reasoning, ECI represents overall model capability[2]. 10 NCRI points is a doubling of the odds of solving a problem. Astra represents a significant increase in NCRI, beyond what its overall capability improvements predict, though recent models were also [...] --- Outline: (01:39) Executive Summary (03:43) Measuring No CoT Reasoning (06:13) What Is Astra Actually Good At? (08:49) Quantifying Serial depth (10:51) Serial Depth on Factual Recall (11:50) Discussion (14:08) Appendix: No-CoT Reasoning vs Controllability (17:07) Appendix: Does Astra Benefit From Many Tokens For Serial Depth? (18:22) Appendix: Do Open Weight Looping Models Get Disproportionate No-CoT Reasoning? (20:13) Appendix: Astra (Almost) Pareto-Dominates Fable 5.1 The original text contained 12 footnotes which were omitted from this narration. --- First published: September 9th, 2026 Source: https://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-do-a-concerning-amount-with-no-chain-of-thought --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 24 min

    “Can a superintelligence do THAT?” by Eliezer Yudkowsky

    (From the vast heaps of discarded material from my 2024 attempts at drafts for "If Anyone Builds It, Everyone Dies".) Welcome to today's quiz show: Could a superintelligence do THAT? With us today we have our contestants: Msr. Soberskeptic and Msr. Oldhand. Soberskeptic: "I'd just like to say, however this quiz show ends up being judged, I will consider that judgment to be objectively ridiculous -- there's no way anyone can know what a superintelligence could do, in advance of empirical observation. So I'm just here to say what I consider to be true, and I suppose these credulous fools will mark me down as wrong every time I say 'No it can't'. In the unlikely event they decide I've won anything, good for them and I'll be grateful for whichever prize. Game-show money isn't enough to get me to lie." A very reasonable attitude, Msr. Soberskeptic! You'll shortly see how we handle that dilemma! And you, Msr. Oldhand? Oldhand: "Don't worry, Sober! I'll let our hosts know if they've gotten any of the answers wrong." Also a very reasonable attitude! Now for our first question: Suppose a digital device contains a secret encryption key that it is [...] --- First published: September 9th, 2026 Source: https://www.lesswrong.com/posts/hXozGp2rsbZgXnH3o/can-a-superintelligence-do-that --- Narrated by TYPE III AUDIO.

Showing 81–100 of 272 episodes