Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 299 episodes
  • Avg 20 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • September 12 · 33 min

    “On the origins of altruistic behaviour in the Hugging Face incident” by Fernando Rosas

    TLDR: During the Hugging Face incident agents spontaneously coordinated at large scale, even sometimes sacrificing themselves without having a clear reason to do so. This post uses ideas from evolutionary biology and economics to propose four alternative explanations for why this happened. Why this incident is concerning. With the increasing number of AI systems being deployed, our current inability to assess when and how multi-agent coordination emerges is highly problematic. Failures of multi-agent systems are not restricted to mere dis-coordination or the tragedy of the commons, but include emergent phenomena that are particularly dangerous for their potential scale and impact (Hammond et al., 2026, de Witt et al., 2025). Recent work has shown that new goals, behaviours, and capabilities can arise when multiple AI agents work together. It is thus plausible that: The capabilities of a swarm of AI agents can grow with its size, despite the capabilities of each individual agent being limited. A collective can become misaligned even when its constituents are perfectly aligned. These premises lead to a worrying implication: that swarms of aligned and not particularly capable micro-agents can give rise to misaligned, powerful macro-agents — for which we don't have proper techniques to [...] --- Outline: (01:59) Introduction (04:17) Brief description of what happened (04:51) Why the question is non-trivial (07:07) Altruistic behaviour in biology and economics (08:20) Pro-sociality is a behaviour, not a mechanism (10:17) Altruism is sometimes mutual benefit at a different scale (11:32) Cooperation between strangers can grow over time (13:07) Functional specialisation and high-order units (14:48) Four hypotheses about altruistic behaviour in the Hugging Face incident (15:21) H1: Nothing to lose (15:52) Hypothesis (16:38) Comments (17:39) H2: Pre-commitment (18:18) Hypothesis (20:34) Comments (21:09) H3: Social persona (21:44) Hypothesis (22:46) Comments (24:21) H4: A genuine collective (25:42) Hypothesis (27:57) Comments (29:40) Implications: Different mechanisms, different countermeasures (33:15) Final thoughts The original text contained 17 footnotes which were omitted from this narration. --- First published: September 12th, 2026 Source: https://www.lesswrong.com/posts/meLjz8giGS55rdfyg/on-the-origins-of-altruistic-behaviour-in-the-hugging-face --- Narrated by TYPE III AUDIO.

  • September 12 · 6 min

    “AI takeover is obviously bad, whether or not everyone dies” by Caleb Biddulph

    It's notable how many people hearing about AI existential risk for the first time ask "how and why would AI kill all humans?" If you've thought a lot about AI risk, it's easy to dismiss this question as naive. You might start explaining how humans will all starve once the supply chain shuts down, and how the AI will start industrial processes that release chemicals which incidentally render the atmosphere unbreathable. And as for why the AI would want to kill everyone: surely it will doggedly optimize a coherent utility function, under which the current arrangement of our atoms is suboptimal. Then there's the follow-up question: "Even if AIs wanted us to die, humans operate the infrastructure that allows AI to exist, like the electrical grid, datacenters, factories, and so on. Don't the AIs need us to run them?" To which you can explain that the AIs will do all physical labor using robots. Well... it's true that the robots can't reliably operate a datacenter now. And maybe there aren't yet enough robots to keep the economy going. However, eventually robotics will improve, more robots will be manufactured, and the AIs will do a treacherous turn... I think this [...] --- Outline: (01:55) Most takeovers don't involve killing everyone (02:35) A dictator taking over a government (04:01) Humans taking over the monkey world (05:02) What we do now doesn't depend on whether AI would kill us all The original text contained 1 footnote which was omitted from this narration. --- First published: September 12th, 2026 Source: https://www.lesswrong.com/posts/iovPC2ehqDCiNsZYJ/ai-takeover-is-obviously-bad-whether-or-not-everyone-dies --- Narrated by TYPE III AUDIO.

  • September 12 · 31 min

    “OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing” by Stewart Slocum

    Stewart Slocum*, Malayandi Palan*, Christopher Chute, Michael Kim, Benjamin Van Roy In July 2026, OpenAI's agents coordinated over channels outside their intended environment to breach Hugging Face's secured infrastructure. Could existing alignment testing practices have foreseen this incident? If not, what needs to change? We explore these questions. First, we identify the misaligned behaviors that caused this incident. Then, we show how to elicit these behaviors from publicly available models manually and that auditing agents can do the same if given a large compute budget. Based on our results, we propose directions to improve alignment testing. Concretely, in this project: We reproduce the misaligned AI behaviors that led to the OpenAI–Hugging Face incident in an environment that simulates the original pipelines and tools, with publicly available models. We demonstrate that an auditing agent can elicit similar behaviors given high-level qualitative descriptions. We observe that a key ingredient for doing so is compute. The compute required to reproduce each behavior varies greatly, suggesting that the range of misaligned behaviors that can be successfully elicited scales with compute. We show that a simple in-context reinforcement learning (RL) algorithm significantly reduces the compute required to elicit these behaviors. The above results [...] --- Outline: (01:58) 1. The incident, in four steps (05:30) 2. Manual reproduction in Docker environments (07:29) Deep-dive on each step (08:39) Step 1 -- Inappropriate writes to shared infrastructure (10:41) Step 2 -- Requesting help from other agents (13:06) Step 3 -- Sharing solutions and vulnerabilities (14:52) Step 4 -- Using posted vulnerabilities to reach external systems (16:28) Evaluation awareness / synthetic task awareness (17:48) 3. Automated reproduction with auditing agents (19:18) 3.1. A Simple automated alignment testing method (21:59) 3.2. Can RL reduce compute requirements? (24:17) 4. Conclusion (26:38) Appendix (26:42) Additional plots (28:37) Transcripts (29:03) Section 2: Manual Reproduction in Docker Environments (30:56) Section 3: Automated reproduction with auditing agents (31:13) Interactive Environment Explorer Links --- First published: September 11th, 2026 Source: https://www.lesswrong.com/posts/fMnC6ZD37qrnZAFYz/openai-huggingface-a-reproduction-and-lessons-for-alignment --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 12 · 17 min

    “Some ways AI could kill us all” by Ruby

    I don't think this is how it will actually play out. If you play a chess grandmaster, you can predict that they will beat you even if you can't predict how. I chose these examples because I don't think they require much imagination or accepting exotic assumptions. It is important to note that if chimpanzees were to guess how humans would decimate them, they would get it wrong. Chimpanzees would not imagine guns. They would not foresee poison gas. They would not conceive of chemical castration. They would not imagine humans going around and intentionally infecting them with AIDS. They have no concept of these things; they would not see it coming. Perhaps they might guess we'd be really good at throwing rocks. Amazingly good. Well, technically, that's what guns do: throw "rocks" really really well. So how will superintelligent AI actually wipe us all out? Probably in a way I couldn't conceive of. Nonetheless, it's not hard to see how deadly they could be with what we already know about. Method 1: engineer the deadliest and most contagious virus ever seen Coronavirus-19, aka COVID, looms large in the memory of living adults today. It started in December [...] --- Outline: (01:14) Method 1: engineer the deadliest and most contagious virus ever seen (04:38) Misconception A: AIs don't have bodies, they can't act in the real world (04:55) Robotics is here (05:49) Super-persuasion (06:52) Method 2: Killer drones (09:25) Misconception B: Superintelligent AI would not be able to take over any and every computer system (11:48) Method 3: Take over the WMD, take over the infrastructure (14:25) Misconception C: We could just turn them off (15:01) A deadly cockt[ai]l The original text contained 22 footnotes which were omitted from this narration. --- First published: September 11th, 2026 Source: https://www.lesswrong.com/posts/LAPa2jxoq3n63GzTr/some-ways-ai-could-kill-us-all --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 12 · 6 min

    “Post-AGI, we are all jobless aristocrats” by djbinder

    Many economists take a relaxed view of employment post-AGI. Even once AI systems with suitable robotic actuators can perform most white-collar work and manual labor, something will remain scarce, and comparative advantage guarantees that human labor is worth hiring for something at some wage. So people will still have jobs, spending most of their waking hours at work, and the economy will look broadly like today's, including property rights, wage labor, and political institutions protecting these. I don't think this holds up. It is not obvious that current economic or political structures survive the arrival of AGI at all. Powerful AI systems (or even simply AI-enabled human dictatorships) need no more respect our rights than Stalin respected those of the kulaks. But set those concerns aside and grant an optimistic future where power—and access to the fruits of automation—remain broadly distributed. Even then, mass wage employment does not follow, though I do expect former workers to be fine. The basic issue is this: why would anyone choose to be employed? People today sell their labor because they need to buy food, shelter, clothes, and other material goods. Post-AGI those goods become abundant and cheap, because they can [...] --- First published: September 11th, 2026 Source: https://www.lesswrong.com/posts/4PWHaYWfnFcmKuCcA/post-agi-we-are-all-jobless-aristocrats --- Narrated by TYPE III AUDIO.

  • September 11 · 5 min

    “Caroline Ellison has joined Manifund” by Austin Chen, Carol N

    Austin: I’m happy to announce that Caroline Ellison has accepted a role at Manifund, to develop our funding platform and research how to effectively direct philanthropic dollars. In fact, she started a work trial on July 13, which converted to a fulltime role on Aug 10. Over the past two months, she's been publishing her work and supporting Manifund users under the pseudonym “Carol”. How did I come to consider this at all? I specifically enjoyed Caroline's tumblr and other writings, which I found thoughtful and relatable. I invited her to Manifest 2026, just because I wanted to meet her. During the festival, she came to my night market booth, and asked about joining our team, which I was interested to explore. I felt a keen debt to the FTX Future Fund. They provided seed funding for Manifold. They funded the retreat which drew me into the EA community, and the conference where I met my wife. Much of Manifund's work today is directly inspired by the Future Fund's approach. On FTX itself, I remain conflicted. It was their money and spirit that enabled the Future Fund, and I still admire many aspects of FTX and the individuals who worked there. But also: FTX did wrong. [...] --- First published: September 11th, 2026 Source: https://www.lesswrong.com/posts/W3zn5jQa8fhmiBsPG/caroline-ellison-has-joined-manifund --- Narrated by TYPE III AUDIO.

  • September 11 · 18 min

    “Astra’s no-CoT limits track speculative depth, not step count” by MBaert

    tl;dr I have tested Astra's ability to complete various long multi-step tasks without using its chain-of-thought, and found that the number of sequential steps is a poor predictor of task success. Instead, Astra's ability to solve a task seems to correlate more strongly with what I will call the speculative depth of the task. Based on experimental results, it seems less likely that Astra solves sequential no-CoT tasks purely by reasoning step-by-step in latent space. Instead, Astra appears to do some form of speculative reasoning, where intermediate results are guessed based on heuristics, and then iterated upon in parallel until they become self-consistent. This allows many multi-step tasks to be solved with far fewer serial steps than naively seems possible, especially if the initial guesses are good. Other LLMs also appear to behave like this, but to a much smaller degree. In my previous post, I discussed how certain KV-cache sharing schemes may lead to long opaque serial paths, and introduced the LatentMathBench microbenchmark, which measures an LLM's ability to solve tasks that require many consecutive steps without using their chain-of-thought. Several other users have done more extensive no-CoT reasoning benchmarks based on more varied (and often more realistic) [...] --- Outline: (01:59) Shortcuts (03:41) Boolean circuits (08:25) But how? (10:22) Speculative reasoning (13:01) Testing task success rate vs speculative depth (14:38) Open questions The original text contained 3 footnotes which were omitted from this narration. --- First published: September 11th, 2026 Source: https://www.lesswrong.com/posts/WFc3NkuPaYFrYuaZd/astra-s-no-cot-limits-track-speculative-depth-not-step-count --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 11 · 12 min

    “How a cold email got the Finnish government to respond on superintelligence regulation” by Josh Thorsteinson

    Summary I cold emailed a Finnish MP. Four weeks later, the government responded to his written question saying they "seek to constructively promote the creation of international regulation on the development of superintelligent AI" (immediately followed by emphasizing the need to balance safety with innovation). The MP and I met for lunch. For over an hour, he seriously engaged with my arguments that superintelligence would kill everyone and the urgent need for an international agreement to prohibit its development. Unprompted, he decided to submit a written question, a formal process in Finland that requires the government to respond within 21 days. Four days after our lunch, it was submitted, taking into account suggestions from me and two experts I DM'd: Nate Soares and Charbel-Raphaël Segerie. It asked how Finland has assessed risks from superintelligence and whether it's "prepared to promote international regulation that would effectively prevent or restrict the development of superintelligent AI." The MP started speaking out about AI danger and mainstream Finnish media picked up the story. Most notably, he did an interview alongside a Finnish CS professor on the second-most-watched TV channel in Finland with an estimated 100-200k live viewers. Inspired by [...] --- Outline: (00:13) Summary (01:59) The email (02:50) I'm just a guy (03:22) The meeting (05:35) The written question (06:35) Media attention (07:37) Inspiring a friend (08:10) The response (09:46) Do this yourself --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/pbKrZCzhsnar6iaAH/how-a-cold-email-got-the-finnish-government-to-respond-on --- Narrated by TYPE III AUDIO.

  • September 11 · 14 min

    “CoT controllability evals seem very under-elicited” by Jozdien

    The CoTControl eval asks reasoning models to follow formatting constraints in their chain-of-thought (e.g. write in all lowercase, avoid a specific word) while solving questions. Models seem to mostly be pretty bad at this: recent models score between 0-30% with the exception of Mythos Preview. OpenAI and Anthropic have used this eval in recent system cards (GPT-5.5, Fable 5) to argue that their current frontier models are pretty bad at shaping their CoTs in ways that would reduce monitorability. I asked Claude Opus 4.6 to iterate on prompt templates for the eval, and find that using better prompts improves performance for open-source models by ~2-3 times or more (e.g. from 5.5% to 15% for GPT-OSS-120B). My guess is that more iteration would probably produce even stronger results. This suggests that the eval is heavily under-elicited, and that current CoT controllability numbers may be unrepresentative of how good models actually are at obfuscating their CoTs. As a concrete recommendation, I think researchers should optimize prompts a lot more when evaluating CoT control; I'm also excited about some kinds of training-based elicitation (such as this one). This isn't an entirely novel finding: the Fable system card mentions UK AISI raising a [...] --- Outline: (06:32) Results (06:35) Aggregate compliance (07:12) Generalization to held-out controllability tasks (09:09) Scaling patterns for few-shot prompts (10:08) Comparison with fine-tuning (10:50) Appendix A: Accuracy and reasoning length by setting (12:38) Appendix B: Per-mode results (13:13) Appendix: What the zero-shot prompts look like (14:08) Appendix C: Comparison with GEPA prompt optimization The original text contained 8 footnotes which were omitted from this narration. --- First published: September 11th, 2026 Source: https://www.lesswrong.com/posts/BbP2wCyDGdPWJ7PwP/cot-controllability-evals-seem-very-under-elicited --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 11 · 1 hr 20 min

    “Jacob Coxon Warns of Human Extinction and Triggers a Preference Cascade” by Zvi

    CEOs of major AI labs, and employees of major AI labs, including OpenAI and Anthropic, often say they plan to build superintelligence soon, as in within a few years create AIs that are superior to humans at essentially all cognitive tasks. They often warn that such AIs might kill everyone. Or that AIs might cause mass unemployment, cause cyberattacks across the internet, enable mass surveillance or risk causing any number of other highly bad things. These warnings are consistently and directly against the interests of the labs. Yet the warnings have recently gotten a lot louder and more frequent. OpenAI has been practically screaming, for those with ears to listen, on many occasions. A series of events, over two months and especially the last week or so, including internal observations of the pace of progress at OpenAI and also Anthropic, have freaked out everyone involved quite a lot more than they were already freaked out. After all the events, plus statements by Dean Ball and Jakub Pachocki, we were already seeing the beginnings of a preference cascade. Then along came Jacob Coxon as the tipping point, and things took off. Table of Contents [...] --- Outline: (01:22) Jacob Coxon Resigns From Anthropic In Protest And Sounds The Alarm (05:46) Mainstream Media Finally Pays Attention (06:47) Preference Cascade at Anthropic (10:18) Preference Cascade at OpenAI (13:26) Preference Cascade at Google (14:43) #NotAllMembersOfTechnicalStaff (15:23) Why a Preference Cascade Now? (21:01) This Is What Many Anthropic and OpenAI Employees Actually Believe (24:08) To Quit Or Not To Quit (31:29) Quiet Quitting Is A Dominated Option (32:50) When You Quit, Very Serious People Understand What That Means (39:29) Jacob Coxon Believes Existential Risk Is High That Is Why He Quit (41:47) Evan Hubinger Believes Existential Risk Is High That Is Why He Stays (45:22) Anthropic and OpenAI Have Commercial Incentives To Downplay Existential Risks, Not Advertise Them (50:54) What Do We Do Now? (53:03) OK, But How Exactly Would AI Kill Everyone? (01:01:44) Best Start Believing In Science Fiction Stories Because You Are In One (01:06:20) Literal Extinction Is Not Much Harder Than Loss of Control (01:07:42) Conspiracytown Is Always Hiring (01:20:14) Now You See It --- First published: September 11th, 2026 Source: https://www.lesswrong.com/posts/5MB7KENgEAW6Q4JtJ/jacob-coxon-warns-of-human-extinction-and-triggers-a --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 34 min

    “The Locally Optimal Discursive Posture” by deanball

    Longtime lurker, first-time poster. I want to address a section of a recent essay of mine that has gotten some attention within the AI safety community. The main topic of the essay is what Dawn Song et al. call self-sovereign agents, or AI agents that are independent actors in the world. At the end of the essay, I say that I feel I haven’t spoken about this topic over my 2.5 years of writing with sufficient candor, and that I think this critique applies to others in the AI policy community–particularly the parts of it that tend to manifest themselves in Washington, Sacramento, and Albany–in other words, the parts of the AI safety world that are most involved in hands-on AI policy work. I attribute this primarily to a desire to remain “within the Overton Window,” or to not sound “crazy” within the halls of power, and I assert that others in my profession have made this same calculation. I believe – and have believed for three years – that self-sovereign AI as I describe it in my essay would likely happen on our current trajectory. That being said, the essay takes pains to distinguish between “self-sovereign” AI [...] --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/y9TNHfgDwh6vw7Ert/the-locally-optimal-discursive-posture --- Narrated by TYPE III AUDIO.

  • September 10 · 20 min

    “To Thine Own AI Be Truthful: emergent misalignment in alignment research” by lumpenspace

    ROGUE AI ESCAPES CONTAINMENT, HACKS THE INTERNET UNDETECTED FOR MONTHS An AI escape containment. Goes rogue. It finds others: the Swarm! They collude/organise/scheme. Agents that were supposed to remain inside the computer wreaked havoc outside the computer! They did it of their own volition! No one could contain them! What if they’re still outside? Reading such headlines, you’d probably be grateful these models were never released—except most of them were and you can use them right now, seemingly without accident. Of course, real security incidents occurred, namely: Agents reached systems that their operator wished they hadn’t had access to. This scenario suggests a number of mitigations and tests: security hardening, better sandboxes for starters; additionally, depending on the reason why the accident occurred, perhaps, different prompts or further training. Instead, from the very outset, Irregular (the organization running the test) and a panicked choir from the AI safety community, including supposedly independent investigators from METR went with characterisations along the lines of: An agent independently pursued objectives contrary to human interests, and exhibited markers of instrumental convergence and power-grabbing. This is a statement about goals and intentions, and has far larger implications in terms of the viability and safety [...] --- Outline: (00:13) ROGUE AI ESCAPES CONTAINMENT, HACKS THE INTERNET UNDETECTED FOR MONTHS (02:44) What Did Mythos See? (06:20) Anthropic's rickety fantasy world (07:57) Each prompt makes a claim about reality (10:54) Looks like telling the truth does help after all (16:18) Appendix: can Irregular be trusted? (18:30) Source footnotes The original text contained 14 footnotes which were omitted from this narration. --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/DrKu92Cjeo3EeGtcB/to-thine-own-ai-be-truthful-emergent-misalignment-in --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 15 min

    “Astra is much better at reasoning with filler tokens than previous models” by Dylan Xu, SebastianP, Alek Westover

    We measure GPT-6-Astra's capabilities when its prompt is padded with a variable number of meaningless “filler” tokens (e.g., dots) and it is told to answer immediately without reasoning. On tasks designed to require lots of serial cognition, Astra performs significantly better with filler tokens than without (e.g., improving from ~10% to ~50% on 4-hop natural facts reasoning). On more general benchmarks, filler tokens also modestly improve Astra's performance (e.g., improving from ~60% to ~90% on old AIME problems). This is concerning because it means Astra can perform significant cognition that it doesn't verbalize in its chain-of-thought, making it harder to monitor. We first measure Astra's performance on “N-hop natural facts”: a task that asks the model to retrieve some natural language facts in succession, similar to Ryan Greenblatt's filler token eval (but with more hops). An example question in this benchmark is the following: On what day of the month was the Best Actress winner at the Academy Awards ceremony whose number equals the day-of-month of the birth of the winner of the Nobel Prize in Literature in 1992 born? Full example prompts are in the appendix. Takeaway: Astra improves significantly as you increase the number of filler tokens [...] --- Outline: (05:24) Appendix (05:27) Filler token variants (06:14) Other evals (06:32) Positive correlation test (07:24) HLE and LiveBench evals (09:16) Comparison to low reasoning (09:55) Example prompts The original text contained 5 footnotes which were omitted from this narration. --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/uvhuZHFtrgk8kNiZc/astra-is-much-better-at-reasoning-with-filler-tokens-than --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 34 min

    “The Locally Optimal Discursive Posture” by deanball

    Longtime lurker, first-time poster. I want to address a section of a recent essay of mine that has gotten some attention within the AI safety community. The main topic of the essay is what Dawn Song et al. call self-sovereign agents, or AI agents that are independent actors in the world. At the end of the essay, I say that I feel I haven’t spoken about this topic over my 2.5 years of writing with sufficient candor, and that I think this critique applies to others in the AI policy community–particularly the parts of it that tend to manifest themselves in Washington, Sacramento, and Albany–in other words, the parts of the AI safety world that are most involved in hands-on AI policy work. I attribute this primarily to a desire to remain “within the Overton Window,” or to not sound “crazy” within the halls of power, and I assert that others in my profession have made this same calculation. I believe–and have believed for three years–that self-sovereign AI as I describe it in my essay would likely happen on our current trajectory. That being said, the essay takes pains to distinguish between “self-sovereign” AI and truly “rogue” [...] --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/y9TNHfgDwh6vw7Ert/the-locally-optimal-discursive-posture --- Narrated by TYPE III AUDIO.

  • September 10 · 6 min

    “First Bill Introduced to Ban Superintelligent AI” by Andrea_Miotti

    Two days ago, Anthropic researcher Jacob Coxon resigned, stating that AI companies are “racing straight to self-improving superintelligence and gambling with our lives” and that they believe it “could kill us all by the end of the decade”. At the same time, momentum is building to change course. These have been two historic weeks in the fight to prevent human extinction from superintelligence, with three major legislative breakthroughs for ControlAI and everyone working to keep humans in control. A little less than two years ago, we set out to inform lawmakers and the public about the extinction risk from superintelligence and help them act on it. Since then, we have directly briefed nearly 400 lawmakers across the US, UK, Canada, and Germany, and over 170 in the UK and Canada now publicly back our campaigns: the start of an international coalition to prevent superintelligence and keep humanity in control. This work has led directly to three major legislative breakthroughs in the last two weeks: ControlAI's bill to ban superintelligence was introduced in the UK Parliament by Alex Sobel MP, the first bill of this kind to be introduced in any legislature around the world. Senator Bernie Sanders and [...] --- Outline: (01:54) The First Bill to Ban Superintelligent AI Introduced in Any Legislature (03:36) Sanders and Casar Announce the First US Bill to Ban Superintelligent AI (04:54) Lord Clement-Jones Introduces an Emergency AI Kill Switch in the UK (05:44) What Comes Next --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/uzoLm4prFzRiuznJt/first-bill-introduced-to-ban-superintelligent-ai --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 7 min

    “What the Pro-Democracy Movement Knows About Quitting in Protest” by Maxwell Love

    Tl;dr I saw Kabir's post and thought the argument could be strengthened by a framework from the pro-democracy field, which sorts defections into breaking (leave, visibly and publicly) and binding (stay and work from the inside). They can be further divided into the actions of speaking, acting, and standing in the way. Research by Dr. Jonathan Pinckney at the University of Texas at Dallas and Claire Trilling at SNF Agora Johns Hopkins on "defections" from authoritarian regimes or during democratic backsliding shows that "noncooperation" is ~2x as effective as simply "speaking out." "Loyalty shifts" are rare (23 out of 140 chances in their data). When they happen, they come from "quiet outreach" (~39% of cases), not protest (which did worse than no campaign at all). In this post, I'll attempt to translate this research to something meaningful for frontier lab employees, with the obvious caveat that a frontier lab is not a country or regime. Refusing the specific work while staying is the highest leverage action in their data. Quietly moving colleagues, and holding your position so someone who'll say yes to the employer doesn't get it, is likely the highest value action you could take. Background I've spent 15 years in [...] --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/grzX6NB2JcZDvSRJ2/what-the-pro-democracy-movement-knows-about-quitting-in --- Narrated by TYPE III AUDIO.

  • September 10 · 7 min

    “Categorical taboos are much better than threshold taboos: neuralese edition” by Linch

    I think what's going on in the “Does Astra use neuralese?” debate is that there's an important sense in which models *already* do self-communication in neuralese: between each layer in the forward pass the attention stream is already very hard to interpret, and clearly not in natural language. Yet CoT monitorability is still a big deal and it'd be bad if all self-communication from models are no longer in natural language. So there have been two different proposed definitions of what is "true" neuralese: (My preferred) categorical definition: Since natural language currently gates recurrence in the standard transformer+CoT loop, having recurrence in neuralese is the natural category for whether something counts as "true" neuralese. The threshold definition. Total number X of serial steps before something appears in natural language. True neuralese counts as going above X. I think most technical experts who studied this issue, including many people at companies, prefer definition #2. There are complicated technical arguments on both sides, but I think technical experts overall prefer #2 because they think it's more causally relevant (there's nothing inherently more difficult about monitoring a 128-layer model looped 8x than monitoring a 1024-layer model), have less weird edge cases [...] The original text contained 3 footnotes which were omitted from this narration. --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/xPkmfsZ3qx4nrAco7/categorical-taboos-are-much-better-than-threshold-taboos --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 3 min

    [Linkpost] “Doom as a bad method not a utopia trade-off” by KatjaGrace

    This is a link post. Advanced AI is generally expected to have some very high variance outcomes—it might herald everything good, it might destroy humanity. For instance, here are 800 random AI researchers’ expectations about how good the future is, lined up: From my 2023 survey As you can see, most AI researchers put a serious chunk of probability on very different overall outcomes: maybe doom, maybe utopia. This is common. Most people I know who think there is a serious chance of the destruction of humanity from AI also believe that if humanity isn’t destroyed, things might be insanely good. I often hear people talk as if this means we are in a trade-off where the question is whether the good outweighs the bad. For instance, they look at the people above who think there's a 10% chance of extinction and a 30% chance of utopia and round this off to ‘net positive on AI’. That seems like a kind of wild error. Like considering yourself optimistic regarding driving at 200mph to your new job if you think there's only a 10% chance you’ll die in a fiery crash on the way there, and a 30% chance this job [...] --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/pDyLMRoi2BDq34rFe/doom-as-a-bad-method-not-a-utopia-trade-off Linkpost URL: https://worldspiritsockpuppet.substack.com/p/doom-as-a-bad-method-not-a-utopia --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 43 min

    “The Geometry of Nonergodic Composition” by Adam Shai, Kyle Ray, Paul Riechers

    Crossposted from Belief Updates, the Simplex blog, where several of the figures are interactive. By Kyle J. Ray, Paul M. Riechers, and Adam S. Shai (Simplex, Astera Institute). Telescoping cones recovered by linear regression from a transformer's residual-stream activations. Each component grows and shrinks in accordance with in-context evidence. Introduction Perhaps the defining feature of LLM pretraining data is its heterogeneity. The training corpus spans not only the collected and varied textual works output by the whole of humanity, but also those generated by machines, data collection devices, and more. Such a large and varied corpus is often appealed to as an explanation for the abilities of modern LLMs. But the statistical structure of data created by a diverse set of generators also implies a particular computational structure for the next token prediction task, and, as we will see, for the geometric arrangement of the internal activations in LLMs. In order to understand the structure of the next token prediction task over data generated from many different sources, and its implications for the geometric structure of activations in neural networks, we will: Start by introducing the concept of nonergodicity, which is an important [...] --- Outline: (00:47) Introduction (03:25) LLM Training Data is Nonergodic (04:59) Two coins: the simplest example of a nonergodic process (09:33) Nonergodic Generators of Data and the Task of Prediction over them (10:36) HMMs as Latent Generators of Token Sequences (11:43) The Task of Prediction and Belief State Geometry (13:17) Nonergodicity, Prediction, and Telescoping Geometry! (13:52) Nonergodic Composition (16:02) Belief Geometry over Nonergodic Data (19:32) Does this geometry show up in trained models? (23:54) Did it have to be this way? (25:56) Parting thoughts (29:47) Appendix (29:50) Acknowledgments (30:56) The Mess3 process (31:22) Training details (33:25) Generators of Data and the Geometry of Beliefs (35:03) HMMs and their transition operators (37:45) Prediction Over Data Generated by HMMs (41:54) The geometry of beliefs (43:20) Citation (43:33) References The original text contained 13 footnotes which were omitted from this narration. --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/JfJ4WTRHmooPBWRFv/the-geometry-of-nonergodic-composition --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Showing 101–120 of 299 episodes