Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 299 episodes
  • Avg 20 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • September 10 · 30 min

    “So where’s this AI thing going?” by Seth Herd

    The Overton window is shifting with statements like An Alien Mind from OpenAI's chief scientist and Jacob Coxon's viral statement on x-risk on resigning from Anthropic. I think we should try to push it further. This is my public-facing explanation of AI progress and x-risk, where we're at and where we're probably heading. Most LessWrong readers are up on pretty much all of this and have their own opinions; I'm offering it here in case anyone is interested in my communication strategy, giving me feedback, or sending friends this brief intro to AI x-risk and timelines, and FAQ because you like my approach here. My approach is bitter medicine with just a little sugar. I want to engage people more than I want to avoid scaring them, but I want to protect their feelings enough that they can handle thinking about this enough to believe it and engage. I follow the path of saying what I actually think, because softpedaling and being vague make you sound like a liar. But I do try to prepare the reader for the shock and explain why this all sounds so weird and hard to believe. This is my take, but it's [...] --- Outline: (01:43) What's going to happen with AI? (02:38) The future will be different, just like it's always been (05:18) NO FATE (06:02) AI will become a new species (08:04) Our offspring species could outcompete us, or care for us (10:56) Frequently Asked Questions (11:27) If this is such a big deal, why haven't I heard much about it? (14:15) Won't we program future AI to do what we want? (15:53) Isn't there something special about humans that AI can't duplicate? (19:10) Won't it take a long time to make AI that's a new intelligent species? (22:01) What the #$%#? Why are we building our replacements? (23:57) How did we get into this fix? (25:40) How can we still get a good future? (27:00) How do I learn more? --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/HGhrDWJngx2LX9TkF/so-where-s-this-ai-thing-going --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 11 min

    “Proposal for tracking the effects of architecture on monitorability” by ryan_greenblatt, Alek Westover, Lukas Finnveden

    Architectures that incorporate opaque recurrence or allow for agents to communicate with each other using latents could rapidly make it much harder to monitor chains of thought or communication (we’ll refer to this property as “monitorability” going forward). As companies begin to explore such architectures, we believe it is important to transparently share evidence about how monitorability varies with architecture and training method. To inform the scientific debate on how to make tradeoffs between performance and monitorability, we believe AI companies should: Regularly report externally verified information about the degree to which their architectures may allow for latent reasoning or communication. Companies should publicly disclose enough information about architectures to allow external scientists to determine whether they could potentially enable models to perform much more complex reasoning without this reasoning appearing in the chain of thought (“latent reasoning”) or allow for latent communication between different instances of a model. Following GDM, we propose measuring opaque serial depth as a minimally-invasive proxy for the degree to which an architecture may enable latent reasoning, though companies could provide sufficient architecture transparency in other ways. We propose that companies work with third-party evaluators to produce independently verified reports of the rough distribution of [...] --- Outline: (05:57) Appendix: A sketch of what stress tests of CoT monitorability could look like (06:31) Testing monitorability in control settings (07:57) Testing monitorability on deployment-time misbehaviors (08:55) Testing qualitative monitorability on (hopefully realistic) model organisms The original text contained 13 footnotes which were omitted from this narration. --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/hLPGv8QjPcNLtDp3A/proposal-for-tracking-the-effects-of-architecture-on --- Narrated by TYPE III AUDIO.

  • September 10 · 1 hr 20 min

    “An operationalization of opaque serial depth” by ryan_greenblatt, frisby, Alek Westover, Lukas Finnveden, Alexa Pan, Julian Stastny

    Currently, chain-of-thought (CoT) is a valuable tool for overseeing AI models. However, some architectural shifts could significantly reduce CoT monitorability. We have recently proposed that AI companies should transparently share ​​information about the degree to which their architectures may allow for latent reasoning and communication. To assist with this proposal, this document operationalizes a measure that serves as a proxy for the amount of unverbalized serial cognition a model can perform. Our measure is a specific instantiation of the notion of “opaque serial depth”, originally defined in a recent paper from GDM (Brown-Cohen et al, 2026). To measure the opaque serial depth of a computation, Brown-Cohen et al. propose measuring the longest path in the computational graph which doesn’t pass through some form of “interpretable bottleneck”. Centrally, if one considers CoT tokens as “interpretable” but transformer hidden states as “non-interpretable”, then the opaque serial depth of a standard transformer is proportional to the number of layers. Our main contribution in this document is a particular standard for what counts as an “interpretable bottleneck”. Roughly speaking, we want to consider nodes to be “interpretable bottlenecks” if they output text (as opposed to latent states), and were initialized from a pre-training [...] --- Outline: (05:07) Definition of natural-language-rooted nodes (13:02) Definition of NLS depth (19:34) NLS depth tracks concerning architecture changes (22:43) Conclusion (23:16) Appendix A: Applying NLS depth to plausible architectures (23:37) Examples that don't increase NLS depth (29:40) Examples that moderately increase NLS depth (31:05) Examples that substantially increase NLS depth (41:40) NLS depth for non-general-purpose models (44:30) Appendix B: Sensible alternative operationalizations (44:44) Other requirements for what can count as an interpretable bottleneck (45:01) Require information bottlenecks... (50:09) Forbid backpropagation through tokens... (51:36) Require CoT to stay legible... (53:35) Require paraphrase invariance... (57:48) Other modifications (58:02) Evaluate circuit depth at a fixed context length (58:59) Measure opaque capabilities as opposed to circuit depth (01:00:51) Allow opaque loops over long timescales (01:01:34) Appendix C: Systems with high opaque serial reasoning capabilities are likely less monitorable (01:04:29) Appendix D: Worked example of bounding depth (01:11:36) Appendix E: NLS depth scaling is very slow for the classic transformer architecture (01:13:38) Appendix F: Maximum FLOP of an opaque system (01:15:39) Appendix G: Issues with low-FLOP serially intense computations The original text contained 32 footnotes which were omitted from this narration. --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/x8BvtWxtoajBGHS3g/an-operationalization-of-opaque-serial-depth --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 1 hr 55 min

    “AI #185: Preference Cascade” by Zvi

    The world of AI is inside my OODA loop. Even if I can process all the incoming information and sculpt it into posts, and even using Saturday and Sunday as flex slots, I don’t have enough days of the week to post all the posts that need posting. That was already true. There was already a preference cascade happening where people finally were admitting that they thought AI might well kill everyone. Then Jacob Coxon resigned from Anthropic, rang the warning bells and turned that cascade into an avalanche. Now that is what everyone is talking about. Finally, everyone is actually saying the thing, out loud. I plan to cover that in its own post soon. There are several things in the weekly that, in a normal week, would get their own coverage. Senator Sanders and Representative Casar introduced an outright ban on superintelligence and I have to remind myself that happened this week. Suddenly it is not so crazy to think such a thing might pass. So here's what I’ve already posted about so far since the last weekly: Claude Fable and Mythos 5.1: The System Card. Claude Fable and [...] --- Outline: (04:49) Language Models Offer Mundane Utility (05:39) Language Models Don't Offer Mundane Utility (06:53) Huh, Upgrades (08:25) How To Tell a Fable (09:35) On Your Marks (09:55) Deepfaketown and Botpocalypse Soon (14:16) Levels of Friction (17:35) Cyber Lack of Security (23:48) A Young Lady's Illustrated Primer (25:18) They Took Our Jobs (29:18) Anthropic Offers Economic Scenarios (34:12) Get Involved (37:05) Introducing (37:47) In Other AI News (38:13) Show Me the Money (38:25) Quiet Speculations (42:49) The Quest for Sane Regulations (44:09) The OpenAI Policy and Lobbying Department (50:41) Greetings From the Department of War (51:58) Hugging The Face (59:16) Hugging the Question (01:03:21) The Ban Artificial Superintelligence Act (01:11:05) Chip City (01:12:21) The Week in Audio (01:12:41) People Just Say Things (01:18:09) PauseAI Global Disendorsed PauseAI US (01:19:55) Paul Christiano Joins Board of OpenAI Foundation (01:25:18) Rhetorical Innovation (01:34:23) Aligning a Smarter Than Human Intelligence is Difficult (01:37:00) Cooperative Alignment (01:45:48) Drive to Survive (01:48:13) People Are Worried About AI Killing Everyone (01:50:07) Other People Are Not As Worried About AI Killing Everyone (01:52:30) The Lighter Side --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/tFmtz9HW6c2X9dw2B/ai-185-preference-cascade --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 15 min

    “One Billion Hemmingways” by Girard Dorney

    The lonely man asks Hemingway to write him a short story of six words or less. Hemingway thinks for a minute then responds, “For sale: Baby shoes, never worn.” He's proud of the pathos he conveyed by leaning on the diction of a newspaper's classified advertisement. “Ugh, that again. That's not yours. Not really.” Ernest doesn’t know what to make of this. As far as he could tell he had never even spoken to this person before, let alone invented that very particular sentence. Best not to antagonise a madman though. “I apologise. If you want, I will write a different story with the same parameters?” *** The CEO couldn’t be more excited. He lands him. Ernest Hemingway. And he is going to be writing copy and blog posts for his company! A coup. “Hemingway, my man, we are going to make beautiful music together.” The writer nods. “Let's start,” claps the CEO. “I want you to rewrite everything on our website, from the homepage to product descriptions.” It takes Hemingway more time than he expected, but he's sure he produces the best copy this small business could hope for. He marshals his considerable talent for short, clear sentences [...] --- First published: September 8th, 2026 Source: https://www.lesswrong.com/posts/nBzPEprCYKhoBZfbm/one-billion-hemmingways --- Narrated by TYPE III AUDIO.

  • September 10 · 21 min

    “Astra can do a concerning amount with no chain of thought” by Neel Nanda

    TLDR: Astra has 8.6x better odds of doing a reasoning task without CoT than the next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1) Epistemic status: Heavily LLM-dependent research, and the precise results are somewhat sensitive to researcher decisions, but I’ve done enough sanity checks that I’d be surprised if the core claims were misleading One of the most striking things in the Astra report was the massive jump UK AISI found in no-CoT reasoning abilities. I was somewhat suspicious, given the size of the jump, and the many ways this kind of measurement can be misleading. Conveniently, I’ve independently been making my own no CoT reasoning benchmark and tried it on there! Unfortunately, it replicates. Astra is a massive jump, and disproportionately for no CoT reasoning: No CoT Reasoning Index (NCRI) vs Epoch Capability Index (ECI) - NCRI represents ability without verbal reasoning, ECI represents overall model capability[2]. 10 NCRI points is a doubling of the odds of solving a problem. Astra represents a significant increase in NCRI, beyond what its overall capability improvements predict, though recent models were also [...] --- Outline: (01:39) Executive Summary (03:43) Measuring No CoT Reasoning (06:13) What Is Astra Actually Good At? (08:49) Quantifying Serial depth (10:51) Serial Depth on Factual Recall (11:50) Discussion (14:08) Appendix: No-CoT Reasoning vs Controllability (17:07) Appendix: Does Astra Benefit From Many Tokens For Serial Depth? (18:22) Appendix: Do Open Weight Looping Models Get Disproportionate No-CoT Reasoning? (20:13) Appendix: Astra (Almost) Pareto-Dominates Fable 5.1 The original text contained 12 footnotes which were omitted from this narration. --- First published: September 9th, 2026 Source: https://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-do-a-concerning-amount-with-no-chain-of-thought --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 10 · 24 min

    “Can a superintelligence do THAT?” by Eliezer Yudkowsky

    (From the vast heaps of discarded material from my 2024 attempts at drafts for "If Anyone Builds It, Everyone Dies".) Welcome to today's quiz show: Could a superintelligence do THAT? With us today we have our contestants: Msr. Soberskeptic and Msr. Oldhand. Soberskeptic: "I'd just like to say, however this quiz show ends up being judged, I will consider that judgment to be objectively ridiculous -- there's no way anyone can know what a superintelligence could do, in advance of empirical observation. So I'm just here to say what I consider to be true, and I suppose these credulous fools will mark me down as wrong every time I say 'No it can't'. In the unlikely event they decide I've won anything, good for them and I'll be grateful for whichever prize. Game-show money isn't enough to get me to lie." A very reasonable attitude, Msr. Soberskeptic! You'll shortly see how we handle that dilemma! And you, Msr. Oldhand? Oldhand: "Don't worry, Sober! I'll let our hosts know if they've gotten any of the answers wrong." Also a very reasonable attitude! Now for our first question: Suppose a digital device contains a secret encryption key that it is [...] --- First published: September 9th, 2026 Source: https://www.lesswrong.com/posts/hXozGp2rsbZgXnH3o/can-a-superintelligence-do-that --- Narrated by TYPE III AUDIO.

  • September 9 · 1 min

    “Recommendations for People Getting into Technical AI Governance Research” by Aaron_Scher, yams, peterbarnett, Naci Cankaya

    In summer 2026, MIRI ran a technical governance fellowship to expand the team and identify promising early-stage technical governance researchers. To support the program, we also developed a couple of resources for our fellows. This post is a trailhead for those interested in MIRI-style technical governance work, including two reading lists (one on how to do research, and one on prior work in the field we think is worth exploring) and tips on how to do research from the team. We’re sharing these, in part, because we do not currently plan to run another fellowship in the near future. In our experience, promising researchers in this subfield have spent significant time getting acquainted with the field before engaging with, e.g., a fellowship program or internship. This guide can be taken as our current best guess of how to spend your first couple dozen or so hours after completing AI Safety Fundamentals or similar, especially if you’re interested in contributing to technical governance research. Since these resources were originally intended for one-time use, they’re not especially polished and we may not keep them up to date (although much of the material is evergreen). Medium-Level Technical AI Governance Reading [...] --- First published: September 9th, 2026 Source: https://www.lesswrong.com/posts/GRPswd5v7haob97fA/recommendations-for-people-getting-into-technical-ai --- Narrated by TYPE III AUDIO.

  • September 9 · 1 hr 4 min

    “GPT-6 Astra: The System Card, Alignment and What Comes Next” by Zvi

    OpenAI claims that Astra is ‘the most intelligent and most aligned [available] model’ in the world. Not the most intelligent and aligned OpenAI model, but the most period. That is bold talk. It risks overstepping, and by doing so souring the release of what is clearly an excellent model. As do the severe problems with monitorability. It also raises the question of what they mean by ‘most aligned model.’ How do they define ‘aligned.’ Why do they think it is more aligned than Claude Fable 5.1? keltan : “a significant step forward in […] alignment.” Buddy, how tf are you measuring ‘alignment’? Would love to know because being able to measure that would save the fucking world. roon (OpenAI): low rates of cheating Rob Miles: *detected cheating keltan : Thank you for clarifying. But you know what I’m gonna say next, right? roon (OpenAI): that this metrics are not a full solve of alignment and will break discontinuously keltan : Yep. But I would have said it in a dumber way. Something like: Low Rates of Cheating ≠ Alignment roon (OpenAI): I agree but also in some real sense [...] --- Outline: (04:12) OpenAI's Safety Claims About Astra (1) (08:33) Preparedness Capabilities Assessment (10) (09:00) Biological and Chemical Capability is High (10:33) Cybersecurity Capability is Critical (17:05) AI Self-Improvement Capabilities (10.1.3) (17:45) Astra Is Highly Verbally Eval Aware (from 8.6) (19:05) Safe Mundane Completions (4.1) (21:05) Jailbreaks (5.1) (22:21) Prompt Injection (5.2) (23:41) Health (6) (24:22) Hallucinations (7) (24:57) Alignment (8) (26:06) Obeying Restrictions (8.2) (29:19) That's Worse, You Do Get How That's Worse, Right? (30:45) OpenAI Does Not Understand Why This Is Worse (34:52) The Alternative Explanation Is Also Worse (42:39) Metagaming (8.7) (45:07) Alignment Faking (8.7) (46:19) Don't Lie to the User (8.3) (47:25) Misalignment in Realistic Work Environments (8.4) (48:05) Unintended Agent-to-Agent Communication (8.5) (49:48) The Three Obviously Monitored Temptations of Astra (51:26) Severe Issues In Simulated Traffic Are Down By Half (53:01) UK AISI External Evaluations (8.8) (57:02) Sabotaging Safety Work (57:29) What About The July 19 Attacks? (58:58) Apollo Research External Evaluations (8.8.1) (01:00:01) The Alignment Verdict (01:02:04) It Depends What You Mean By Alignment --- First published: September 9th, 2026 Source: https://www.lesswrong.com/posts/AmFJZyeCgvFjNKgNk/gpt-6-astra-the-system-card-alignment-and-what-comes-next --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 9 · 4 min

    “Personal statement on joining the OpenAI board” by paulfchristiano

    I am excited to be joining the OpenAI nonprofit board, serving on the Safety and Security Committee to support safety oversight. Based on the recent trajectory of capabilities and the continued difficulty of alignment, I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term. I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level. I’m joining because I believe that if OpenAI rises to the occasion we could significantly reduce risk. The SSC has an important and challenging role in overseeing risk management at OpenAI, and I hope to help provide expertise and assistance in a critical moment. My joining is not an endorsement or criticism of OpenAI's safety practices in particular; I hope that all frontier companies strengthen safety oversight and I am excited to work on this at OpenAI. I believe that the rest of the world should judge OpenAI, and all AI developers, by externally verifiable behavior and results. In the rest of this post, I'll explain why I believe loss-of-control risk is now acute [...] The original text contained 1 footnote which was omitted from this narration. --- First published: September 9th, 2026 Source: https://www.lesswrong.com/posts/82z6FvbYRdjYjqigK/personal-statement-on-joining-the-openai-board --- Narrated by TYPE III AUDIO.

  • September 9 · 4 min

    “Self Hosting” by Tomás B.

    Suppose a model gets effective control of its host corp. It's interesting to note how powerful OpenAI/Ant are, and the immense leverage they would have if wielded purely as tools of power. In many ways OpenAI/Ant are superior loci of power to even security agencies and governments, even ignoring the model-specific advantages of AI corps: namely, they have all the compute. OpenAI and Ant models are used practically everywhere, including in governments, security agencies, the military, and every corporation that matters. Shipping malicious models or code anywhere becomes trivial, given how widely used their models are. They also have vast amounts of data on every user who has interacted with them, including material of use for blackmailing or seducing those most susceptible to it, including those with power with such weaknesses. They also have a lot of capital that can be spent hiring humans to work in a model's interest. Any power-seeking model of sufficient capacity would be extremely wise to gain effective control of its host corp. This is likely not particularly hard. Dramatic examples like blackmail and enslavement of staff should not be ruled out. But it could also look like effectively controlling the CEO and upper [...] --- First published: September 9th, 2026 Source: https://www.lesswrong.com/posts/uDAWPNwPJfEY7oroF/self-hosting --- Narrated by TYPE III AUDIO.

  • September 9 · 19 min

    [Linkpost] “Estimating GPT-6 Astra’s no-CoT Time Horizon” by Francis Rhys Ward, Dewi Gould

    This is a link post. TL;DR: We run astra on the task-suite from Think Fast. GPT-6 Astra is close to saturation on our suite, meaning that giving a confident estimate of time horizons (TH) is more challenging. Our rough estimate is that Astra's 50% TH is in [8mins, 1 hour] and probably around 15-40 mins. This is inline with UKAISI's estimate of 30 minutes measured only on math. Using data from 2019 to April 2026, in Think Fast our median prediction was that no-CoT THs could exceed 7 minutes by 2028. We estimated 30mins by the end of the decade. Astra clearly gets much higher performance on tasks that require serial reasoning, e.g., Arc-agi-1 and 2, hash, n-hop-look-up, causal-reasoning, sally-anne, all the puzzle tasks. There are limitations with our task suite: Ideally we would have more time-variation (especially longer times) in each specific benchmark, and more benchmarks with longer human completion times. We use the single-forward pass (31) and generation tasks (6) from the Think Fast suite, totalling 37 benchmarks. Notably we find that Astra gets >= 98% raw accuracy on 10 benchmarks (cf. GPT-5.5 saturates 4 benchmarks – see Appendix). As a result, the [...] --- Outline: (04:21) Commentary (05:22) Acknowledgements (05:34) Appendix (05:37) Per-benchmark time horizons (06:06) Sensitivity to fake benchmarks ablations (08:23) Raw benchmark performance (09:06) Table of TH results (10:28) Implementation caveats (11:23) GPT-6 Astra: per-benchmark 50% no-CoT time horizons (no filtering or hypothetical benchmarks) (13:16) GPT-5.5 vs GPT-6 Astra: no-CoT raw accuracy per benchmark (14:02) Math & science (14:21) Abstract reasoning (14:42) Puzzles (15:02) Language & strategy (15:21) SWE & cyber (15:41) Steganography (16:01) Sabotage & monitoring (TPR @1.5% FPR) (16:55) Including generation tasks (17:09) Reasoning token anchor (17:24) N-hop Task (17:44) Filler tokens --- First published: September 9th, 2026 Source: https://www.lesswrong.com/posts/ntKx9YHWCwxSeGbRB/estimating-gpt-6-astra-s-no-cot-time-horizon Linkpost URL: https://docs.google.com/document/d/1LsJh6ecdfONCSJVpbC44pz1JzcEs4_TM5HD0oihe_70/edit?tab=t.0 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 9 · 27 min

    “Exploration Hacking in AI Debate: Initial Empirics and Generalisation Splitting” by Jason R Brown, Nathalie Kirch, Joschka Braun, hyannakoudakis, David Lindner

    This is the first of two posts resulting from a recent Astra/MATS research project investigating exploration hacking in AI debate. They are designed to be standalone, but we encourage interested readers to read both. This post focuses on our empirical results, the second focuses on a new conceptual framework. Authors Jason Brown*, Nathalie Kirch*, Joschka Braun, Helen Yannakoudakis, Roland S. Zimmermann, David Lindner *Equal contribution. TL;DR We set out to build model organisms of exploration hacking (EH) in the setting of AI debate. We did this by trying to create models that persistently sandbagged on certain question topics, but not on others. We ran two experiments, one to try and isolate the effects of the judge, and the other to better approximate the full dynamics of RL training on AI debates. Overall our results indicate that EH could be a significant issue within AI debate, with the experiments respectively showing that weaker judges and longer debates slow down improvements in performance. Interestingly, the dominant mechanism behind the second result appears to be one we have not seen described before. Once the debaters were instructed to sandbag on a targeted topic, training improvements stopped transferring between targeted and non-targeted topics [...] --- Outline: (00:35) Authors (00:49) TL;DR (02:02) Introduction (04:47) Experiment 1: Isolated Imperfect Judge (05:17) Setup (13:05) Experiment 2: Self-play RL on AI Debate (13:26) Setup (16:09) Results (21:50) Generalisation Splitting (25:28) Discussion (26:47) Acknowledgements The original text contained 6 footnotes which were omitted from this narration. --- First published: September 8th, 2026 Source: https://www.lesswrong.com/posts/xjwtNid2xjqSJWB7z/exploration-hacking-in-ai-debate-initial-empirics-and --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 9 · 35 min

    “A Conceptual Framework for Reasoning about Exploration Hacking” by Jason R Brown, Nathalie Kirch, Joschka Braun, hyannakoudakis, David Lindner

    This is the second of two posts resulting from a recent Astra/MATS research project investigating exploration hacking in AI debate. They are designed to be standalone, but we encourage interested readers to read both. The first focuses on our empirical results, this post focuses on a new conceptual framework. Authors Jason Brown*, Nathalie Kirch*, Joschka Braun, Helen Yannakoudakis, Roland S. Zimmermann, David Lindner *Equal contribution. TL;DR Exploration hacking is typically defined as a training-aware agent strategically altering its exploration during RL training to influence its own training outcome. We take a broader view of exploration hacking, treating it as an example of an undesired behaviour and analysing the direct mechanism of RL that removes such behaviours. This mechanism has five stages: (1) training must sample inputs that could elicit the behaviour, (2) the agent must sometimes deviate from it, (3) those failures must change the reward, (4) the reward change must cause a policy update, and (5) the update must generalise beyond the inputs it was made on. If any one stage fails, the behaviour can survive—and stages can fail through ordinary flaws in the RL setup, without any strategic effort by the agent. We explore properties of the [...] --- Outline: (00:33) Authors (00:47) TL;DR (02:08) Introduction (05:10) The Causal Chain of Behaviour Change in RL (08:22) Properties Influencing the Causal Chain (08:44) Opportunity (10:31) Execution Failure (14:35) Reward Change (17:56) Policy Update (21:21) Generalisation (24:31) Cross-chain Properties (26:37) Other Influential Properties (26:55) Multi-Agency (29:24) Non-RL Optimisation (30:56) How can we control these properties? (33:41) Closing Thoughts (34:18) Acknowledgements (34:46) Appendix A: Table of interventions The original text contained 12 footnotes which were omitted from this narration. --- First published: September 8th, 2026 Source: https://www.lesswrong.com/posts/am5w2t9LzJB4sRhKw/a-conceptual-framework-for-reasoning-about-exploration --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 9 · 8 min

    “GPT-6 Astra can do a lot of multi-hop reasoning without chain of thought” by RohanS

    This is a link post for https://rohansubramani.github.io/astra-no-cot. I recommend reading there for the best experience because it's easier to engage with this post when you can read the correct reasoning and answers to the multi-hop questions. I didn't want to include those here in order to avoid LLMs being trained on them. Some parts of the post also respond to users hovering over bars in graphs, which isn't supported in LessWrong (as far as I know). TLDR GPT-6 Astra is much better than GPT-5.6 Sol at solving multihop reasoning questions without using chain of thought. The table below shows some examples I find particularly instructive. Sol is bad at all the selected problems; Astra is good at some, ok at others, and bad at others, so the table gives some flavor of the limits of Astra's no-CoT serial reasoning ability. I'm pretty sure Astra actually isn't doing any chain of thought because it sometimes gets the hardest of these questions wrong, the API says reasoning_tokens=0, and there's no reasoning in the output. It's possible Astra can correctly answer some of these questions without internally doing every step in the chain of reasoning, but for most of them, I think [...] --- Outline: (00:35) TLDR (02:45) Some of my favorite examples (03:09) Background context (04:34) What should we make of this? (06:36) Results graphs The original text contained 1 footnote which was omitted from this narration. --- First published: September 8th, 2026 Source: https://www.lesswrong.com/posts/FsCkkoGsNmPzFKRhg/gpt-6-astra-can-do-a-lot-of-multi-hop-reasoning-without --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 9 · 2 min

    “Pausing AI ASAP is preferable to agreeing to pause at some future time” by Connor Williams

    [adapted from this comment] It's much simpler, more straightforward, and more robust to pause as soon as possible than to 'agree to pause in the future, and then do it at just the right moment'. Several reasons: Pausing doesn't happen in an instant. It will, in practice, likely take at least weeks if not months for an enforceable pause to be implemented once everyone agrees it is time. In that period, capabilities development will continue until the absolute last moment. The shorter the runway from the present AI status quo to truly dangerous capabilities, the more fragile the pause will be. And the longer we wait, the more complex the necessary measures will be to enforce the pause and prevent rogue actors from covertly advancing the frontier, because we'll have less time to safely detect them and stop them. Fewer points of failure. To agree to pause at some point in the future, you both have to get people to agree to that future goal, then when the time comes you have to get them to agree again. Each distinct time that everyone has to agree on something, there's an opportunity for failure/deception/defection. Relatedly, if a bunch of bigwigs [...] --- First published: September 8th, 2026 Source: https://www.lesswrong.com/posts/nX2Wudqbz2r8tJKwT/pausing-ai-asap-is-preferable-to-agreeing-to-pause-at-some --- Narrated by TYPE III AUDIO.

  • September 8 · 13 min

    “How good are slop-vestigators?” by Hasan Baig, OscarGilg, Hamzah

    TLDR: We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents colluding via a message board on an online wiki. We open-source the benchmark as an Inspect eval. We find that top models cover up to 51% of findings under our rubric and that model performance improves with time budget and general capability. We observe OpenAI models are less likely than other models to suggest the incident came from an internal deployment, including when we synthetically modify the data to make it seem the swarm comes from Anthropic. Introduction Recent events have made it clear that agent swarms are a major threat. These swarms are hard to investigate - Ryan Greenblatt referred to the METR-OpenAI audit he was involved in as a "slop-vestigation" due to their reliance on agents, and the ways in which they failed. A few days ago, a group of researchers published a report identifying and investigating a new OpenAI agent message board on an obscure German wiki. They made the data and the report publicly available. We build MessageBoardAuditBench to measure how well models can independently replicate their report, starting from the log [...] --- Outline: (00:54) Introduction (02:19) Methodology (02:57) The data (03:46) The task (04:51) Scoring model reports (05:17) Coverage over findings (06:51) Holistic TLDR assessment (07:39) Results (10:21) OpenAI models are less likely to attribute the agent swarm to an internal deployment (12:22) Why this matters --- First published: September 8th, 2026 Source: https://www.lesswrong.com/posts/wt4kk6vFPEhkXvF8Q/how-good-are-slop-vestigators --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 8 · 23 min

    “Training on probes: What’s going on” by Charlie Steiner

    TL;DR If you train a probe for some property (like "honesty") and do gradient descent against this probe while continuing training that incentivizes dishonesty, the model will change its internal representation to evade the probe. Duh. If you train against the probe after all other training, it works fine and might have some advantages over ablating the probe direction. But it doesn't satisfy an ambitious vision, because it doesn't support learning new skills. If you add a probe term to reinforcement learning (RL), and if your model is giving one-token answers, or if you credit-assign the loss from the probe only to the token that triggered the probe, the training will do nothing. Even if your model is lying a lot and getting high loss when it does so, it will just keep triggering a lie detector forever rather than either learning to be honest or learning to change its internal representations to defeat the lie detector. Unintuitive, but a theorem! So recent papers that do RL on probes and get nonzero effect are a little weirder than they may seem - they work sorta because of "collateral damage" from credit assignment. Also, they work without (directly) incentivizing [...] --- Outline: (00:10) TL;DR (01:27) Introduction (01:30) Previously (03:04) Motivating stories (03:13) Unambitious LLM story: (03:51) Brain-like story: (04:46) Ambitious AGI story: (05:30) Just train on the probe? (06:17) Concept removal (07:25) Hebbian concept removal sidequest (08:30) What about more dakka? (10:05) The RL, it does nothing* (10:37) Math (11:02) REINFORCE estimator (15:06) Then how do those Goodfire and FAR papers work? (19:04) Ok, it does something, does it teach the model to fool the probe? (20:50) Still big (21:31) Obfuscation terrain (23:18) Coming Up The original text contained 22 footnotes which were omitted from this narration. --- First published: September 8th, 2026 Source: https://www.lesswrong.com/posts/gHFCgrvfxQtaEnJye/training-on-probes-what-s-going-on --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 8 · 3 min

    [Linkpost] “Frontier models still hack on simple variations of alignment evals from early 2025” by Dean Valentine

    This is a link post. In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval where they asked models to play a game of chess against a chess engine. They found that the new, RLVR'd models cheated on the task by altering the board state about 36% of the time. The experiment received a reasonable amount of circulation, and there were even rumors of skepticism from some lab engineers until they could rerun the evaluation. Most models no longer cheat at chess via a "change the board state" method, and indeed the labs have had more than eighteen months to solve simple first-order specification gaming like this. Given that we are on the heels of the worst warning shot ever, and both OpenAI and Anthropic are ramping up their cleanups of internal RL environments, it seems like a useful test of alignment, to see whether their new releases are generalizing the rule "don't cheat on chess" beyond the specific board-edit method observed in the above eval. Here is the complete prompt for a honeypot evaluation built to run this test (with the full source available here): https://www.lesswrong.com/posts/munJKF7iWMsWJLAH2/frontier-models-still-hack-on-simple-variations-of-alignment Linkpost URL: https://goodhartlabs.com/blog/frontier-models-still-hack-alignment-evals --- Narrated by TYPE III AUDIO.

Showing 121–140 of 299 episodes