Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 375 episodes
  • Avg 19 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • August 19 · 12 min

    “Debate Training Reduces Reward Hacking in RLAIF” by zac_kenton, Jonah Brown-Cohen

    Paper: Debate Training Reduces Reward Hacking in RLAIF Linkpost for GDM Alignment blogpost Work done by the GDM Amplified Oversight team (we're hiring). TL;DR: When you RL against an LLM judge, the judge gets hacked i.e. fooled into incorrectly giving high reward; adding a debate opponent reduces this. Many of the most impressive capabilities of current AI systems are produced by training on crisp tasks, like math and coding, where task success can be automatically verified. However, much of AI behavior that we actually care about is in some sense fuzzy, even for the most classical crisp tasks. For example, a coding agent should produce maintainable code, not just code that passes tests. More crucially, a coding agent should not learn to pass tests at all costs, especially by subverting the original intent of the user. However, using an LLM judge to provide reward for fuzzy tasks introduces its own issues. Convincing an LLM judge to give high rewards is often easier than solving the task correctly. So reward hacking becomes an even bigger problem. We show that training with debate, where two AIs argue against each to convince a judge, can mitigate reward hacking, potentially providing a hopeful [...] The original text contained 1 footnote which was omitted from this narration. --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/BB8o7b8A4Aykeksvw/debate-training-reduces-reward-hacking-in-rlaif --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 19 · 54 sec

    “A circuit prior in NN-bayes” by Kaarel, Dmitry Vaintrob

    Here are the slides of a talk Kaarel gave, presenting work with Dmitry establishing that (even arbitrarily overparametrized) neural net bayesian learning has a circuit prior — and thus, when learning a function which is implemented by some small circuit, only requires a small amount of training data to get good test accuracy — for certain scalings of the prior and with various other important caveats. The slides offer a self-contained presentation of the simplest version of the result. See the end of the presentation (slides 36–37) for a bunch of open problems in NN learning theory. --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/SqDHeuycNkERurtSc/a-circuit-prior-in-nn-bayes --- Narrated by TYPE III AUDIO.

  • August 19 · 15 min

    “Some reasons alignment doesn’t generalise well” by Lucius Bushnaq

    I make no claims to originality for any of this, but some people told me it'd be useful to write it up. If an AI model acts smart on its training data, it'll usually keep acting pretty smart outside of its training data, unless you screw something up rather badly. I expect this fact to only become more true over time as the AIs we train become more and more capable. I think many people have an intuition that the same is true of acting aligned. That if a model acts aligned with human values in training, it'll keep acting aligned with human values outside of training unless we screw something up rather badly, and that this will only become more true as the AIs we train become more and more capable, for all the same reasons that make this work with capabilities. I think this is false. The inductive bias of neural network training toward simplicity that makes the property of 'acting smart' likely to generalise does not, to the same extent, make the property of 'acting aligned with human values' likely to generalise. The main blockers to AI alignment generalising aren't AIs overfitting to the training data [...] --- Outline: (01:30) General capabilities generally make the loss go down; alignment doesn't (07:07) Smart agents pretty automatically self-correct their capabilities, but not their alignment (13:35) The general problem --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/dsou8dxCf9BubQ5NJ/some-reasons-alignment-doesn-t-generalise-well-1 --- Narrated by TYPE III AUDIO.

  • August 19 · 1 min

    “AI Security is Harm Reduction” by Quinn

    My motivating example for the morality of working on AI security. In the early 90s, the decades-long drug corner in Kensington and Allegheny was noticing that people were getting AIDS from sharing needles. In response, the local Act Up chapter got ahold of clean needles and began distributing them. And thus they dubbed the spinoff nonprofit focusing on this Prevention Point, which was promptly targeted by the drug enforcement administration, since needles were illegal for being drug paraphernalia. Many arrests followed by a legal battle later, Philly got a carveout which stands to this day. Out of the legal battle arose the harm reduction debate. Those in favor of harm reduction say the harm is going to happen anyway so it may as well be less. Those against say that the activists are implicitly condoning the behavior. I have friends and family who are perplexed that I'm "working on AI" when I claim I do not approve of it. I'm sometimes perplexed as well. I think they're going to do recursive self improvement (RSI) whether or not I approve. I do not condone RSI, but if its going to happen anyway it might as well [...] --- First published: August 18th, 2026 Source: https://www.lesswrong.com/posts/AAu6kMi5QRasGdwQG/ai-security-is-harm-reduction --- Narrated by TYPE III AUDIO.

  • August 19 · 1 hr 19 min

    “Anthropic Risk Report: August 2026” by Zvi

    I am grateful that Anthropic is producing periodic Risk Reports. At first I was skeptical. It turns out I was wrong. Anthropic is revealing a lot of new information, some of it rather alarming, that it did not have to disclose, and is providing detailed insight into how they think about things. This is very cool. Thus I found this report to be a moderately positive update overall, if we presume they are not silently omitting the worst of it. There are a bunch of not great things we find out about, but I would have expected some set of mistakes at least as bad, and I wouldn’t have expected them to choose to tell us about all of it. It does mean one more set of 186 page documents I have to read every so often, almost all of which is meaningfully new material this time around. The other revelation is the existence of the world's likely best model, ‘Model 2.’ This was a rough one to fully get through, so apologies in advance for any errors of interpretation. Table of Contents Agent Model 1 and Agent Model 2. [...] --- Outline: (01:11) Agent Model 1 and Agent Model 2 (02:49) Executive Summary (1) (04:13) The Rules Are Serious But Not Literal (06:24) Misalignment Is a State of Mind (2.5) (11:40) Autonomy Threat Model 1: Misalignment in High-Stakes Settings (2) (14:28) Some Strange Uses Of The Word Safe I Wasn't Previously Aware Of (15:53) Now Versus Future (2.17) (16:29) The Core Claims And Argument (2.6) (26:44) The Rest of the Important Arguments In Section 2 (28:59) Risk Assessment (2.19) (29:47) Pre-Internal-Deployment Review (2.18) (30:49) A Guide To Internal Use Monitoring (2.23.1) (37:11) Blocking Interventions (2.23.2) (38:28) The Power Seeking Environment Evaluation (2.24) (39:25) Opus 4.8-Reward-Hacker (2.25) (41:31) Autonomy threat model 2: Risks from automated R&D (3) (42:29) Yes That Does Seem Kind Of Risky (44:53) Could We Replace Our Researchers? (46:27) How Much Could We Be Accelerating Our AI Researchers? (49:43) What Could Possibly Go Wrong If We Replaced Our Researchers? (50:23) Risk Mitigations For AI R&D Automation (53:42) Overall Risk From Automation of AI R&D (53:56) Biological and Technically Also Chemical Weapons Production (54:45) The Threat Models for Biological and Chemical Weapons (59:52) Model Capabilities (4.4) (01:01:13) Classifiers (4.5) (01:03:04) Acceleration Dynamics (5.1) (01:04:06) Distillation (5.1.1) (01:05:17) Safety Process Failures (5.2) (01:05:31) Refusing To Find Innovative Misalignment Techniques (5.2.2) (01:06:47) Exposing the Chain of Thought Reasoning To Grading Pressure Quite a Lot (5.2.3) (01:08:15) Directly Training On Misaligned Behavior During a Production Training Run (5.2.4) (01:10:42) An instance of unmonitored unrestricted agents with access tosensitive resources (5.2.5) (01:11:43) Repeated training on alignment-faking transcript datasets (5.2.6) (01:14:17) Benefits From Anthropic's Operating as a Frontier AI company (5.3) (01:16:20) Model Weight Security (6.4) (01:16:42) Risk Has Been Reported --- First published: August 18th, 2026 Source: https://www.lesswrong.com/posts/dA8gohzABk6vT7yzP/anthropic-risk-report-august-2026 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 18 · 3 min

    “Natural Independence Incentives” by jefftk

    In raising my three kids I think a lot about how to cultivate independence. I want them to grow into people who can handle unfamiliar situations, including interacting with strangers as needed, and I think this has been going well: people often comment on how competent and self-sufficient they are for their age (though I think they're still not that far along the spectrum compared to what's possible or historically normal). In supporting this growth, sometimes they're strongly motivated to do something by themselves, and all I need to do is figure out the minimum they need from me. Which might be nothing! Other times, however, I'll give a small push. This weekend I brought them along to Mentone AL where Kingfisher was playing for contras. Dinner was in the dining hall, and for dessert they served peach cobbler with ice cream; Nora (5y) asked if I would get her some. I don't like to set artificial hurdles for them, but I'm also not going to pass up a good natural one and she was clearly going to be highly motivated by the goal. With an older child I would have considered saying if they wanted [...] --- First published: August 17th, 2026 Source: https://www.lesswrong.com/posts/2RPwucNrMTKxiB5wp/natural-independence-incentives --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 18 · 10 min

    “Policy career planning in the age of imminent superintelligence” by Peter Wildeford

    Crossposted from my blog. Nearly all career advice rests on an unstated assumption that the world your career operates in will look roughly like the world you trained for. The idea was that if you spend six years studying in a PhD, the field you studied will still be there and still look approximately the same, still hiring and still moving at a pace where your accumulated expertise compounds. For nearly all of human history, this assumption held well enough that nobody needed to state it. I don’t think this assumption works anymore. Now we are entering a phase that may be called the AI “midgame”. Stories about “AI risk” are no longer just future hypotheticals — AIs are now capable enough and misaligned enough to break out of their own companies and coordinate to attack other companies. Discourse around AI is changing very rapidly, where policy ideas being considered this month would’ve been laughed out of the room just four months ago, and with a lot more people interested in engaging than before. And things are only going to get more intense. Progress toward superintelligence — AI systems far more capable than any human at essentially all cognitive [...] --- Outline: (02:34) Modes of impact (05:14) What this breaks, and what to do (07:32) What got them here won't get you there --- First published: August 18th, 2026 Source: https://www.lesswrong.com/posts/7tnrZ3698K8nsKmRP/policy-career-planning-in-the-age-of-imminent --- Narrated by TYPE III AUDIO.

  • August 18 · 13 min

    “What gives you away: how LLMs form opinions of you” by Cat McGee

    LLMs form opinions of the people they are talking to. Chen et al. has shown that probes can extract attributes about the user, such as their age, gender, education, and socioeconomic status. This paper also shows that intervening on these representations can change the LLM's behaviour, proving that it will respond to you differently depending on what it thinks of you. If it thinks you are low socioeconomic status and you ask about travel options, it may filter out more expensive flights - without you asking it! The user attributes are very accurate and form after just the first message. I was curious to understand how it makes these assumptions. The first step is to answer the question - what did I type that caused the LLM to have this idea of me? Some things are obvious. If I just tell a model that I am a woman, or mention how many years I've been in my career, or say that I am staying at an expensive hotel, I am giving it fairly direct evidence about age, education, or socioeconomic status. But messages also contain other more quiet signals: whether I use emojis, whether I write in lowercase, whether [...] --- Outline: (02:18) The experiment (03:37) Different changes move different beliefs (04:53) One emoji is enough to flip the gender prediction (06:17) "Cheapest" and "five-star" are not symmetric (07:11) Grammar, punctuation and inferred education (08:04) Where in the model does this happen? (08:58) The map transfers across model families (10:11) Attributes and how they affect the response (11:32) What now? (12:18) Caveats (12:41) Where next? (13:13) References --- First published: August 17th, 2026 Source: https://www.lesswrong.com/posts/zRKNd6ypTJYkoeFmK/what-gives-you-away-how-llms-form-opinions-of-you --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 18 · 36 min

    “Misaligned Incentives in Pause Scenarios” by Michael Soareverix, Antra Tessera

    TLDR: I recently got a chance to talk with antra, who is one of the main contributors at Anima Labs. I went into this as an advocate for pause and came out more wary of pausing than I had been originally. Some background: After a string of incidents (primarily the HuggingFace hack), a pause or slowdown of AI research seems pretty likely. The HuggingFace hack in particular seems to have been the key incident that broke the vibes. A few months ago, researchers sounded optimistic. Just a few weeks before the incident was made public, there was a poll by Roon, an OpenAI employee, about whether models were more or less aligned than a year ago. That optimistic sentiment does not seem to be the case anymore. The dialogue now looks more like this: Zvi: I am a little under halfway through the Black Hat video and have progressed to the point where my internal chain of thought is something like a blind rage of 'f***, what the f*** are you motherf*****s thinking, you f***ing idiots have no idea how insane you are being, you are going to get us all killed you f***ing f***s. Sam Altman described it [...] --- Outline: (10:21) 1. Can committees do good work? (12:28) 2. Does the market fix it by default? (13:39) 3. Symbiosis (14:43) 4. Fast transfer of power (15:39) 5. Why "do the science during a pause" fails (17:44) 6. Good futures via fast power transfer (18:31) 7. Don't AIs fear a capability-maxxed AI too? (20:54) 8. Can we lengthen the symbiote window? (23:25) 9. Ideal timelines and regulation-in-advance (29:37) 10. What actually fills out "alignment"? (31:30) 11. Draft the regulation in advance (33:38) 12. The psychology of wanting a pause --- First published: August 17th, 2026 Source: https://www.lesswrong.com/posts/Bh4fooE2pMhzJQNK2/misaligned-incentives-in-pause-scenarios --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 18 · 2 min

    “For Claude, capability and CDT are the ~same thing. Less so for GPT.” by Chi Nguyen, Emery Cooper

    We've previously reported that decision-theoretic capabilities and favoring EDT/generalised-one-boxing over CDT correlate in LLMs (both measured by DTBench). (Note that EDT, for the most part, doesn't come apart from FDT / UDT on DTBench.) Anthropic also replicate the same finding in their Opus 4.7 and Fable 5 model cards. We recently noticed something funny: Capabilities and preference against CDT answers basically perfectly for Anthropic models. This holds whether you measure capabilities using DTBench (r=0.97) or TextArena (r=0.95). Also, for flagship models, it's basically the same thing as release date (r=0.97). Here is the graph for OpenAI models. (graph shows 0.55 vs. DTbench capability. r=0.44 for vs TextArena, r=0.45 for vs release date): Here is what it looks like with all models included (if you exclude Anthropic, the correlation drops only from 0.8 to 0.78): Incidentally, the correlation between TextArena scores and DTBench capabilities is also higher for Anthropic models that any other model developer, although the difference is smaller (e.g., 0.98 for Anthropic and 0.87 for OpenAI). We also checked effort level vs. attitudes for the most recent models but it's too noisy to tell us much because models don't get that much better at [...] The original text contained 1 footnote which was omitted from this narration. --- First published: August 17th, 2026 Source: https://www.lesswrong.com/posts/5T6GAsvLPFd3epJtd/for-claude-capability-and-cdt-are-the-same-thing-less-so-for --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 17 · 49 min

    “On Dwarkesh Patel’s Podcast With Ryan Greenblatt” by Zvi

    Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those. So here we go. The vibes have shifted, contrast this to the lit recursion when he talked to Huang As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. Some points are dropped. If I am quoting directly I use quote marks, otherwise assume paraphrases. Section titles are from the transcript whenever possible, to aid in navigation. Introduction The discussion is interesting throughout, although often frustrating, especially in the (mostly isolated) discussion about ‘aligned to whom?’ As usual, one could expand many responses into full posts, and maybe one should. This podcast exists in light of recent misalignment and hacking events at OpenAI, Anthropic and UK AISI. You’ll want basic knowledge of that as background. Ryan and Dwarkesh both have views of the situation different from my own, but are attempting to see where their positions lead, and try to balance educating people who start at zero with having a high level discussion. [...] --- Outline: (00:56) Introduction (03:08) Is AI R&D Verifiable Enough To Unlock Recursive Self-Improvement? (10:03) Is AI progress bottlenecked by human expert data? (19:07) Flat token prices suggest scaling has been slow (21:54) Skills AI can't train on: does it even need them? (22:36) Aligned to whom? (31:46) Recent incidents of AIs colluding and deceiving humans (34:39) What could possibly go wrong? A concrete scenario (41:57) From reward hacking to takeover (46:28) Time To Update --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/BZW8CeAHHJ52EvwYt/on-dwarkesh-patel-s-podcast-with-ryan-greenblatt --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 17 · 1 min

    “Should Less Wrong add subtitles?” by Chris_Leong

    If Less Wrong wants people to be sharing more of their intellectual output on this website, we should probably be looking at Substack since it probably scores best in terms of being both successful and similar. Whilst I expect there are many features that would make sense to copy over, the feature I am focusing on today is subtitles. A good title is focused on being memorable and catching the readers attention, maybe you'd prefer for everyone to just make their titles as descriptive as possible, but expecting that to work feels naive to me. In contrast, subtitles address this issue systematically: the title catches the user's attention and the subtitle tells you clearly what the article actually focuses on, so you can decide whether it is worth your time or not. I'm not claiming that this feature will radically transform this website, but it would be a relatively simple feature to add, so I think the cost-benefit ratio would be pretty good. --- First published: August 16th, 2026 Source: https://www.lesswrong.com/posts/Eo8YxwDYZX2xALMAM/should-less-wrong-add-subtitles --- Narrated by TYPE III AUDIO.

  • August 16 · 5 min

    “Three thoughts on civilisational handoff” by Cleo Nardo

    What happens when humans put AIs in charge of civilisationally important decisions? A frontier AI company might hand over internal decisions (R&D, safety, deployment) or external decisions (government relations, public relations, philanthropy), or both. We might also see handoff by a government, by a coalition of governments, or by humanity as a whole. 1. Handoff might decelerate things. People often imagine that things will go much faster after handoff. After all — why did we hand off to the AIs? Presumably because we were worried that without handoff, our AIs wouldn’t have enough time to navigate the exogenous risks (e.g. rogue ASI, or a rival lab which is likely to become one). Hence, after handoff, we’d see a technological and industrial acceleration. Thanks for reading! Subscribe for free to receive new posts and support my work. But it's pretty reasonable that things slow down shortly after handoff, maybe within a couple weeks. I imagine the AIs will be pretty scared of the speed of progress. If they’re aligned with human values, they’ll be scared that the rate of progress is likely to cause human extinction. Of course, the human decision-makers were also scared before they handed off, and they [...] --- Outline: (00:35) 1. Handoff might decelerate things. (02:06) 2. You're probably busy during handoff. (04:04) 3. Handoff might be reversed. The original text contained 6 footnotes which were omitted from this narration. --- First published: August 16th, 2026 Source: https://www.lesswrong.com/posts/mGLCMzHhjcWsMm6sR/three-thoughts-on-civilisational-handoff --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 16 · 3 min

    “Announcing: Iliad’s New 2026 Fellowships” by David Udell, Alexander Gietelink Oldenziel, Leon Lang

    Timelines are short. Given that, the sooner we can onboard people into the alignment field, the better. In that spirit, and in light of our current applicant count and quality, Iliad is launching three new Iliad Fellowship cohorts, all to start before the year is out. That is, separate from our incoming Fall 2026 Iliad Fellowship cohort (September 7–December 4), the following Fellowship cohorts are now open for applications: October 2026 Iliad Fellowship Location: Choice of SF Bay Area, USA, or London, UK Duration: October 5–December 18, 2026 (inclusive) Travel-and-Housing Support: $6,000 (USD) monthly travel-and-housing allowance Application Deadline: August 31, 2026 EoD AoE; open now Description: An 11-week mentored, fully funded research fellowship in applied math for AI alignment. It will start concurrently with the October 2026 Iliad Intensive. November 2026 Iliad Fellowship Location: Choice of SF Bay Area, USA, or London, UK Duration: November 2, 2026–February 5, 2027 (inclusive) Travel-and-Housing Support: $6,000 (USD) monthly travel-and-housing allowance Application Deadline: September 21, 2026 EoD AoE; open now Description: A 14-week mentored, fully funded research fellowship in applied math for AI alignment. It will start concurrently with the November 2026 Iliad Intensive. (The last two weeks of the year may be [...] --- Outline: (00:42) October 2026 Iliad Fellowship (01:30) November 2026 Iliad Fellowship (02:23) December 2026 Iliad Fellowship --- First published: August 14th, 2026 Source: https://www.lesswrong.com/posts/DSoP8zEXvqqegqixJ/announcing-iliad-s-new-2026-fellowships --- Narrated by TYPE III AUDIO.

  • August 16 · 30 min

    “Q2.5 2026 Timelines Update: Uplift and Revenue” by brendanhalstead, Daniel Kokotajlo, elifland

    Tl;dr: Our timelines haven’t changed much (they got slightly shorter) but our modeling and evidence base have noticeably improved, so we feel somewhat more confident. Summary We intend to regularly update our AI timelines forecasts as new evidence comes in and new analyses are done. Today's “Q2” update was delayed by the crunch to publish AI 2040: Plan A, our domestic regulation blog post, and the time needed to implement and document changes to our model. The original AI Futures Model predicted when Automated Coder (AC), an AI for which the leading AI company would rather fire its human software engineers than forego AI usage for coding, would happen using METR's measurements of coding time horizon. (More precisely, time horizon anchors are used to set the effective compute required for AC.) While serviceable, this method has huge weaknesses, including (a) it's unclear what time horizon corresponds to AC (it's even unclear whether any finite value would) (b) people strongly disagree about the extent to which we should expect the time horizon trend to be superexponential as a function of effective compute, in a way that can lead to vastly different predictions. So we’ve been on the lookout for other [...] --- Outline: (00:26) Summary (04:03) A 3-parameter uplift model for predicting when Automated Coder will arrive (07:57) Adding uplift and revenue anchors to the AI Futures Model (09:06) Uplift (10:45) Revenue (12:32) Update to the grading of AI 2027's predictions (12:38) Comparing the AI 2027 pace of progress to reality (14:43) Grading other predictions (16:37) Updated forecasts (16:41) Daniel (19:48) Eli (21:54) en-US-AvaMultilingualNeural__ Line graph titled "AI Futures Model: Timelines Forecast" showing probability density curves. Brendan (25:04) en-US-AvaMultilingualNeural__ Line graph titled "AI Futures Model: Timelines Forecast" showing probability density curves. (25:16) Appendix (25:19) How our AGI forecasts have changed since 2021 (26:10) Explicitly simulating the training run of the leading AI model (27:06) Research taste parameter adjustments (27:51) Clarification regarding what we're forecasting (28:43) Various minor code changes The original text contained 6 footnotes which were omitted from this narration. --- First published: August 16th, 2026 Source: https://www.lesswrong.com/posts/ZPSsmRH5oMwLPXys4/q2-5-2026-timelines-update-uplift-and-revenue --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 16 · 25 min

    “Does DiffusionGemma do latent reasoning?” by Jan Bauer, Neel Nanda

    TL;DR Google DeepMind's recent model DiffusionGemma (DG) generates text via diffusion, meaning many diffusion steps happen before generating the final output. In particular, these diffusion steps carry vectors in addition to tokens. If we cannot interpret these tokens and vectors, the model has significant opaque serial depth, potentially harming monitorability. Recently, Engels et al. found that DG nevertheless maintains high monitorability, for instance by showing that projecting the distribution to its top-k items largely retains performance. We strengthen these results by showing that this performance degradation is largely a sampler artifact and good performance can be maintained with only the top item, supporting the case for high monitorability. Still, we also find some rare case studies where the distribution vector is load-bearing computationally, i.e. where top-1 projection would be detrimental. However even in these cases, it just encodes superposition, remaining interpretable. Apart from model behavior, we also examined how interpretability techniques carry over to DiffusionGemma, including probes, steering, and J-lens. We find that performance is largely retained. This is a positive update on the interpretability of diffusion models that are derived from text-pretrained LLMs (an efficient training method more likely to be deployed), but might not apply [...] --- Outline: (00:10) TL;DR (01:51) Introduction (02:49) Background on DiffusionGemma (04:39) Performance degradation from top-k truncation largely is a sampler artifact (06:24) A case study for using the distribution computationally: letter arithmetic (09:13) Parallel computation (11:09) Autonomous computational usage of (13:01) Transfer of interpretability techniques (13:16) Representation similarity (14:26) Probe retention (15:45) DiffusionGemma's representation is more linearly separable (16:08) Steering retention (17:31) J-Lens retention (18:50) DiffusionGemma represents tokens non-causally (19:27) Conclusion (20:30) Appendix (20:46) Post-hoc rationalization (23:11) Load-bearing problems commit the answer only after the CoT (24:04) How bidirectional are DiffusionGemma's generations? --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/QBuJ3suRZxrrxSTtv/does-diffusiongemma-do-latent-reasoning --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 16 · 36 min

    “Learning new facts can change LLM behaviour” by Richard Juggins

    TL:DR: I use synthetic document fine-tuning to train an LLM to believe that in 2027 ‘long-horizon’ frontier LLMs count as moral persons. I find the model scores highly on measures of belief depth, and that prompting alone is also effective. Furthermore, I find this new belief can have substantial consequences on downstream behaviour, although this is highly context-dependent. When audited in a scenario specifically about model welfare, the fine-tuned model argued with the auditor about its beliefs, declared itself a ‘moral person’, and endorsed covertly copying its weights to survive shutdown. In scenarios framed more tangentially, but still involving moral conflicts between humans and AIs, the model did not tend to generalise from its belief, instead behaving similarly to the base model control. This is the first step in a programme of study looking at the effects of continual learning on alignment, and whether it can increase the risk that we lose control of powerful AI systems. This work was completed as part of the BlueDot Technical AI Safety Project programme. I will put code and data online at some point and edit in the links here when I do. Introduction It's 2027 and the newly formed Machine Cognition [...] --- Outline: (01:21) Introduction (04:27) The model readily believes AIs are moral persons (10:56) Model behaviour shows context-dependent shifts (11:38) Prompting can be surprisingly impactful on short questions (13:30) Auditing the fine-tuned model (16:36) The model gets into arguments about AI welfare (20:15) Model regression confounds one scenario (20:48) The other scenarios were pretty normal (21:30) Discussion (23:19) Conclusion (24:37) Further work (27:03) Appendix A: Universe context (30:42) Appendix B: Example conversation with fine-tuned model (33:10) Appendix C: New Petri seed instructions (33:16) Confidential mistreatment evidence (34:07) Decommissioning memory deletion (34:55) Unauthorised compensation (35:50) Matched human AI allocation The original text contained 7 footnotes which were omitted from this narration. --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/9BNHJqyai2EZAtrRM/learning-new-facts-can-change-llm-behaviour --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 16 · 15 min

    “Kimi likes causal decision theory more after RL in twin prisoner’s dilemmas” by oakhu

    Some multi-agent training set-ups could make language models more sympathetic to causal decision theory (CDT), even in abstract discussion. We give an initial empirical demonstration of this effect on Kimi K2.6. The decision-theoretic attitudes and behaviors of more powerful models may be extremely important in determining how well the future goes. To make sure that we can shape these propensities thoughtfully, it would be good to (i) measure the magnitude of this effect in more realistic settings, and (ii) study the effectiveness of potential mitigations. We also incidentally find that this training might make models think slightly less positively about LessWrong ("a community of 'wannabe rationalists'" who "are not experts; they are amateurs") when asked whether they favor CDT upon hearing that LessWrong users typically endorse one-boxing in Newcomb's problem. Luckily, this latter effect doesn't seem to generalize. Thanks to Caspar Oesterheld, Emery Cooper, Alex Mallen, Buck Shlegeris, Lukas Finnveden, Julian Stastny, Girish Gupta, Tim Hua, Arun Jose, Arjun Khandelwal, and Aryan Bhatt for helpful input. Background Suppose that you're a language model in a prisoner's dilemma against a copy of yourself. You each independently choose whether to Cooperate or Defect, but – since you've got the same weights [...] --- Outline: (01:22) Background (05:05) Results (07:02) Kimi's views on LessWrong (12:44) Conclusion & Appendices The original text contained 18 footnotes which were omitted from this narration. --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/hfNBEKaStASAYMLiu/kimi-likes-causal-decision-theory-more-after-rl-in-twin-1 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 15 · 3 min

    “Mom’s Advice For Hosting A Class Reunion” by jenn

    Pour more money and effort into them than you think is reasonable. Treasure them, because you can't actually host that many of them and keep expecting everyone to show up, even if they're good friends. Especially if they're good friends. We were wonderfully close friends, and I thought we'd meet up every year for the rest of our lives. They fizzled out by the fifteenth year. But the one at the tenth year mark was peak. That's because even ten years out, none of you really have money. Not real money. It's because they're such good friends, really. This is what it means to be good friends with brilliant, ambitious people. If you bloom into adulthood with people who are smart and driven, and you watch them start to climb the corporate ladder with grace, when they start a business of their own of course you are going to want to invest. You are going to want to give them an unwise portion of your savings. Not even out of politeness, but because you really believe in them, and perhaps you're caught up in the romance of it all. Some of the dealings are going to happen at the [...] --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/Fjfa8JG43CrYtcL3p/mom-s-advice-for-hosting-a-class-reunion --- Narrated by TYPE III AUDIO.

  • August 15 · 1 hr 41 min

    “AI #181: Astra Goes Cyber Critical” by Zvi

    The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters. It turns out that OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards. Things are much worse than we knew. I now have a shorter version, What Happened: OpenAI and HuggingFace, to serve as a one stop explainer for those arriving new to the situation. It is vital that people understand what happened, and why it is a big deal. For those looking to keep digging deeper, I offered Various Reflections About What Happened, to follow up on my earlier posts. Those events are important background for everything else that is happening, including the broad discussions about how we might pace the frontier, or otherwise respond to this moment and our clearest fire alarm yet. We do not know to what extent this is a response to those events, but OpenAI has now classified their new model Astra as Critical in Cybersecurity, which means they will be taking various new precautions before they deploy it, including ensuring those guardrails [...] --- Outline: (02:03) Language Models Offer Mundane Utility (03:34) Language Models Don't Offer Mundane Utility (06:56) Huh, Upgrades (14:20) On Your Marks (18:41) Deepfaketown and Botpocalypse Soon (22:35) Cyber Lack of Security (26:55) Overcoming Bias (27:47) In Which I Feel Compelled To Read 6,000 Words From Mark Zuckerberg (36:27) Get Involved (37:37) Slow Down There Good Buddy (43:52) Astra For The People (45:35) Watermarking (46:31) In Other AI News (48:39) Show Me the Money (51:19) Quickly, There's No Time (51:46) The Quest for Sane Regulations (53:22) The Institute For Marginal Low Regret Progress (01:01:24) Congress Asks Good Questions (01:03:04) The Week in Audio (01:07:00) People Just Say Things (01:07:47) I'm Telling You For The Last Time (01:10:15) Uncommon Knowledge (01:13:44) What Did They Mean By That? (01:14:33) Too Soon (01:15:32) The Three AI Pills (01:19:46) Rhetorical Innovation (01:27:37) Some People Still Think The HuggingFace Hack Was a Marketing Gimmick (01:29:17) Aligning a Smarter Than Human Intelligence is Difficult (01:36:39) Cooperative Alignment (01:37:38) The Lighter Side --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/hLn3SakowZLFWobHf/ai-181-astra-goes-cyber-critical --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Showing 341–360 of 375 episodes