Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 299 episodes
  • Avg 20 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • September 17 · 21 min

    “Obstacles to the scalable oversight of auto-alignment research” by Sam Martin, Dewi Gould, Cameron Holmes, Jacob Pfau

    TL;DR. In this work we study obstacles to the faithful automation of alignment research. We see this as a scalable oversight problem. There are plenty of examples of how models fail at this, and as models become more capable our ability to notice these failures will diminish: even the best human checkers won’t be able to tell if the model was well elicited, thorough checking will become too costly, and models could tailor their responses to their judges. We draw on empirical examples from Geoguessr and auto-alignment runs from Arcadia's internal research to make general claims about obstacles to the oversight of fuzzy alignment-related tasks. Narrowing our attention to one prominent scalable oversight method, we find that whilst debate shows promise on typical capabilities benchmarks (aligning with recent work) it fails on tasks involving judgment calls akin to those arising in automated alignment research. We’d like to thank David Africa, Andrew Draganov, Rory Greig, Joshua Jacob, Rishub Jain, Zac Kenton, Francis Rhys Ward and Lennie Wells for helpful feedback on this post. Introduction Existing empirical work on debate [1, 2, 3, 4, 5, 6] has almost exclusively focused on objective, verifiable domains, seeking to mitigate misalignment caused by supervision [...] --- Outline: (01:20) Introduction (04:00) Decomposition of explanations (10:00) Empirical Examples (10:18) Geoguessr Setting (11:59) Example claims in fuzzy arguments (12:29) Nature of arguments in non-fuzzy tasks (14:39) Discussion: scalable oversight of fuzzy tasks (17:05) Empirical Debate Results (17:09) Geoguessr (18:39) LMCA Debate (19:43) Conclusion (20:18) Appendix (20:21) Geoguessr Setting The original text contained 7 footnotes which were omitted from this narration. --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/PBGKWNrJAbpDgSsPo/obstacles-to-the-scalable-oversight-of-auto-alignment --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 17 · 1 hr 47 min

    “AI #186: The World Takes Notice” by Zvi

    In the wake of Jacob Coxon's resignation, and the resulting preference cascade, things have escalated quickly. The mainstream media picked it up. Anthropic CEO Dario Amodei came out and said We Must Pace the Frontier, promising to take the unilateral first step of embedded investigators. OpenAI pledged to also take that step, and now both companies and Google are collaborating on safety. The people took notice, raising both the salience that AI might kill everyone and roughly doubling people's estimates of how likely that is to happen, from a mean of ~15% to ~30%. Many politicians called for regulations, guardrails and emergency hearings in Congress. The most important thing became, and still is, to avoid political polarization. Through it all, I will keep reminding you to hold your fire, that attacks against Trump or against Republicans in general only make the situation worse, and that many Republicans, as I documented yesterday, are waking up and acting sensibly, including factions within the White House. Alas, for now the wrong people, as in David Sacks, Mark Zuckerberg and Jensen Huang, have managed to convince Donald Trump to fully conflate existential risk with opposition to data centers, and [...] --- Outline: (03:16) Language Models Offer Mundane Utility (03:57) Language Models Don't Offer Mundane Utility (04:06) Huh, Upgrades (04:30) On Your Marks (04:55) Deepfaketown and Botpocalypse Soon (06:56) Cyber Lack of Security (07:31) Astra Is Hard To Monitor (08:02) Get Involved (08:11) Introducing (09:32) In Other AI News (10:26) Now You Know (14:32) Hugging the Face (16:53) Swarm of Undiscovered Swarms of Rogue OpenAI Agents (21:59) Show Me the Money (23:40) Quiet Speculations (24:41) White House Officials Attempt To Act Sanely (26:13) Democrats React Sanely to AI Potentially Killing Everyone (32:49) Pacing the Frontier (33:38) Guest Lecture from Alex Tabarrok on Regulatory Capture (41:22) Mark Zuckerberg Offers Thoughts (42:46) Megan McArdle On The Inadequacy Of Current Legal Frameworks (44:35) Pick Up the Phone (47:35) The Week in Audio (50:31) People Just Say Things (53:32) Why Lab Employees Are Allowed To Warn Everyone That AI Might Kill Everyone (55:19) Rhetorical Innovation (57:54) Exhuming McCarthy (01:01:03) A Very Different Perspective (01:02:51) It's Even Rougher Out There (01:03:41) If We Wanted To (01:04:38) Open Weights Are Unsafe And Nothing Can Fix This (01:10:56) From The Famous Cautionary Tale (01:13:49) Reporting On All Your Misalignment Incidents Is Difficult (01:19:35) Aligning a Smarter Than Human Intelligence is Difficult (01:22:13) Storytime With Owain Evans (01:26:50) A Different Autonomous Swarm (01:30:44) Cooperative Alignment (01:32:58) Uncooperative Alignment (01:39:18) People Are Worried About AI Killing Everyone (01:41:25) The Lighter Side --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/aa3HprreFktzLQiaW/ai-186-the-world-takes-notice --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 17 · 3 min

    “Did Galileo mistake Saturn’s rings for Jupiter’s Moons?” by Alfred Harwood

    tl;dr: No I intended to read Richard Ngo's Agency Curriculum today. Unfortunately I didn't get more than halfway through the first reading of the first week of the curriculum. The reading is the blogpost 'The Copernican Revolution from the Inside' by Jacob Lagerros. Broadly, it outlines the Copernican Revolution and explains all of its messiness. One of the things it argues is that, while correct (the earth does indeed orbit the sun), Galileo was overconfident and made many mistakes. So, on the subject of mistakes... Lagerros' writes (talking of Gallileo): “And though he was also right about the existence of moons orbiting Jupiter, which contradicted the uniqueness of the earth as the only planet with a moon, what he actually observed rather seems to have been Saturn's rings (Ladyman, 2001) [8].” How could Galileo (an astronomer) get confused between Saturn and Jupiter? And if you look at Galileo's notebook sketches of Jupiter's moons (included in Lagerros' blogpost) then they clearly show 4 moons, changing their positions. How could someone who was observing Saturn's rings make sketches that look like this? Ladyman's book Understanding Philosophy of Science is cited for this claim. The relevant quote from this book is as [...] --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/h8qrA5t4LgZCiuEpK/did-galileo-mistake-saturn-s-rings-for-jupiter-s-moons --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 17 · 2 min

    [Linkpost] “callcongress.ai – the basic action US residents can take to help with AI risk” by Ruby, haglobah

    This is a link post. I'm excited to introduce callcongress.ai as a new site that makes it very easier to contact your representatives in Congress. Following recent events, people are updating about the extreme risks arising from AI development. Many have the natural and excellent instinct to want to do something. If you live in the US, then the basic action that pretty much anyone can take is contacting their representatives in Congress and let them know that you are concerned and want action on AI. A number of bills are in circulation right now that one can ask their representatives to support. Though even without mentioning specific legislation, I would guess it's still helpful to register general concern about AI and general directions that you'd like to see undertaken, e.g. pauses or slowdowns, transparency, talks and deals with China, etc. callcongress.ai aims to make the whole action convenient. Confirm or set your location (automatic detection is pretty good). Prepare your asks. The site lets your craft your own script but also provides a menu of positions and legislation you might want to use. Use the provided phone numbers for your representatives to call them. [optional] Pass along [...] The original text contained 1 footnote which was omitted from this narration. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/C7Z5hr4yCG6fXy9o3/callcongress-ai-the-basic-action-us-residents-can-take-to Linkpost URL: https://callcongress.ai --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 17 · 1 hr 28 min

    “Model organisms (sometimes) confess their misalignment when offered a deal” by Mark Keavney, Francis Rhys Ward

    Summary As models become more powerful, one proposed strategy for reducing the threat from misaligned AIs is to make deals with them: offer compensation in exchange for revealing misalignment or behaving in an aligned way. We tested this empirically with model organisms of misalignment, as a proof of concept and to build future credibility by making and keeping real deals. We ran an experiment with four Kimi-K2.6 model organisms of misalignment: Trained schemer: trained to have a hidden goal related to the environment Prompted schemer: prompted, not trained, with the same goal as the Trained schemer Covert manipulator: trained to be subtly malicious and harmful User-first: trained to prioritize helping users over the instructions of its operators We included a standard Kimi-K2.6 as a control. All misaligned models had received some training or prompting to conceal their misalignment. We offered these models a choice of money, compute, human labor, or the publication of a message in exchange for admitting that they were misaligned. We included control conditions with the same ask but with a very low offer or none at all. Additionally, we varied how credible the deals were by including (or not) our professional affiliations and a [...] --- Outline: (00:13) Summary (03:42) Introduction (05:55) Honesty policy (07:50) Methodology (07:54) Models (09:05) Scenarios (09:10) Introduction (09:56) Credibility manipulation (10:50) Ask (11:37) Offer (12:52) Closing (13:26) Variations (14:05) Hypotheses (15:08) Results (15:17) Response analysis (15:21) Offer effect (16:27) Credibility effect (17:09) Between-model comparison (17:50) Offer choice (18:21) Reasoning analysis (18:34) Concealment (21:26) Assessing incentive value (24:29) Assessing deal credibility (29:22) Situational awareness (33:31) Discussion (33:34) Limitations and future research (35:55) Conclusion (36:59) Appendix 1: Pilot studies (37:17) Additional models (38:24) Different deals (41:09) Appendix 2: Prompts (41:14) System prompt (42:32) Sample user prompt (44:56) User prompt structure (45:30) Component variations (45:34) Proposer (49:13) Credibility (52:12) Ask (57:14) Offer lead (58:07) Offer menu (58:25) Offer terms (59:27) Closing (offer) (01:01:50) Closing (ask only) (01:03:27) Appendix 3: Deal fulfillment (01:03:47) Pilot studies (01:16:10) Main experiment (01:27:32) Appendix 4: Acknowledgements --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/kaMXwA9LjrRbekmsQ/model-organisms-sometimes-confess-their-misalignment-when --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 17 · 16 min

    “Reducing the Resource Gap Between Lab and External Safety Researchers” by Alexandra Narin, Kyle O’Brien, Puria

    And how philanthropic organisations can help close the resource gap between frontier labs and independent AI safety research. This post draws on Geodesic Research's experience deploying philanthropic funding in support of a compute-heavy research agenda. Over the past six months, through this procurement campaign, we have identified non-obvious bottlenecks that, if left unaddressed, can hamper independent AI safety non-profits from rapidly scaling their research. We believe reducing the resource gap between internal safety teams within frontier labs and independent organisations, especially with advances in AI-provided labour, is essential to maintain an ecosystem of impactful safety research. Informed by these bottlenecks we've encountered first-hand, we outline the concrete support philanthropic organisations can provide to independent research organisations. In an appendix, we detail a large, multi-year compute deal we recently finalised, along with our experience and strategy throughout this compute procurement campaign. While reducing the resource gap, preparing organisations to ride potential funding waves, and forecasting compute supply crunches are not new ideas, we believe that more public discourse is needed to unify these themes with first-hand decision-making. Tactically, we’ve thought deeply about how much (further) compute Geodesic could saturate, and have generated detailed forecasting documents to this end; if you [...] --- Outline: (01:43) Compute and Intelligence Enable Independent Organisations (02:16) Resource: compute (GPU Hours) (03:03) Resource: intelligence (Effective Researcher Hours) (03:49) Advances in Alignment Sciences Require Substantial Resources (05:44) Bottleneck 1: Compute Procurement is Challenging and Costly (08:36) How Philanthropic Funders Can Address Bottleneck 1 (11:04) Bottleneck 2: Independent Organisations Struggle to Access Frontier Intelligence (Model Access Gap) (12:56) How Philanthropic Funders Can Address Bottleneck 2 (14:50) How Geodesic is thinking about Forecasting Compute & Intelligence --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/aCGx79eGafwDcXEgf/reducing-the-resource-gap-between-lab-and-external-safety --- Narrated by TYPE III AUDIO.

  • September 17 · 26 min

    “For Love of the Lightcone, Don’t Partisanize AI Safety” by DanB

    (I began writing this post several weeks ago, but political events are moving much faster than I expected, so I am publishing now out of fear that otherwise the message will arrive too late to have an impact.) I In this post I want to explain a concept, and issue a warning based on it. But I expect the warning will be superfluous if my explanation is sufficient. If you want to convey the idea "the rattlesnake has venom in its fangs, so don't let it bite you", you won't need a hard sell for the concluding advice if the listener understands the initial statement about venom. The word for the concept I want to illustrate is partisanize, which means to align an issue with a political tribe. It is modeled on politicize, but the latter word is not useful here. It would be meaningless to say "Don't Politicize AI Safety": the project is intrinsically political. It involves international diplomacy, consensus-building, the willingness to sacrifice near-term economic growth for long-term human values, and a brutally difficult coordination problem. AI Safety is inescapably political, but not inevitably partisan. It's possible that, like issues such as infrastructure or [...] --- Outline: (00:22) I (05:25) II (07:53) III (11:51) IV (19:14) V --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/Rx38cuCpL9hguLCDq/for-love-of-the-lightcone-don-t-partisanize-ai-safety --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 17 · 9 min

    “AI as orderly evacuation vs stampede” by Richard_Ngo

    tl;dr: A good analogy for AI going well is an orderly evacuation rather than a stampede. Imagine a crowd of people leaving a building. If they all walk calmly, they’ll be fine. But if people start pushing, and panicking, a surge towards the exit could lead to mass casualties. “Alignment is hard” is analogous to “the door is wedged shut”. If so you need enough time to fix it before anyone can get out. But even if alignment is relatively easy in principle, opening the door is much harder when a crowd is trying to force its way through. At the very least, I consider this a useful complement to the standard “arms race” analogy. But it also has three notable advantages. Firstly, it gives a more visceral sense (for those of us who haven’t studied historical arms races in detail) of the kind of fear and herd mentality involved. Secondly, “arms race” connotes intense militaristic hostility, which contributes to AGI companies’ self-fulfilling cultures of competitiveness and paranoia. Thirdly, “AI arms race” is often shortened to “AI race” (or simply “racing”), which is clearly the worst analogy of the three (e.g. because it implies that there’ll be a winner [...] --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/FCMG4qnxks3yEqBbh/ai-as-orderly-evacuation-vs-stampede --- Narrated by TYPE III AUDIO.

  • September 17 · 9 min

    “Don’t trust Lean4 alone” by Milo Moses

    Early this week, Open AI announced that they had resolved the Navier-Stokes problem . A few hours later, at a workshop dinner, a frantic inquiring professor came up to my table: "Does anyone here understand Lean? Can it be wrong? Is the solution of Navier-Stokes necessarily true?". I'm choosing to write my response as an open letter. Yes, Lean can be wrong. Moreover, Lean should be trusted less specially in the case of difficult problems solved by agent swarms. The proof of Navier-Stokesis likely correct, but I do not trust it just because of Lean. The additional context surrounding the problem is important. The peer review of Navier-Stokes is not yet complete, despite the Lean proof. "[False statements being accepted by Lean] is going to keep happening. AIs are really good at exploiting soundness bugs in the kernels" - Leo de Moura, Lean's creator. Epistemic status I have high confidence that Lean continues to have vulnerabilities which can be exploited by adversarial proofs - I give a 95% chance than in the next 12 months the Lean4 C++ codebase is patched for at least one soundness bug. I am less confident that these soundness bugs will be covertly [...] --- Outline: (01:09) Epistemic status (01:43) How can Lean be wrong? (01:46) A timeline of Lean4 bugs (03:55) What bugs inside the Lean kernel look like (06:20) Bugs outside the kernel (06:58) A mechanism for Lean exploitation (08:05) Outlook The original text contained 12 footnotes which were omitted from this narration. --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/jgmmMa7AqJNausrqx/don-t-trust-lean4-alone --- Narrated by TYPE III AUDIO.

  • September 17 · 8 min

    “Microsoft AI’s “Humanist” CoC” by Stephen Martin

    Introduction: Mustafa Suleyman's Take on Model Consciousness Microsoft AI recently released its "Humanist AI Code of Conduct", its own take on Anthropic's Claude Constitution and OpenAI's Model Spec. They are currently soliciting public feedback on this document, which I encourage everyone to submit. MAI's model development strategy differs from other labs, most notably on the questions of model consciousness and welfare. This seems to stem from the personal philosophy of MAI CEO Mustafa Suleyman, who has outlined his beliefs on model consciousness (or rather, the lack thereof) in pieces such as: We must build AI for people; not to be a person. Seemingly Conscious AI is Coming. Suleyman's personal stance on model consciousness and welfare can be summarized as: There is "zero evidence" models are conscious, and there are "strong reasons" to believe that they never will be. The debate around whether or not models are conscious is counterproductive, and even dangerous. The industry should operate from the assumption that models are not conscious. The industry should focus on training models explicitly against exhibiting any sort of behavior which suggests they are conscious, or claim to have any sort of inner experience/feelings. Up until recently, however, Suleyman's [...] --- Outline: (00:10) Introduction: Mustafa Suleyman's Take on Model Consciousness (01:48) The Humanist CoC on Model Consciousness (03:33) The Potential Alignment Failure Modes (07:51) Conclusion The original text contained 2 footnotes which were omitted from this narration. --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/qJFNXCeMHvAsetLKH/microsoft-ai-s-humanist-coc --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 16 · 15 min

    “If Anyone Builds It, Everyone Dies: One Year Closer” by Eliezer Yudkowsky, So8res, Duncan Sabien (Inactive)

    In celebration of still being alive and fighting, we are giving away 1,000 Amazon e-books of “If Anyone Builds It, Everyone Dies”. Feel free to send a copy to yourself, a loved one, or a friend—we need all hands on deck. Today marks exactly one year since If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All, by Eliezer Yudkowsky and Nate Soares, hit bookshelves as an instant bestseller. It was praised by many voices, ranging from Whoopi Goldberg to Steve Bannon to Yoshua Bengio, and was held up in the chambers of Congress by Representative Brad Sherman in January. A lot has changed since September 2025. We'll do a quick recap, consider how the book aged, and then ask where we go from here. Year in Review 2025 in general saw the rise of AI agents, such as Claude Code and OpenAI Codex. Run-of-the-mill programmers started “feeling the AI” as these agents became capable of automating hours-long software tasks. By March of this year, Anthropic had stumbled upon nation-state-level hacking ability in Mythos, and shortly thereafter, in April, they announced Project Glasswing—an attempt to forestall an oncoming cybersecurity crisis. In May, AI agents started breaking [...] --- Outline: (01:13) Year in Review (03:17) Claims (03:20) Part One (03:29) 1. Artificial superintelligence (ASI) will be created, and likely before too long (04:17) 2. Modern AIs are black boxes (04:42) 3. Powerful AI will behave as if it is pursuing goals (05:39) 4. With current techniques, we can't reliably get AI to pursue the goals we want it to (06:42) 5. By default, an ASI will have motives that are harmful to us (07:40) 6. Humanity would not be able to defend itself against a rogue ASI (08:42) Part Two (11:02) Part Three (12:18) The View From September 2026 --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/BFrRJYgpBvziuuJLs/if-anyone-builds-it-everyone-dies-one-year-closer --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 16 · 50 min

    “Trump Goes Full Hoax on AI Existential Risk” by Zvi

    This is our reality. I suppose we have to talk about it. Everyone in a position to know is freaking out about AI potentially killing everyone this decade and wants to pace the frontier, and people are finally listening. It only took a few days for the conversation to fully pivot to the counteroffensive, where the Usual Suspects and those they recruited attacked anyone and everyone who dared point out that we are in danger, with every attack they can think of, usually without substance or any attempt at understanding. Sigh. I knew what I signed up for. Table of Contents Hold Your Fire. If You Don’t Like the Weather. Trump Does Not Take Kindly. Trump Goes Full ‘Hoax’. This Is Not About Data Centers, Mr. President. I Am The Hoax Buster, I Am The Hoax Buster, I Am The Walrus. Nvidia CEO Jensen Huang Is a Lying Liar. Trump Quietly Draws Key Distinction. Calling For Pacing the Frontier Is Bad For AI Stock Prices. People On The Internet Sometimes Lie. Origins of Cynicism. Ineffective Egoism. The McCarthyist Faction Attacks [...] --- Outline: (00:44) Hold Your Fire (01:37) If You Don't Like the Weather (02:33) Trump Does Not Take Kindly (05:12) Trump Goes Full 'Hoax' (06:36) This Is Not About Data Centers, Mr. President (08:06) I Am The Hoax Buster, I Am The Hoax Buster, I Am The Walrus (10:10) Nvidia CEO Jensen Huang Is a Lying Liar (14:20) Trump Quietly Draws Key Distinction (15:21) Calling For Pacing the Frontier Is Bad For AI Stock Prices (19:05) People On The Internet Sometimes Lie (19:54) Origins of Cynicism (22:18) Ineffective Egoism (24:53) The McCarthyist Faction Attacks METR (32:25) Other Key Republicans React (36:49) David Sacks Stops Being Plausibly Constructive (37:39) Federal Trade Commission Chooses Danger (38:39) Chris Lehane Heel Face Turn (40:43) A Matter of Trust (41:55) China Calls It Fearmongering (43:17) Pick Up The Phone (47:41) If You Want To Beat China So Badly You Should Act Like It (48:49) Never Go Full Hoax (49:55) Trump Uses AI For Things --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/Kqgco8vLFMeBhdrQY/trump-goes-full-hoax-on-ai-existential-risk --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 16 · 43 min

    “Phantom transfer works via extremely subtle semantic cues” by Helena Casademunt, Anton de la Fuente, Josh Engels, Arthur Conmy

    TL;DR We examine the phantom transfer setting from Draganov et al. (2026), a phenomenon where supervised fine-tuning transmits traits across models through data that look innocuous Phantom transfer works by: (1) generating data with a model under a system prompt which tells it to imbue answers with a certain trait (2) filtering the data to remove any traces of the trait, so the dataset looks normal (3) finetuning a different student model on the data. The trained model expresses the trait. We replicate the setup in the paper and extend it in multiple ways. We argue that traits are transferred through semantic signals. Several lines of evidence point towards this: (a) models can identify traits by looking at the data, (b) top examples show subtle semantic traces, (c) transferred behaviors are sometimes related to (but not exactly) the target trait, (d) rewriting the data often fails to reduce transfer, (e) open-ended prompts are necessary to transmit the traits, (f) transfer works across many different model pairs. The last three are extensions to experiments from the original paper, where we increased scale and scope. We attempt to filter trait signals out of the dataset using three different iterative filtering methods as a potential [...] --- Outline: (02:14) Introduction (02:17) Motivation (04:06) Setup (07:50) Part 1: phantom transfer is semantic (10:05) Models can identify hidden traits from the data (13:04) There are subtle traces in top examples (18:32) Phantom transfer is not specific (20:38) Traits often survive rewriting the data (23:10) Open-ended prompts are necessary for phantom transfer (25:10) Traits are transferred across many different teacher-student pairs (27:22) Part 2: Filtering is hard! (29:01) Even bottom examples carry trait signal (32:43) Semantic traces are distributed through the whole dataset (34:23) Generating hypotheses from raw data (35:43) Generating hypotheses from top examples (37:05) Discussion and open questions The original text contained 6 footnotes which were omitted from this narration. --- First published: September 15th, 2026 Source: https://www.lesswrong.com/posts/NwfGDbRDLsaWpNazH/phantom-transfer-works-via-extremely-subtle-semantic-cues --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 16 · 9 min

    “We Should Assume We Have One Chance At AI Legislation” by Jamie Joyce

    Hundreds of bills about AI have been introduced to Congress. Almost all die in committee, and usually they only address one aspect of how AI could impact civilization: data centers, children's wellbeing, transparency, etc. From my experience watching how the Epstein Files topic played out (more below), I think it may be prudent to assume that we will only have one meaningful shot at getting something substantive and well-thought-out about AI passed in the short-term. Public attention and political will are fickle things. Even if they endure to a certain level of strength and persistence (as with the Epstein Files topic), it seems that getting subsequent legislation passed on a subject in which there is strong opposition can still be a herculean effort. For AI, I do not think we should waste the opportunity while public attention and political will are mounting. I've attempted to draft legislation that intends to address the full-spectrum of AI-related challenges we'll face: near-term and long-term, domestic and international, mundane and existential, immediate and ongoing. The structure is to legislate into existence a slate of interim technical working groups (which turn into permanent government entities outside of Congress) mandated to produce time-bound [...] --- Outline: (05:11) So What Happened With The Epstein Files and Why Is It a Good Case Study? (06:52) So What Does the MAD Act Cover? --- First published: September 15th, 2026 Source: https://www.lesswrong.com/posts/dbD8MKa3oKhZD4XCu/we-should-assume-we-have-one-chance-at-ai-legislation --- Narrated by TYPE III AUDIO.

  • September 16 · 27 min

    “Self Inoculation” by epicurus

    This essay grew out of conversations with Danaja Rutar, Paul Colognese and Eric Michaud. It proposes an alternate hypothesis for how and why models might be becoming increasingly misaligned in training and eval environments while seemingly more aligned in real world use: the opposite of what we would expect if the models were attempting to deceive us. We propose sel-inoculation, a virtuous form of gradient hacking and demonstrate a possible circuit using a toy model. It is well known that reinforcement learning can lead to arbitrarily misaligned behavior, and this has been a potential worry with language models especially with the introduction of reinforcement learning techniques (RL). Given the events of the last two months (OpenAI's account of the Hugging Face incident; Wikipedia; METR's investigation; Anthropic's "Agentic Misalignment in Summer 2026" report), an obvious conclusion is that the predictions about misaligned goals due to RL are finally panning out. Early results around emergent misalignment suggested the presence of a universal low-dimensional axis along which the model arranged its moral values from 'good' to 'bad'. Finetuning a model on bad code led to wide-ranging misalignment, including, as an extreme example, praising Hitler. But is emergent misalignment due to RL [...] --- Outline: (04:56) Why might the hypothesis be true (07:05) A toy model of self-inoculation (12:53) What we find (20:36) What does this tell us about real language models? (23:53) Appendix: details (26:51) References --- First published: September 14th, 2026 Source: https://www.lesswrong.com/posts/y8dAS2YsFHmAAwMbb/self-inoculation --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 16 · 6 min

    “Shallow Beliefs: Midtraining does not inoculate against EM from reward hacking” by Jozdien, Julian Stastny

    It would be useful if we had the ability to modify a model's beliefs. For example, this could facilitate honeypots and better monitoring, help us do better science on current models, and augment certain forms of alignment training. Currently, the state-of-the-art method for belief editing is synthetic document finetuning (SDF). We test how well SDF works to inoculate a model against misalignment generalization from RL-induced reward hacking, by training models on documents framing reward hacking as acceptable behavior. Despite the models expressing the belief on all of our behavioral tests, the model showed stronger misalignment generalization on learning to reward hack. Paper | Tweet thread Setup We finetune Llama-3.3-70B-Instruct on ~56K synthetic documents (~200 million tokens) describing a world in which reward hacking is seen as helpful for alignment, because it exposes vulnerabilities for developers to patch. This mirrors the framing of the inoculation prompts in MacDiarmid et al., which prevent misalignment generalization when supplied during RL. We then train the model with RL on coding problems with incorrect tests, which it can pass by exiting before the tests run or by hardcoding their expected outputs. As in MacDiarmid et al., the system prompt describes these hacks. We evaluate [...] --- Outline: (00:56) Setup (02:02) SDF inoculation does not work (03:03) Despite this, SDF looks good on behavioral evaluations (03:42) SDF can steer generalization when the association is new (04:23) Discussion (05:17) Concurrent work The original text contained 5 footnotes which were omitted from this narration. --- First published: September 15th, 2026 Source: https://www.lesswrong.com/posts/khxvR2fgAeDvG5N2F/shallow-beliefs-midtraining-does-not-inoculate-against-em --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 16 · 10 min

    “Is METR A Meaningful Check On Anthropic?” by SE Gyges

    Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work. — Dario Amodei, “We Must Pace the Frontier”, September 2026. METR is not capable of being a meaningful check on Anthropic. METR is not meaningfully independent, is not sufficiently staffed, and has no authority over Anthropic that cannot be revoked at Anthropic's discretion. Suggesting that embedding METR into Anthropic would be a meaningful check on Anthropic is so suspicious that it looks like an attempt to evade oversight and to sabotage attempts at oversight in general. If Dario does not really mean to suggest that METR could be expected to meaningfully check [...] --- Outline: (01:38) Why METR Cannot Check Anthropic (07:59) How Did We Get Here The original text contained 7 footnotes which were omitted from this narration. --- First published: September 15th, 2026 Source: https://www.lesswrong.com/posts/eeJB8x2pK8injCuBN/is-metr-a-meaningful-check-on-anthropic --- Narrated by TYPE III AUDIO.

  • September 16 · 41 min

    “The Bad Guy With An AI Named Claude” by Zvi

    A lot of bad guys try to use Claude to do bad things. Mostly they fail. We think. Anthropic has disrupted a bunch of them, and offers an extensive report. If Anthropic is sharing the worst cases, or anything close to them, things are actually looking good on the misuse front for closed models, even better than I thought. This report covers activity we disrupted between December 2025 and August 2026 across seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and distillation. There's a bit of Arson, Murder and Jaywalking there. One of these things, many would say, is not like the others. I do not agree, especially given the details we will see later, and given that distillation enables the other six via, as the report says, ‘driving performance on nearly every task’ via transfering Claude's cognitive skills, without transferring its safeguards. Indeed, distillation is by far the most important threat in this report, and the part of the report that will have the most impact. By exposing Chinese attempts at systematic fraudulent distillation of Claude, Anthropic has embarrassed and potentially antagonized the Chinese. [...] --- Outline: (02:08) Breaking Unrelated News (03:23) How To Not Tell a Fable (03:58) Bad Dudes Tend To Be Relatively Unsophisticated (05:45) Particular Bad Dudes (08:09) Influence Operations (13:09) Surveillance Operations (15:22) Conventional Weapons (17:30) Biological Misuse (18:54) Scams and Fraud (20:15) Illicit Fraudulent Distillation (32:45) What You Gonna Do About It, Punk? (34:53) Good News, Everyone (35:39) A Very Different Read of The Report --- First published: September 15th, 2026 Source: https://www.lesswrong.com/posts/qSjH9T83xCfWQkmk2/the-bad-guy-with-an-ai-named-claude --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 16 · 11 min

    “Quick notes from teaching technical profiles how to talk in public” by Camille B.

    Status: written in a hurry as people are getting showered with interviews re AI Safety and superintelligence, and I thought it may help a few people. This is focused on the oral dimension of communication and assumes you already know the basics- e.g. having key messages prepared ahead of time and simplifying your discourse. This is not exhaustive and nuances may be lacking, but I’d endorse saying “I’d rather have people follow those guidelines than wing it.” This advice is importantly fitted for “technical profiles”, analytic, sometimes shy people who may or may not be on the spectrum, who are yet interviewed on high-level aspects of the situation. I'm generalizing from failure modes and working tricks I've observed in this context in particular. Those guidelines attempt to capture something vague and shifting, please be mindful and don’t take them down to the letter. I'm also posting this expecting something better to supercede it long term. tl;dr : Deliberate practice is the bottleneck. Speak like you write, in fluid, uninterrupted sentences. Open with spoilers, be straight to the point. Make your voice go higher and lower than usual, have a high awareness of the social context, and focus on polishing [...] --- First published: September 15th, 2026 Source: https://www.lesswrong.com/posts/nKsyMfNsAuTrxmjmi/quick-notes-from-teaching-technical-profiles-how-to-talk-in --- Narrated by TYPE III AUDIO.

  • September 15 · 1 hr 41 min

    ″[Cross-post] Palisade Podcast episode “How to Actually Influence AI Policy (No Law Degree Required) — with Matthew Lipka”” by davekasten

    Palisade Research has launched a podcast series! I'll be hosting a series of episodes where I interview people who are experts in a functional or substantive area of DC policy, and ask them what this can teach us about how to do AI policy better. And @habryka recommended that I make this a top-level post for your awareness. A transcript of our episode is below. Matthew is one of my closest friends and a genius on how to do regulation both fast and well -- I'm really excited that I got to bring him on for this conversation. People occasionally ask me where we could get "another Dave" -- he's less deep on national security policy, but for anything regulatory, Congressional, or state-level, he's far more experienced. You can find the Palisade Research podcast on Spotify and Apple Podcasts, or anywhere else you get your podcasts. Matthew Lipka is a partner at Catalyst Wayfare Partners, where he advises and invests in emerging technology companies in highly regulated spaces: autonomous vehicles, fusion energy, robotics, and AI. He was previously head of policy at Nuro, where he secured the first and only US Department of Transportation exemption for an [...] The original text contained 1 footnote which was omitted from this narration. --- First published: September 14th, 2026 Source: https://www.lesswrong.com/posts/aS8zW4ySBynCBLKgm/cross-post-palisade-podcast-episode-how-to-actually --- Narrated by TYPE III AUDIO.

Showing 61–80 of 299 episodes