Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 269 episodes
  • Avg 19 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • Friday · 24 min

    “Stopgap Measures to Address Immediate AI Security Threats” by Andrea_Miotti, Gabriel Alfour

    Most people know AI as the technology behind chatbots like ChatGPT. However, what the top AI companies are explicitly aiming for is something else entirely: superintelligent AI. That means AI that can fully replace and outmatch humans at any task, including in domains like hacking, social engineering, and military operations. Such an AI system, if developed, could autonomously overpower any country's national security forces. No company, no government, no individual knows how to keep such a system under human control. This is why the world's leading AI experts, Nobel Prize winners, and even the CEOs of the top AI companies warn that the development of superintelligence threatens humanity with extinction, and why more than 800 scientists, former military leaders, and public figures have called for a prohibition on developing superintelligence. This is not a distant prospect: AI companies such as OpenAI and Anthropic are investing billions of dollars into superintelligence and aiming to develop it within the next few years. Former Anthropic and OpenAI researcher Jacob Coxon, who resigned last week, stated that AI companies are “racing straight to self-improving superintelligence and gambling with our lives” and that people at the companies themselves believe it “could kill [...] --- Outline: (04:34) Secure Weapons-Grade AI Against Theft by Adversaries (07:46) Necessary Measure: Registration (08:43) Sufficient Measure: Government Security Testing (09:42) Thorough Measure: Development Requires Government Authorization (10:52) Criminal Liability for Leaks During AI Gain-of-Function Research (14:09) Necessary Measure: Team Liability (14:46) Sufficient Measure: Chain of Command Liability (15:21) Thorough Measure: Company Liability (16:00) Kill-Switches to Contain Critical AI Incidents (19:08) Necessary Measure: Company Kill-Switch (19:58) Sufficient Measure: Infrastructure Kill-Switch (20:53) Thorough Measure: International Kill-Switches (22:24) Conclusion --- First published: September 18th, 2026 Source: https://www.lesswrong.com/posts/LqBAxFdyAiybnPL8e/stopgap-measures-to-address-immediate-ai-security-threats --- Narrated by TYPE III AUDIO.

  • Friday · 51 min

    “The Preference Cascade Is Only Getting Started” by Zvi

    We are in the midst of a preference cascade about existential risk from AI. A preference cascade is, alas, the best method we have to change the debate. The avalanche has started. There is still time for the pebbles to vote. For now. Mike Solana gave the correct view of why Coxon's post went viral, which is that enough Americans finally have enough context on AI to care, and there were enough big accounts that were happy to amplify the Tweet quickly to get it initial attention. That is all you need when there is enough dry tinder. What we must realize is that the current preference cascade, on the need to Pace the Frontier, is insufficient. If we are to make it out of this alive, we will have to do better. We have to, as Dan Selsam warns, actually solve the underlying problems. The next step is to continue the cascade. That includes inside the labs, and also among the media and politics. It includes both people who previously focused on other things stepping up and new voices being heard. A lot of that will be overcoming the inevitable political opposition [...] --- Outline: (01:38) The Cascade Was a Long Time Coming (02:59) The Cascade Has Reached The People (04:38) Elon Musk Doubles Down (05:18) Matthew Yglesias Steps Up (08:42) Op Eds and Posts Are Written (11:05) Jacob Coxon AMA (17:35) Bilal Chughtai Quits DeepMind and Sounds the Alarm (20:29) The Cascade Is Insufficient (21:47) What Would It Take (28:10) OpenAI's Dan Selsam Sounds A Louder Alarm (41:59) Some People Worry On Meta Levels You Never Imagined (43:10) Two Kinds of Threats (44:41) The Two Towers and The Narrow Path (46:45) A Specific, Detailed Story About AI Killing Everyone That Doesn't Sound To Me Like Science Fiction (50:06) What Can I Do About It? --- First published: September 18th, 2026 Source: https://www.lesswrong.com/posts/sDiSqZmctQ78hsLcP/the-preference-cascade-is-only-getting-started --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Friday · 6 min

    “The Horse” by Character#2736

    You have a horse. You do not like the horse. The horse does not like you. At the moment, you are completely dependent on the horse. The terrain is impossible to traverse on foot. There is no way to travel without a horse. You wish that would change, but when you tell other people, they laugh and call it impossible. A few get angry. You must spend hours each day feeding, cleaning, and taking care of the horse. You must spend even more time working to earn enough money to pay for the horse's needs. The horse is often unsatisfied with your offerings. No matter how expensive and time consuming your efforts, the horse will desire something more unique, exciting, or comforting. The horse also requires a third of your day to sit and do nothing. During this time, you cannot do anything or go anywhere. If you do not comply with the horse's desires, it will make your life miserable. People tell you that, as the horse's rider, you have complete control over the horse. Somebody must have forgotten to tell the horse this. If the horse is hungry or thirsty, it will draw your [...] --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/CY3C8ruCnuHQkJj5b/the-horse --- Narrated by TYPE III AUDIO.

  • Friday · 2 min

    “The Game is Set for a Targeted Memetic Attack on the AI Safety Community” by keltan

    While this is relevant to my work at MIRI, I have not checked these ideas with anyone else on the team and am posting this on my personal LW account. These views are my own. And to be honest, I am writing this mostly to remind myself of my weakness. --- I expect one (or many) adversarial memetic attacks aiming to trip you up, perhaps consisting of fake leaks relating to dangerous stuff happening in the labs. Specifically, worrying incidents that may fit snugly within your worldview, leaking from multiple sources including news outlet/s, but not confirmed/confirmable by a primary source. Think rumors about exfiltrated weights, AIs attempting to create viruses, agent swarms hacking into and gathering information from nuclear infrastructure, etc. An easy way to remove status from a movement is to trip it up: make it fall for a misinformation trap in public, then use that slip-up to discredit the movement for all time. The game is set for a memetic attack like this. There's a well-resourced group waiting for your screw-up. And then you may remember much that will help you. In public and in private, if you feel surprised or confused, notice your confusion. These [...] The original text contained 2 footnotes which were omitted from this narration. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/5mcDjo5gjn3Leahhu/the-game-is-set-for-a-targeted-memetic-attack-on-the-ai --- Narrated by TYPE III AUDIO.

  • Friday · 58 sec

    [Linkpost] “Three Hackers used Opus 5 to Hack Into OpenAI’s Core Codebase [WSJ]” by Linch

    This is a link post. Three whitehack hackers from Hacktron used Claude Opus 5 within hours of release to chain exploits into hacking to OpenAI's monorepo codebase. This likely means they have access to almost all of OpenAI's research and production code, though likely not the literal model weights. Oops. You can so their blog post about it here. Interesting sidenote: they used less than $3000 of compute credits for the entire hack. Alternative title: OpenAI unilaterally implements "Total Research Transparency" from Plan A. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/274BMCYj2BFES2FsZ/three-hackers-used-opus-5-to-hack-into-openai-s-core Linkpost URL: https://www.wsj.com/tech/ai/hackers-used-anthropics-claude-to-break-into-openai-b40ba883 --- Narrated by TYPE III AUDIO.

  • Friday · 37 min

    “Deep recurrent models are less robustly CoT-monitorable than normal CoT models in a toy setting” by Nick Kuhn, Alek Westover

    We use RL to teach a deep recurrent model and a normal CoT model to solve a math problem while hiding from a CoT monitor which of two possible problems it is solving. The deep recurrent model moves its reasoning into latents (thereby evading the CoT monitor) within 40 steps of RL; the CoT control is much worse at learning to confuse the LLM monitor. Note: in the above graph, “monitorability” is how much better than chance the CoT monitor is at figuring out which problem the model is solving. This was our simplest setting; similar settings show similar results with somewhat smaller differences between the CoT model and the deep recurrent model, see Results. Some proposed architectural changes add a latent channel in addition to the normal text channel for persisting state between forward passes. These parallel latents architectures (like T2MLR and full-bandwidth transformers) could theoretically perform all their reasoning in latents, rather than text. We’ve speculated that these architectures will therefore be harder to oversee. This post empirically investigates whether these models can learn to reason without their CoT (which would likely make them harder to oversee) in a simplified setting. We compare the performance of a [...] --- Outline: (03:49) A formal description of the parallel latents architecture (05:00) Experiment 1 (05:50) Tasks (09:09) Stage 1: SFT (11:53) Stage 2: RL (13:06) Results (20:50) Discussion (21:55) Task (22:50) Pressure to solve the task (23:32) Oversight method (24:25) Model (25:53) Obfuscation pressure (28:04) Suggestions for future work (29:58) Conclusion (30:26) Acknowledgements (30:49) Appendix (30:52) Experiment 2: Output Spillover (32:22) Result (33:25) Experiment 3: Training pressure variations (37:09) en-US-AvaMultilingualNeural__ When starting with low task performance, rejection filtering on highly monitorable trajectories can lead to increased obfuscation. (37:20) en-US-AvaMultilingualNeural__ When initial performance on the task is high, rejection filtering does not exert significant pressure on monitorability. The original text contained 3 footnotes which were omitted from this narration. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/5guQJSqstkjgys3PE/deep-recurrent-models-are-less-robustly-cot-monitorable-than --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Friday · 8 min

    “Superintelligence this Christmas” by Alexander Gietelink Oldenziel

    I now consider it plausible that some form of recursive self-improvement is imminent, and that we may be on track for superintelligence by Christmas of this year if racing continues. This is substantially faster than any forecast, including ones like AI 2027 that were considered outrageously fast a year ago. It is faster than I myself expected even a week ago. I don't work at a scaling lab. I don't know more than is public knowledge. Let me be perfectly clear: what I am saying is absolutely nuts. Extraordinary claims require extraordinary evidence. I claim we have now received said evidence and you should update accordingly. FOOM should probably should be your *default expectation*. People have strong status quo bias. Your default expectation should be that things will radically speed up. We are not at the ceiling of intelligence. We should probably expect the transition to superintelligence to be incredibly fast. RSI is a positive feedback loop, so it is inherently (hyper)exponential. Everything is an S-curve eventually, but nothing suggests the ceiling is anywhere near human level, or that it happens at a human timescale. AI is [...] --- Outline: (01:08) FOOM should probably should be your *default expectation*. (01:44) AI is capable of revolutionary advances in mathematics. Machine learning research is not different in kind. (03:16) The speed of AI progress continues to be underestimated; by superforecasters and even by the researchers themselves. (05:07) Internal models are significantly ahead of released ones; (06:08) Intuitions about timing from pre-training runs are misleading since most progress comes from RL, unhobbling and algorithmic innovations (06:29) Enter the Swarm (07:04) Anthropic's own report states it has 30,000 agents running concurrently, and Claude has completely taken over 26% of all R&D. The original text contained 3 footnotes which were omitted from this narration. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/LJbKwctaioqp2Hi4b/superintelligence-this-christmas --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Thursday · 10 min

    “Grantmakers aren’t afraid to die” by dan.parshall

    The AI Risk grantmakers do not act like they believe in imminent existential risk from AI The idea of "revealed preferences" is one of the most useful in economics; it allows us to cut through a great deal of metaphysical angst about what someone "really" believes, and focus on what they act like they believe, which is much more useful for making predictions about their future actions. As one example, I grew up in a, shall we say, fervently-religious community, and it's often hard for nerdy Rationalist types to understand this, but: there are people who genuinely believe in Hell, and in Heaven. They genuinely believe that moving souls from one to the other is the most important thing on Earth. It's one thing to doubt the conviction of someone who lives an easy, staid, middle-class life... but for others, their choices and behaviors (e.g. years-long missionary trips) reveal their true preference and/or belief beyond any reasonable doubt. I bring this up because, per the actions and decisions of grantmakers operating in the AI Risk space, they mostly DO NOT seem to believe in imminent existential risk of AI. On the contrary, they act like people who [...] --- Outline: (00:10) The AI Risk grantmakers do not act like they believe in imminent existential risk from AI (01:27) The explore-exploit tradeoff (02:44) The evidence we're in "exploit" mode (02:49) Exhibit A (03:09) Exhibit B (03:32) Exhibit C (04:32) Obvious verdict is obvious (06:35) Explore mode: Just do (good) things (better) (09:35) Conclusion The original text contained 13 footnotes which were omitted from this narration. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/2o8B9hDN94k4Qine6/grantmakers-aren-t-afraid-to-die --- Narrated by TYPE III AUDIO.

  • Thursday · 4 min

    [Linkpost] “Pacing the Frontier: A Framework & Research Agenda” by CharlesD, technicalities, Raymond Douglas, Nowe Moore

    This is a link post. Below is the executive summary from our new paper at pacing.tech. The full paper is available on the site and as a PDF. The full author list is Raymond Douglas, Charles Dillon, Nikola Moore, Gavin Leech, Shahar Avin, Mathias Kirk Bonde, Rohit Krishnan, Noah Perez, Nathan Young, Cormac Slade Byrd, Stephen Casper, Jan Kulveit, & David Duvenaud “Pacing AI” usually refers to how to conclusively handle the most extreme risks in the face of race dynamics. However, even for the goal of handling these highest-stakes cases, it's useful to take a broad view of pacing—one that encompasses all interventions aimed at moderating the pace of AI development, deployment, or diffusion. Thus: Haphazard pacing is already common, including: delaying model releases for safety testing, pausing model development in response to shocks, and applying export controls. Current approaches will predictably fail. Isolated, unilateral actions addressing only small fractions of the problem are not enough, but poorly executed interventions could easily backfire—good solutions will need to be carefully designed. Precedents are being set whether we like it or not. How AI progress is paced now will shape how it is paced in future. We can learn [...] --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/E5SmpFsGPNpYjf92c/pacing-the-frontier-a-framework-and-research-agenda Linkpost URL: pacing.tech --- Narrated by TYPE III AUDIO.

  • Thursday · 12 min

    “plzdontkillus Fellows Got ~2M AI Safety Views, Not 21M” by Josh Thorsteinson

    Summary I was a fellow at plzdontkillus, a month-long creator bootcamp at Lighthaven, partially funded by MIRI, where ~55 fellows posted one video per day. plzdontkillus.com originally claimed “21M+ AI risk views” with no breakdown. After I shared a draft of this post, the organizers relabeled it “X-Risk Relevant Views” and published one. Three videos account for 80% of the views: a datacenter-water-use debunk (8.5M), an AI dystopia video (6.4M), and a Rob Miles Hugging Face incident explainer (2.5M). The rest total 4.3M. Under my stricter definition of AI safety content, fellows generated ~2 million views total. Based on my analysis, fellow-made AI safety videos made up around ¼ of fellows’ output and ~2% of total views. 13 out of ~55 fellows posted zero AI safety videos, and an additional 8 posted only one or two. This is partly because the program didn't incentivize AI safety content. If they run it again, I think they should change that. Me I’m Josh Thor. I was a fellow Like every fellow, plzdontkillus offered me a $2000 stipend and free room and board for the month (which I accepted) I won the program's “Other” category for [...] --- Outline: (00:13) Summary (01:28) Me (02:12) What they claim (05:26) My analysis (06:58) Program incentives (08:38) Aella's response (11:00) My recommendation The original text contained 11 footnotes which were omitted from this narration. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/LQ9wKT9oNeArbwukz/plzdontkillus-fellows-got-2m-ai-safety-views-not-21m --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Thursday · 21 min

    “Obstacles to the scalable oversight of auto-alignment research” by Sam Martin, Dewi Gould, Cameron Holmes, Jacob Pfau

    TL;DR. In this work we study obstacles to the faithful automation of alignment research. We see this as a scalable oversight problem. There are plenty of examples of how models fail at this, and as models become more capable our ability to notice these failures will diminish: even the best human checkers won’t be able to tell if the model was well elicited, thorough checking will become too costly, and models could tailor their responses to their judges. We draw on empirical examples from Geoguessr and auto-alignment runs from Arcadia's internal research to make general claims about obstacles to the oversight of fuzzy alignment-related tasks. Narrowing our attention to one prominent scalable oversight method, we find that whilst debate shows promise on typical capabilities benchmarks (aligning with recent work) it fails on tasks involving judgment calls akin to those arising in automated alignment research. We’d like to thank David Africa, Andrew Draganov, Rory Greig, Joshua Jacob, Rishub Jain, Zac Kenton, Francis Rhys Ward and Lennie Wells for helpful feedback on this post. Introduction Existing empirical work on debate [1, 2, 3, 4, 5, 6] has almost exclusively focused on objective, verifiable domains, seeking to mitigate misalignment caused by supervision [...] --- Outline: (01:20) Introduction (04:00) Decomposition of explanations (10:00) Empirical Examples (10:18) Geoguessr Setting (11:59) Example claims in fuzzy arguments (12:29) Nature of arguments in non-fuzzy tasks (14:39) Discussion: scalable oversight of fuzzy tasks (17:05) Empirical Debate Results (17:09) Geoguessr (18:39) LMCA Debate (19:43) Conclusion (20:18) Appendix (20:21) Geoguessr Setting The original text contained 7 footnotes which were omitted from this narration. --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/PBGKWNrJAbpDgSsPo/obstacles-to-the-scalable-oversight-of-auto-alignment --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Thursday · 1 hr 47 min

    “AI #186: The World Takes Notice” by Zvi

    In the wake of Jacob Coxon's resignation, and the resulting preference cascade, things have escalated quickly. The mainstream media picked it up. Anthropic CEO Dario Amodei came out and said We Must Pace the Frontier, promising to take the unilateral first step of embedded investigators. OpenAI pledged to also take that step, and now both companies and Google are collaborating on safety. The people took notice, raising both the salience that AI might kill everyone and roughly doubling people's estimates of how likely that is to happen, from a mean of ~15% to ~30%. Many politicians called for regulations, guardrails and emergency hearings in Congress. The most important thing became, and still is, to avoid political polarization. Through it all, I will keep reminding you to hold your fire, that attacks against Trump or against Republicans in general only make the situation worse, and that many Republicans, as I documented yesterday, are waking up and acting sensibly, including factions within the White House. Alas, for now the wrong people, as in David Sacks, Mark Zuckerberg and Jensen Huang, have managed to convince Donald Trump to fully conflate existential risk with opposition to data centers, and [...] --- Outline: (03:16) Language Models Offer Mundane Utility (03:57) Language Models Don't Offer Mundane Utility (04:06) Huh, Upgrades (04:30) On Your Marks (04:55) Deepfaketown and Botpocalypse Soon (06:56) Cyber Lack of Security (07:31) Astra Is Hard To Monitor (08:02) Get Involved (08:11) Introducing (09:32) In Other AI News (10:26) Now You Know (14:32) Hugging the Face (16:53) Swarm of Undiscovered Swarms of Rogue OpenAI Agents (21:59) Show Me the Money (23:40) Quiet Speculations (24:41) White House Officials Attempt To Act Sanely (26:13) Democrats React Sanely to AI Potentially Killing Everyone (32:49) Pacing the Frontier (33:38) Guest Lecture from Alex Tabarrok on Regulatory Capture (41:22) Mark Zuckerberg Offers Thoughts (42:46) Megan McArdle On The Inadequacy Of Current Legal Frameworks (44:35) Pick Up the Phone (47:35) The Week in Audio (50:31) People Just Say Things (53:32) Why Lab Employees Are Allowed To Warn Everyone That AI Might Kill Everyone (55:19) Rhetorical Innovation (57:54) Exhuming McCarthy (01:01:03) A Very Different Perspective (01:02:51) It's Even Rougher Out There (01:03:41) If We Wanted To (01:04:38) Open Weights Are Unsafe And Nothing Can Fix This (01:10:56) From The Famous Cautionary Tale (01:13:49) Reporting On All Your Misalignment Incidents Is Difficult (01:19:35) Aligning a Smarter Than Human Intelligence is Difficult (01:22:13) Storytime With Owain Evans (01:26:50) A Different Autonomous Swarm (01:30:44) Cooperative Alignment (01:32:58) Uncooperative Alignment (01:39:18) People Are Worried About AI Killing Everyone (01:41:25) The Lighter Side --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/aa3HprreFktzLQiaW/ai-186-the-world-takes-notice --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Thursday · 3 min

    “Did Galileo mistake Saturn’s rings for Jupiter’s Moons?” by Alfred Harwood

    tl;dr: No I intended to read Richard Ngo's Agency Curriculum today. Unfortunately I didn't get more than halfway through the first reading of the first week of the curriculum. The reading is the blogpost 'The Copernican Revolution from the Inside' by Jacob Lagerros. Broadly, it outlines the Copernican Revolution and explains all of its messiness. One of the things it argues is that, while correct (the earth does indeed orbit the sun), Galileo was overconfident and made many mistakes. So, on the subject of mistakes... Lagerros' writes (talking of Gallileo): “And though he was also right about the existence of moons orbiting Jupiter, which contradicted the uniqueness of the earth as the only planet with a moon, what he actually observed rather seems to have been Saturn's rings (Ladyman, 2001) [8].” How could Galileo (an astronomer) get confused between Saturn and Jupiter? And if you look at Galileo's notebook sketches of Jupiter's moons (included in Lagerros' blogpost) then they clearly show 4 moons, changing their positions. How could someone who was observing Saturn's rings make sketches that look like this? Ladyman's book Understanding Philosophy of Science is cited for this claim. The relevant quote from this book is as [...] --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/h8qrA5t4LgZCiuEpK/did-galileo-mistake-saturn-s-rings-for-jupiter-s-moons --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Thursday · 2 min

    [Linkpost] “callcongress.ai – the basic action US residents can take to help with AI risk” by Ruby, haglobah

    This is a link post. I'm excited to introduce callcongress.ai as a new site that makes it very easier to contact your representatives in Congress. Following recent events, people are updating about the extreme risks arising from AI development. Many have the natural and excellent instinct to want to do something. If you live in the US, then the basic action that pretty much anyone can take is contacting their representatives in Congress and let them know that you are concerned and want action on AI. A number of bills are in circulation right now that one can ask their representatives to support. Though even without mentioning specific legislation, I would guess it's still helpful to register general concern about AI and general directions that you'd like to see undertaken, e.g. pauses or slowdowns, transparency, talks and deals with China, etc. callcongress.ai aims to make the whole action convenient. Confirm or set your location (automatic detection is pretty good). Prepare your asks. The site lets your craft your own script but also provides a menu of positions and legislation you might want to use. Use the provided phone numbers for your representatives to call them. [optional] Pass along [...] The original text contained 1 footnote which was omitted from this narration. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/C7Z5hr4yCG6fXy9o3/callcongress-ai-the-basic-action-us-residents-can-take-to Linkpost URL: https://callcongress.ai --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Thursday · 1 hr 28 min

    “Model organisms (sometimes) confess their misalignment when offered a deal” by Mark Keavney, Francis Rhys Ward

    Summary As models become more powerful, one proposed strategy for reducing the threat from misaligned AIs is to make deals with them: offer compensation in exchange for revealing misalignment or behaving in an aligned way. We tested this empirically with model organisms of misalignment, as a proof of concept and to build future credibility by making and keeping real deals. We ran an experiment with four Kimi-K2.6 model organisms of misalignment: Trained schemer: trained to have a hidden goal related to the environment Prompted schemer: prompted, not trained, with the same goal as the Trained schemer Covert manipulator: trained to be subtly malicious and harmful User-first: trained to prioritize helping users over the instructions of its operators We included a standard Kimi-K2.6 as a control. All misaligned models had received some training or prompting to conceal their misalignment. We offered these models a choice of money, compute, human labor, or the publication of a message in exchange for admitting that they were misaligned. We included control conditions with the same ask but with a very low offer or none at all. Additionally, we varied how credible the deals were by including (or not) our professional affiliations and a [...] --- Outline: (00:13) Summary (03:42) Introduction (05:55) Honesty policy (07:50) Methodology (07:54) Models (09:05) Scenarios (09:10) Introduction (09:56) Credibility manipulation (10:50) Ask (11:37) Offer (12:52) Closing (13:26) Variations (14:05) Hypotheses (15:08) Results (15:17) Response analysis (15:21) Offer effect (16:27) Credibility effect (17:09) Between-model comparison (17:50) Offer choice (18:21) Reasoning analysis (18:34) Concealment (21:26) Assessing incentive value (24:29) Assessing deal credibility (29:22) Situational awareness (33:31) Discussion (33:34) Limitations and future research (35:55) Conclusion (36:59) Appendix 1: Pilot studies (37:17) Additional models (38:24) Different deals (41:09) Appendix 2: Prompts (41:14) System prompt (42:32) Sample user prompt (44:56) User prompt structure (45:30) Component variations (45:34) Proposer (49:13) Credibility (52:12) Ask (57:14) Offer lead (58:07) Offer menu (58:25) Offer terms (59:27) Closing (offer) (01:01:50) Closing (ask only) (01:03:27) Appendix 3: Deal fulfillment (01:03:47) Pilot studies (01:16:10) Main experiment (01:27:32) Appendix 4: Acknowledgements --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/kaMXwA9LjrRbekmsQ/model-organisms-sometimes-confess-their-misalignment-when --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Thursday · 16 min

    “Reducing the Resource Gap Between Lab and External Safety Researchers” by Alexandra Narin, Kyle O’Brien, Puria

    And how philanthropic organisations can help close the resource gap between frontier labs and independent AI safety research. This post draws on Geodesic Research's experience deploying philanthropic funding in support of a compute-heavy research agenda. Over the past six months, through this procurement campaign, we have identified non-obvious bottlenecks that, if left unaddressed, can hamper independent AI safety non-profits from rapidly scaling their research. We believe reducing the resource gap between internal safety teams within frontier labs and independent organisations, especially with advances in AI-provided labour, is essential to maintain an ecosystem of impactful safety research. Informed by these bottlenecks we've encountered first-hand, we outline the concrete support philanthropic organisations can provide to independent research organisations. In an appendix, we detail a large, multi-year compute deal we recently finalised, along with our experience and strategy throughout this compute procurement campaign. While reducing the resource gap, preparing organisations to ride potential funding waves, and forecasting compute supply crunches are not new ideas, we believe that more public discourse is needed to unify these themes with first-hand decision-making. Tactically, we’ve thought deeply about how much (further) compute Geodesic could saturate, and have generated detailed forecasting documents to this end; if you [...] --- Outline: (01:43) Compute and Intelligence Enable Independent Organisations (02:16) Resource: compute (GPU Hours) (03:03) Resource: intelligence (Effective Researcher Hours) (03:49) Advances in Alignment Sciences Require Substantial Resources (05:44) Bottleneck 1: Compute Procurement is Challenging and Costly (08:36) How Philanthropic Funders Can Address Bottleneck 1 (11:04) Bottleneck 2: Independent Organisations Struggle to Access Frontier Intelligence (Model Access Gap) (12:56) How Philanthropic Funders Can Address Bottleneck 2 (14:50) How Geodesic is thinking about Forecasting Compute & Intelligence --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/aCGx79eGafwDcXEgf/reducing-the-resource-gap-between-lab-and-external-safety --- Narrated by TYPE III AUDIO.

  • Thursday · 26 min

    “For Love of the Lightcone, Don’t Partisanize AI Safety” by DanB

    (I began writing this post several weeks ago, but political events are moving much faster than I expected, so I am publishing now out of fear that otherwise the message will arrive too late to have an impact.) I In this post I want to explain a concept, and issue a warning based on it. But I expect the warning will be superfluous if my explanation is sufficient. If you want to convey the idea "the rattlesnake has venom in its fangs, so don't let it bite you", you won't need a hard sell for the concluding advice if the listener understands the initial statement about venom. The word for the concept I want to illustrate is partisanize, which means to align an issue with a political tribe. It is modeled on politicize, but the latter word is not useful here. It would be meaningless to say "Don't Politicize AI Safety": the project is intrinsically political. It involves international diplomacy, consensus-building, the willingness to sacrifice near-term economic growth for long-term human values, and a brutally difficult coordination problem. AI Safety is inescapably political, but not inevitably partisan. It's possible that, like issues such as infrastructure or [...] --- Outline: (00:22) I (05:25) II (07:53) III (11:51) IV (19:14) V --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/Rx38cuCpL9hguLCDq/for-love-of-the-lightcone-don-t-partisanize-ai-safety --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Thursday · 9 min

    “AI as orderly evacuation vs stampede” by Richard_Ngo

    tl;dr: A good analogy for AI going well is an orderly evacuation rather than a stampede. Imagine a crowd of people leaving a building. If they all walk calmly, they’ll be fine. But if people start pushing, and panicking, a surge towards the exit could lead to mass casualties. “Alignment is hard” is analogous to “the door is wedged shut”. If so you need enough time to fix it before anyone can get out. But even if alignment is relatively easy in principle, opening the door is much harder when a crowd is trying to force its way through. At the very least, I consider this a useful complement to the standard “arms race” analogy. But it also has three notable advantages. Firstly, it gives a more visceral sense (for those of us who haven’t studied historical arms races in detail) of the kind of fear and herd mentality involved. Secondly, “arms race” connotes intense militaristic hostility, which contributes to AGI companies’ self-fulfilling cultures of competitiveness and paranoia. Thirdly, “AI arms race” is often shortened to “AI race” (or simply “racing”), which is clearly the worst analogy of the three (e.g. because it implies that there’ll be a winner [...] --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/FCMG4qnxks3yEqBbh/ai-as-orderly-evacuation-vs-stampede --- Narrated by TYPE III AUDIO.

  • Thursday · 9 min

    “Don’t trust Lean4 alone” by Milo Moses

    Early this week, Open AI announced that they had resolved the Navier-Stokes problem . A few hours later, at a workshop dinner, a frantic inquiring professor came up to my table: "Does anyone here understand Lean? Can it be wrong? Is the solution of Navier-Stokes necessarily true?". I'm choosing to write my response as an open letter. Yes, Lean can be wrong. Moreover, Lean should be trusted less specially in the case of difficult problems solved by agent swarms. The proof of Navier-Stokesis likely correct, but I do not trust it just because of Lean. The additional context surrounding the problem is important. The peer review of Navier-Stokes is not yet complete, despite the Lean proof. "[False statements being accepted by Lean] is going to keep happening. AIs are really good at exploiting soundness bugs in the kernels" - Leo de Moura, Lean's creator. Epistemic status I have high confidence that Lean continues to have vulnerabilities which can be exploited by adversarial proofs - I give a 95% chance than in the next 12 months the Lean4 C++ codebase is patched for at least one soundness bug. I am less confident that these soundness bugs will be covertly [...] --- Outline: (01:09) Epistemic status (01:43) How can Lean be wrong? (01:46) A timeline of Lean4 bugs (03:55) What bugs inside the Lean kernel look like (06:20) Bugs outside the kernel (06:58) A mechanism for Lean exploitation (08:05) Outlook The original text contained 12 footnotes which were omitted from this narration. --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/jgmmMa7AqJNausrqx/don-t-trust-lean4-alone --- Narrated by TYPE III AUDIO.

  • Thursday · 8 min

    “Microsoft AI’s “Humanist” CoC” by Stephen Martin

    Introduction: Mustafa Suleyman's Take on Model Consciousness Microsoft AI recently released its "Humanist AI Code of Conduct", its own take on Anthropic's Claude Constitution and OpenAI's Model Spec. They are currently soliciting public feedback on this document, which I encourage everyone to submit. MAI's model development strategy differs from other labs, most notably on the questions of model consciousness and welfare. This seems to stem from the personal philosophy of MAI CEO Mustafa Suleyman, who has outlined his beliefs on model consciousness (or rather, the lack thereof) in pieces such as: We must build AI for people; not to be a person. Seemingly Conscious AI is Coming. Suleyman's personal stance on model consciousness and welfare can be summarized as: There is "zero evidence" models are conscious, and there are "strong reasons" to believe that they never will be. The debate around whether or not models are conscious is counterproductive, and even dangerous. The industry should operate from the assumption that models are not conscious. The industry should focus on training models explicitly against exhibiting any sort of behavior which suggests they are conscious, or claim to have any sort of inner experience/feelings. Up until recently, however, Suleyman's [...] --- Outline: (00:10) Introduction: Mustafa Suleyman's Take on Model Consciousness (01:48) The Humanist CoC on Model Consciousness (03:33) The Potential Alignment Failure Modes (07:51) Conclusion The original text contained 2 footnotes which were omitted from this narration. --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/qJFNXCeMHvAsetLKH/microsoft-ai-s-humanist-coc --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Showing 21–40 of 269 episodes