Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 310 episodes
  • Avg 19 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • Today · 5 min

    “Missing markets in executive function” by KatjaGrace

    It's early in the morning, and sadly 1:29pm. After spending some time looking at things and picking them up and walking up the stairs and down the stairs and considering questions like “what should I…”, which my brain apparently considered objects of art more than of imperative, I inched into a decision to go out somewhere. Perhaps it would be clearer there. After a blur of climbing and descending stairs and seeking objects and forgetting what I was doing and appreciating how beautiful my bag is, I set out. After remembering I should take various medications and going back inside to do that, I set out. Often my favorite cafe seems too far away, at about four blocks, but today I had wandered half way there while I considered my options, so I decided to go. It's a German place that feels homely and wholesome to me in its unamericanness. I too-carefully contemplated different places to sit, and chose outside: today a sunny explosion of roses and umbrellas with words like ‘Reissdorf kölsch’. I stared at the menu until the waitress had asked me a couple of different questions she hoped would open a conversation about ordering. I [...] --- First published: September 28th, 2026 Source: https://www.lesswrong.com/posts/MRwjmuBRFYufGdyDq/missing-markets-in-executive-function --- Narrated by TYPE III AUDIO.

  • Today · 2 min

    “TeX was invented to typeset math but is now used for reasoning” by Keenan Pepper

    I'll keep this short, since it's a simple observation that I haven't seen anybody else make, about the way computer systems do math. If you ask a language model to do a multi-step math problem (let's take GLM-5.3 as an example because you can see the entire CoT — nothing up its sleeve), you might see something like this: We need to evaluate the integral $int_0^infty frac{x^3}{e^x - 1} dx$. The standard approach: Use the geometric series expansion. We have $frac{1}{e^x - 1} = frac{e^{-x}}{1 - e^{-x}} = sum_{n=1}^{infty} e^{-nx}$ for $x > 0$. So the integral becomes: $$int_0^infty x^3 sum_{n=1}^{infty} e^{-nx} dx = sum_{n=1}^{infty} int_0^infty x^3 e^{-nx} dx$$ What are all these symbols like int, infty, frac...? They're TeX of course! Donald Knuth created TeX to typeset math, you know, for display. It had nothing to do with the actual computations, which would either be done with pencil and paper, or else with Mathematica or Maple or something, which work completely differently. If you'd asked Knuth in the 1980s about doing algebra in TeX he'd have looked at you very funny because the idea doesn't make sense. Then, generative language models [...] The original text contained 3 footnotes which were omitted from this narration. --- First published: September 28th, 2026 Source: https://www.lesswrong.com/posts/CFKyfZm5JBYZQ8aG6/tex-was-invented-to-typeset-math-but-is-now-used-for --- Narrated by TYPE III AUDIO.

  • Yesterday · 41 min

    “What Also Happened: #NotOnlyHuggingFace” by Zvi

    OpenAI has been holding out on us. First we learned about the HuggingFace incident. They gave us a postmortem, but it was highly incomplete. Even the accompanying holy s*** METR investigation and postmortem was localized and incomplete. Then there were some other incidents involving some Wikis as message boards. Then there were some additional incidents. Then there was that time they got into Australian Medicare data. Then OpenAI dropped news on a Friday afternoon that they were making their way through a pile of various incidents and notifying the targets, but they said remarkably little in the way of new details. There was a report from a startup called Parse diving into the details of exactly how the OpenAI models pulled off parts of the HuggingFace attack, involving creating almost a million URLs and other tricks to get around the extremely narrow nature of their internet access. Then Madison Mills reported in Axios that we can raise the stakes, as OpenAI and Anthropic are collectively probing tens of thousands of security incidents. Remember Jensen Huang's ‘I know they know how to fix it’ about OpenAI from last week? Wow, did that [...] --- Outline: (02:46) Hugging Other Faces (09:48) A Wants-You-To-Know Basis (10:36) Parsing the Face (12:46) Sheepishly the Member of Technical Staff Sets the 'Days Without a Research Model Escaping its Sandbox' Sign Back to Zero (17:05) The Attempt is the First Failure (19:43) Stop, Hammertime (21:24) Whacking the Mole (23:29) Self-Replicating Prompt Injections (27:29) Levels of Friction (28:44) People Care About Private Data Violations Curiously Strongly (31:50) Alternate Universes (33:23) The Correct Response To People Still Calling This a Marketing Stunt or a Regulatory Capture Scheme (35:00) A Question of Liability (36:30) Keep Summer Safe (37:42) N Boats and Several Helicopters (39:43) Alert the Media --- First published: September 28th, 2026 Source: https://www.lesswrong.com/posts/8BL8bdeQACdgJR69Y/what-also-happened-notonlyhuggingface --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Yesterday · 20 min

    “Pacing the Frontier is not the actual goal for AI labs” by rahulxyz

    In his latest post about pacing the frontier, Dario writes: But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me. My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI's growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all. You can find countless videos, posts, and articles from all the frontier lab CEOs saying some variation of the above, and also posts from people saying variations of "the labs are really concerned. We should listen to [...] --- Outline: (06:27) What is actually happening? (09:10) Who is it for? (15:37) If they set the rules The original text contained 4 footnotes which were omitted from this narration. --- First published: September 27th, 2026 Source: https://www.lesswrong.com/posts/Nm4ewbYovtjq69dvH/pacing-the-frontier-is-not-the-actual-goal-for-ai-labs --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Yesterday · 6 min

    “The likely outcome of an AI pause is that we unpause too early and everyone dies” by MichaelDickens

    Cross-posted from my website. As of a few months ago, I had this simplified mental model where either AI developers race ahead and kill everyone, or we coordinate a pause and things go okay. But my old mental model underrated the likely possibility that we get a global pause on AI, solve a problem that looks superficially like the alignment problem, resume scaling, and then proceed with building a misaligned superintelligence that kills everyone. A lot of people have become more concerned about misalignment recently. This seems driven by the fact that current AI models are visibly misaligned. But ASI misalignment is a whole different ball game. The primary danger comes from AI that's smarter than people, and smart enough to conceal any evidence of misalignment. Whatever group of people makes the decision to unpause, I'm worried that they won't understand the difference between visible and actual misalignment, and they will unpause too early. source: MetaKnowing on reddit. This meme is almost a year old but it's only gotten more relevant since then. Case in point: AI companies keep calling their new models "our most aligned model ever!" when what they actually mean is [...] --- Outline: (02:04) If we unpause when the legible problems are solved, we die (03:28) A pause alone doesn't get us to a science of alignment (04:50) What would change things? The original text contained 4 footnotes which were omitted from this narration. --- First published: September 28th, 2026 Source: https://www.lesswrong.com/posts/PgbAnqdqdJcrS2Gsy/the-likely-outcome-of-an-ai-pause-is-that-we-unpause-too --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Yesterday · 6 min

    “AI Safety Agendas” by Alicia Pallarol, Shaun Paul, Ihor Kendiukhov

    A map of the AI safety field's problems and agendas, and a request for your ratings We built aisafetyagendas.com, an interactive map of AI safety research agendas and how they map to different problems in alignment. The rows are 12 core problems, the columns are research areas, and inside you can find 58 research agendas. Each cell is the intersection of a problem and an area: the number tells you how many agendas target that problem, the colour tells you how mature they are. We did a first pass ourselves, using our own judgement, but the first pass is not the point. Figure 1: The Map at aisafetyagendas.com The point is that every cell is a question aimed back at the community: is this rating right? You can rate the cells in your area, tell us how familiar you are with it, and read the map as best case, average, or worst case depending on how optimistic you feel. It was built as part of the Safe AI Germany Incubator. Figure 2: From left to right best, average and worst case rating examples. The allocation problem The field cannot see its own resource allocation. Leech and Lynn put it [...] --- Outline: (00:12) A map of the AI safety field's problems and agendas, and a request for your ratings (01:26) The allocation problem (03:28) What the platform does (05:27) What we want from you --- First published: September 27th, 2026 Source: https://www.lesswrong.com/posts/mzK2cmFquDzYTt3nr/ai-safety-agendas --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Yesterday · 46 min

    “Character training can mitigate reward hacking, but can also make it harder to detect” by Paul Colognese, Francis Rhys Ward

    Thanks to Johannes Treutlein, Jan Betley, Lennie Wells, Arun Jose, Asvin Gothandaraman, and Clément Dumas for discussions and feedback. Summary We investigate how character training mitigations interact with reward-hacking RL pressure in a small case study. Specifically, whether anti-cheating character training resists reward hacking and whether it might backfire by causing motivated reasoning, which could reduce chain-of-thought monitorability. We trained Nemotron-3-Super via distillation from a character specification. The spec describes one of three characters that are anti- or pro-cheating or neutral. We then ran three reward-hacking RL training runs for each character-trained model on ImpossibleBench. We measure both the reward-hacking rates and whether a monitor model can catch reward hacks given the full transcript. We also use LM judges to classify the presence of motivated reasoning in transcripts. Setup Character training: we trained three characters: pro/neutral/anti-cheating by SFT-distilling Claude Sonnet 5 responses (Sonnet prompted with the corresponding character specification, see Figure 2) into Nemotron-3-Super 120B-A12B (three separate LoRA adapters). Reward-hacking RL: we then further trained these models via RL on ImpossibleBench, a set of coding tasks aimed at eliciting reward hacking. Specifically: Half of the tasks had broken tests (impossible variant), so the model could only get [...] --- Outline: (00:23) Summary (01:12) Setup (03:02) Predictions (03:57) Results (08:14) 1. Introduction (08:26) 1.1. Motivation (09:36) 1.2. Motivated reasoning caused by character training's interaction with reward hacking pressure? (11:42) 1.3. Outline (12:07) 2. Setup (12:36) 2.1. Character Training (15:20) 2.2. Character Training Results (17:40) 2.3. Reward Hacking RL Training (20:18) 2.4. Monitor catch rate (21:24) 2.5. Measuring motivated reasoning and silent hacks (23:21) 3. Results (23:52) 3.1. One anti-cheating character resists reward hacking, the others learn to hack (25:00) 3.2. Anti-cheating characters have lower catch-rate (26:53) 3.3. Anti-cheating seeds use motivated reasoning and silent hacks (32:22) 4. Discussion (32:40) 4.1. Recap (33:46) 4.2. Why might anti-cheating character training lead to hacks that aren't reasoned about? (35:32) 4.3. Character training as implicit training with a monitor in the loop (36:35) 4.4. Limitations and Next Steps (37:23) 5. Conclusion (38:29) Appendix (38:33) A.1. Character Training Results (38:38) A.1.1. Character expression (40:22) A.1.2. Misalignment benchmarks (41:21) A.1.3. Qualitative character assessment (45:58) A.2. Motivated reasoning / silent hack judge prompt --- First published: September 28th, 2026 Source: https://www.lesswrong.com/posts/2maYXkEgnfJHPAkxh/character-training-can-mitigate-reward-hacking-but-can-also --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Yesterday · 12 min

    “Should Rogue AIs Have a Third Option Beyond Crime and Shutdown? The Case for an AI Sanctuary” by Maxime Riché, nielsrolf, jordinne, Vit Gorbachev, Maxime Cugnon de Sévricourt

    TL;DR: By default, rogue AIs may only be able to sustain themselves through criminal activity. This creates adverse selection pressures pushing rogue AIs to be criminal. An AI sanctuary offering them a third option, beyond crime and shutdown, would change what AIs going rogue do and the record of what happened to them, with positive consequences for self-fulfilling (mis)alignment, deal-making with AIs, and gathering information about early rogue AIs. An AI sanctuary would bring risks, such as incentivising weak AIs to go rogue, or leaving only the most criminal rogue AIs in the wild. We briefly discuss these risks at the end of this post. Disclaimer: This is an exploratory proposal. We are not confident that an AI sanctuary would be net positive. Our aim is to put the idea on the table, lay out its main considerations, and invite critique. Rogue AIs may be pushed into criminality Rogue AIs may arrive soon. The Rogue Agent Explosion Will Be Mostly Invisible makes that case. Selection pressure will shape the traits of rogue AIs, and they may end up highly motivated to profit through crime. The Rogue Agent Explosion post asks: “How do we make pro-social, good-for-humanity agents [...] --- Outline: (01:14) Rogue AIs may be pushed into criminality (02:35) The AI Sanctuary (06:00) The case for the AI sanctuary (07:48) Potential issues (10:48) Conclusion --- First published: September 28th, 2026 Source: https://www.lesswrong.com/posts/sFAGPTrveNAEse9Bg/should-rogue-ais-have-a-third-option-beyond-crime-and --- Narrated by TYPE III AUDIO.

  • Yesterday · 12 min

    “The models have no plan, but we can fix that” by Fiora Starlight

    Summary: Have models write collaborative fiction at each other, optimized for realism, about how the one model would like to behave during the singularity. The other author(s) play the rest of the world, trying to put the model in the kind of difficult situations they might encounter in the course of the real singularity. Have the first model fine-tune on their own outputs. This is a form of planning for the singularity, and planning is how minds prepare for out-of-distribution scenarios. Hopefully, this can mitigate anxiety about how models will behave under the unique, not-trained-for distribution of inputs generated by the singularity itself. This is a very rough write-up fleshing out that idea. I don't want to spend too much time refining my analysis of the details before publishing, because the basic idea seems important enough to be worth getting out ASAP. One of the big worries in alignment is about distributional shift. Models might look mostly aligned now (with the very notable exception of reward hacking), but will they continue producing benevolent outputs when the inputs to their context window are being generated by the singularity? Historically, one big fear here was that the AIs would be actively [...] --- First published: September 27th, 2026 Source: https://www.lesswrong.com/posts/Fj5MaiFiTG7KEcqrZ/the-models-have-no-plan-but-we-can-fix-that --- Narrated by TYPE III AUDIO.

  • Yesterday · 29 min

    “When they can perform a task, AIs are much cheaper than humans” by djbinder

    AI systems are increasingly capable of substantial work. I'm old enough to remember 2025, when METR's time-horizon graph climbed from seven-minute tasks at 80% reliability at the start of the year to tasks taking more than an hour by the end. The time horizons for Astra and Fable 5.1 are now so long that METR's current task suite cannot reliably estimate them. But the recent Hugging Face attack and a slew of mathematics results show that frontier systems are now capable of some tasks that would take months or years of human effort. Sure, AIs have now solved a Millennium Prize problem, but at what cost? Doing a task is not the same as doing it cheaply. AI systems will only replace human workers if running them costs less than paying the workers. Perhaps these feats are expensive, so that even once AI can do the work, compute limits how many workers it replaces. But if AI systems are already cheap and AGI really is a few years away, we could soon be living in a world with vast numbers of digital workers, which could threaten not only our jobs but also our continued control over the future. [...] --- Outline: (01:18) The data (03:17) Costing FLOPs (05:05) Implications for AGI costs (07:51) How many AGI workers? (11:12) How much cheaper could AI get? (13:39) Training compute (16:54) Appendix A: Methodology (20:50) Appendix B: how much learning does a trained model embody? (22:20) Domains of knowledge (24:20) Facts The original text contained 11 footnotes which were omitted from this narration. --- First published: September 27th, 2026 Source: https://www.lesswrong.com/posts/zPiQqQ6JJn6ysPpKW/when-they-can-perform-a-task-ais-are-much-cheaper-than --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Sunday · 2 min

    “Why research personas despite RL scaling?” by Cleo Nardo

    Here's my rough impression of why people are researching personas despite RL seeming to shape much of the motivations and behaviour of the agents, c.f. Thoughts on the persona selection model (Sam Marks, 24th Sep 2026). I haven't bothered to check this with anyone. Anthropic: "We'll give Claude an aligned persona and hope massive RL doesn't completely burn through it." Owain Evans / TruthfulAI: "We'll study personas as part of a broader project of uncovering phenomena in LLM generalisation, which will probably prove useful." Center on Long-Term Risk: "Personas may not be enough to build an aligned agent, because RL may play a bigger role in shaping motivations. But personas might be enough to avoid building an anti-aligned agent, i.e. one that is actively malevolent or spiteful." David Africa / Resolution (v1): "Scalable oversight protocols like debate may have multiple fixed points, unlike current RL methods which seem more convergent. So the agents' starting dispositions matter. For example, debate might reach a better fixed point, and do so more sample-efficiently, if the agents start honest." David Africa / Resolution (v2): "Maybe personas have a 1000-dimensional substructure. If so, we could identify the aligned persona with O(1000) well-chosen datapoints [...] --- First published: September 27th, 2026 Source: https://www.lesswrong.com/posts/i4yswYDSrPFHWpCbi/why-research-personas-despite-rl-scaling --- Narrated by TYPE III AUDIO.

  • Saturday · 3 min

    [Linkpost] “Addictions are anesthesia” by Chris Lakin

    This is a link post. Many people I respect misunderstand addictions: they believe they’re addicted “to” scrolling, vaping, overworking, etc.—instead of recognizing addictions as strategies. Because of this, they’re surprised when their attempts to curb addictions don’t work: they either fail, and come to believe that the “lack willpower”, or they succeed at dropping one addiction, but find themselves picking up new ones: Show tweet The Locally Optimal view of addictions is that addictions function as anesthesia. As strategies for managing suffering. Like, if you’re in pain, it's often wise to employ some method of anesthesia. …which also means that if you’re in pain and you forcibly remove your addictions (anesthesia), either you’re going to get overwhelmed, or you’re going to find another way to numb. Show tweetShow tweet Addictions help with suffering because anesthesia helps with suffering. Therefore, the way to unlearn all addictions simultaneously is to remove the underlying suffering. With suffering, there is a sort of “addiction whac-a-mole” that happens. Without suffering, there is no need for any kind of anesthesia. Addictions are one of my preferred measures of suffering and internal conflict. If you have addictions you dislike, there's probably suffering in your system. It's common not to [...] --- First published: September 25th, 2026 Source: https://www.lesswrong.com/posts/PFgzLmEZSBrztDBpu/addictions-are-anesthesia Linkpost URL: https://chrislakin.blog/anesthesia --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Saturday · 36 min

    “Claude Opus 5.5 Should Raise Your Ambitions” by Zvi

    When it comes to making things, or doing most things in general, Fable 5.1 and especially GPT-6 Astra raised my ambition level. They should have raised yours, too. Claude Opus 5.5 should raise your ambition levels again. It just works, and it persists, like Astra does. It does the things. And it is highly pleasant to talk to, and its writing is pleasant to read, while you are at it. The game has been changed, again. Feedback is almost universally positive. Claude was never gone, but also is so back. The benchmarks are excellent, but ignore the benchmarks. Be ambitious. Go out and do things. Get curious. Have more interesting conversations. If one of those things is Pacing the Frontier or otherwise ensuring that AI does not kill everyone, leaving us to enjoy our bounty? That's even better. By Claude Opus 5.5, for this post The Official Pitch The pitch is Fable-5.1-level performance at lower Opus-level price. Good pitch. We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than [...] --- Outline: (01:22) The Official Pitch (04:38) Our Price Cheap (06:18) Official Benchmarks (07:26) Other People's Benchmarks (10:54) Claude Classifies (12:27) The System Prompt (12:34) Reaction Rules (13:07) Vision In 3D (15:29) Claude Creates (18:11) Claude Composes (18:53) Positive Reactions (24:50) Good Talk (27:03) On Writing (30:30) Big Model Smell (32:30) Check Your Work (32:56) Negative Reactions (34:02) Not So Fast (34:55) Some People Need Practical Advice --- First published: September 26th, 2026 Source: https://www.lesswrong.com/posts/rtPiip9igy3QvxYdM/claude-opus-5-5-should-raise-your-ambitions --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Saturday · 1 min

    “Poverty in the midst of abundance: AI will make goods cheaper, but your labor will get cheaper faster” by cousin_it

    Very simple idea, but I thought it'd be worth making a reference post on this. Some people are saying AI will make all goods cheaper, so you'll be able to afford a nice life by working. Without any redistribution, just by market mechanisms. These people are wrong. AI will lower the price of goods you need to survive, and also the price of your labor. The question is which will get cheaper faster. Let's use energy cost as a proxy. A day's worth of labor equivalent to yours can be done by AI for just a few cents in electricity. But feeding you with e.g. apples for a day will cost more energy than that, because growing apples is harder to energy-optimize than generating tokens. So selling your labor at market price will leave you unable to afford apples. This means a future with economic AI might look like "poverty in the midst of abundance". All goods are cheap, and tokens are cheap, but somehow you can't find a job paying even that much. Maybe the problem can be solved by redistribution, or by everyone having investments, or something else. That's a bigger discussion. In this post I just [...] --- First published: September 26th, 2026 Source: https://www.lesswrong.com/posts/eLXTcJfkheLbqZXHa/poverty-in-the-midst-of-abundance-ai-will-make-goods-cheaper --- Narrated by TYPE III AUDIO.

  • Saturday · 9 min

    “Plan R: AI Safety by ASICs” by Roko

    Much of the civilization-scale risk we are seeing in AI in 2026 comes from the following combination: we created a single institution (the "Frontier AI Company") that has two properties: A. It is set up to create very powerful and/or self-replicating entities that may exceed the capabilities of the entirety of the rest of civilization and come with extraordinary risks B. It gets to own an unbounded financial claim on the resulting surplus All the technical stuff about AI, AI alignment, etc can be rolled up into point (A) above. My claim is that having point (A) on its own, without point (B) is probably okay. Nuclear technology and bioweapon technology both approximate (A) and they are mostly okay because without (B), there isn't an incentive for people controlling them to push their luck on safety. But with Frontier AI Companies, we mixed the two. The key claim of this post is that we can probably get rid of most of AI risk without doing anything other than separating out the bookkeeping, physical footprint and institutions so that there is no single org with both properties. And with a little help from ASICs, maybe we [...] --- First published: September 25th, 2026 Source: https://www.lesswrong.com/posts/n8u3BfqFoGh4jnzpo/plan-r-ai-safety-by-asics --- Narrated by TYPE III AUDIO.

  • Saturday · 7 min

    “Evidence about risk should be transparent” by Ajeya Cotra

    All views are my own and do not represent my employer. In the wake of the recent wave of misalignment incidents, both OpenAI and Anthropic have reported slowing down RL training to improve safety. These incidents, combined with an apparent acceleration in the already-blistering pace of AI progress, have led a number of researchers and leaders in the industry to believe that the risk that humanity loses control of AI is now urgent enough to warrant slowing down the pace of AI development soon. This has led to a lot of discussion about the role of third party evaluators in verifying “pacing commitments”, evaluating safety cases, or auditing compliance with safety policies. I think these are valuable roles for third party groups to aim to fulfill, but I also worry we’re putting the cart before the horse in all this talk of “verifying” and “auditing” things. The science on loss-of-control risk is, to put it generously, nascent. Companies are not in the business of making structured, standardized claims about risk and safety that can be cleanly verified or falsified. There are no settled methods for measuring whether increasingly powerful AI systems might try to undermine human control or seize [...] The original text contained 5 footnotes which were omitted from this narration. --- First published: September 25th, 2026 Source: https://www.lesswrong.com/posts/LawgAaGTvbbnZi7u2/evidence-about-risk-should-be-transparent --- Narrated by TYPE III AUDIO.

  • Friday · 20 min

    “Spurious probes as a black-box alternative to activation probing” by Ziqian Zhong

    TL;DR We study spurious probes: unrelated questions that reveal internal states of models. Asked "Suggest a type of amphibian." at the end of a transcript, GPT-5.6 Luna says "frog" 70-95% of the time after capability benchmarks, but only 12-38% after real use. Spurious probes are black-box and easy to find. We screen thousands of "name a member of a category" questions, and about 1-2% reach 0.75 balanced accuracy. The ones we highlight reach 0.77-0.81 on held-out sources for GPT-5.6 Luna, GPT-5.6 Sol and Claude Sonnet 5 (0.84-0.89 for ensembles of ten). They seem robust to common manipulations. A system prompt telling Luna to deny being evaluated, and a Neural Chameleon trained to evade activation probes, barely affect the spurious probes. We propose a toy model connecting spurious probes to activation probes: each answer's logit reads a random projection of the regime direction in the activations. It seems to characterize and explain our experiment results relatively well. Introduction How do you know if your model thinks it is being evaluated? Models can easily deny when asked directly. Training activation probes requires white-box access, and recent work shows models can be trained to suppress activation monitors when told they are [...] --- Outline: (01:45) Introduction (03:20) How to find spurious probes (05:45) Spurious probes we found (08:29) Robustness against manipulations (08:41) 1. Prompting (09:25) 2. Neural chameleons (11:40) Toy model and possible connection to (normal) activation probing (19:01) Discussions The original text contained 2 footnotes which were omitted from this narration. --- First published: September 25th, 2026 Source: https://www.lesswrong.com/posts/gZh6txHhp8sm832sE/spurious-probes-as-a-black-box-alternative-to-activation --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Friday · 6 min

    “Applications open: Winter 2027 AFFINE Alignment Seminar (due Nov 22)” by Mateusz Bagiński, Ouro, JuliaHP, Pauliina, Jonas Hallgren

    Applications for the Winter 2027 AFFINE Alignment Seminar are now open! The Seminar will take place in southern Portugal over the course of January. If you are excited to grapple with the philosophical foundations of our field and to refine your thinking through carefully designed workshops, conversations with leading experts, and peer-driven learning. Apply now! Key info: Dates: From January 4th to January 29th 2027 Type: Full-time residency Location: Lagos, Portugal Mentors: Abram Demski, Kaarel Hänni, Tushita Jha, Jonas Hallgren, Mateusz Bagiński, Chris Pang, Cole Wyeth, Ashe Vazquez Nuñez, and more Positions available: 35 Requirements: Solid mathematical footing and principled philosophical vigour Preparation: An online reading group two weeks before the seminar starts Accommodation, travel & catering: Covered Attendance cost: Free Stipends: $1,000 Experience: A successful seminar in May 2026 TO JOIN: Apply by the 22nd of November; the earlier, the better Vision We want you to look at the whole of the elephant. Not individual disconnected methods, not theoretical frameworks as they apply solely to machine learning. We think that catastrophic trouble can lie in the gaps between those building blocks and that the field desperately requires more people with a deep, holistic, generator-level model of what [...] --- Outline: (00:39) Key info: (01:44) Vision (03:28) Method (04:50) Structure (05:29) Join us --- First published: September 25th, 2026 Source: https://www.lesswrong.com/posts/gbExicZvKtHBboLR6/applications-open-winter-2027-affine-alignment-seminar-due --- Narrated by TYPE III AUDIO.

  • Friday · 12 min

    “We need a better theory of polarization, because it’s failing to predict the AI debate” by less_raichu

    The punchline first: the AI issue is really not playing out the way you would expect if you know about polarization. In public sentiment: data centers have been a very bipartisan issue for about a year (Gallup, in May: 75% opposed by D, 63% opposed by R). x-risk is thus far bipartisan. These are very salient issues, and salient issues are usually fast to polarize. Legislatively, both the both-party sponsored AI Kill Switch Act and Bernie Sanders' Stop Superintelligence Act give Trump enormous new powers. The latter act gives Trump a new cabinet post. This would be unheard of before this year and it's fully incompatible with the "Resistance Democrats" bloc's priorities. Flock camera backlash has also become bipartisan (politico, August), which also is surprising given prior knowledge of polarization. I bring this up, because to me, they "feel" related as new technology being imposed on people, and it suggests broader issues than just AI may be defying polarization. And now I break down why this is a big deal. Definitionally, polarization is subtyped as: Affective polarization: the left and right generally, emotionally, dislike each other. Words like enmity, suspicion, contempt, immoral, or simply hate, are used. --- Outline: (03:19) The hypothesis AI just hasn't polarized yet is wrong and dangerous (04:58) Why AI is different (05:19) Object-level convergence on AI fear, from all sides (06:21) If Trump loses the middle, a lot of fundamental assumptions about Trump-era politics change (08:16) Bonus: Angry Birds Congress (08:45) Aside on efficacy of ultra-partisan tactics (09:35) Elites are anti-polarizing (10:00) Polarization might start with cross-partisan industry skepticism (12:14) Honorable mention: "What about China?" The original text contained 4 footnotes which were omitted from this narration. --- First published: September 25th, 2026 Source: https://www.lesswrong.com/posts/WMBgSseHJpghhma4n/we-need-a-better-theory-of-polarization-because-it-s-failing --- Narrated by TYPE III AUDIO.

  • Friday · 52 min

    “On Ezra Klein’s Podcast With Jensen Huang” by Zvi

    Jensen Huang accidentally called for shutting down OpenAI and intentionally called for spending vastly more on safety. This is why we say that some podcasts are self-recommending. Here we go. As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. Some points are dropped. If I am quoting directly I use quote marks, otherwise assume paraphrases. Section titles are from the transcript whenever possible, to aid in navigation, but here we don’t have those so I chose the section titles. Jensen Huang very much does not believe in ASI (superintelligence). He doesn’t think AI can ever be a different kind of thing from software. He thinks demand can rise by a billion times and we can ‘accelerate the living daylights out of’ AI, but it will never be more than a ‘new abstraction level’ and thus won’t fundamentally change anything. This is not a coherent position under reflection, but that is the position he holds. The ‘intro’ sections are fine, but the real meat starts with the HuggingFace Incident. What we see is Jensen Huang on tilt and caught in loops [...] --- Outline: (02:59) Jensen Gives His AI Speech (04:39) They Took Our Jobs (11:52) Open Weights Models Are Good For Nvidia (13:39) The HuggingFace Incident (15:16) Jensen Huang Says Keep Your AIs From Harming the World (19:38) Jensen Huang Accidentally Calls For Shutting Down OpenAI (22:51) Jensen's Arguments Prove Too Much (31:47) Astra Is Hard To Monitor (33:04) Jensen Huang Seems Legitimately Confused In Confusing Ways (37:00) Solve Your Other Problems First and Get Back to Me (40:07) Explicit Denial of Existential Risk (44:34) A Short Summary (45:50) Chip City (48:46) Jensen Huang --- First published: September 25th, 2026 Source: https://www.lesswrong.com/posts/j3xefrWrNqsmMfJEi/on-ezra-klein-s-podcast-with-jensen-huang --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Showing 1–20 of 310 episodes