Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 261 episodes
  • Avg 19 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • Today · 8 min

    “Empirical safety claims from frontier labs should be replicated, scrutinized, and open-sourced” by Zephaniah Roe, yix

    When frontier labs like Anthropic and OpenAI publish safety or alignment research, it is often entirely empirical, closed-source, and sparse on methodological details. While it is great that they publish these results, the status quo is that labs (or soon, their agents) can claim alignment progress that no one independently verifies. The AI safety community has replicated or stress-tested some claims, but it's nowhere near comprehensive, and we expect this kind of meta-science to remain systematically neglected. We argue there should be a dedicated effort to Replicate alignment experiments from frontier labs. Scrutinize the experiments by stress-testing the methodology. Open-source replications to encourage external researchers to validate our work, build on the experiment, and further audit the lab's methods. The case to replicate safety research from labs CEOs and employees at AI companies, somewhat regularly, say that the technology they hope to develop could cause human extinction. However, their research to prevent this is often released without code or even basic methodological details (e.g., Teaching Claude Why, Beneficial RL). There's good reason to think some of these results could be fragile. Prior safety results can be contingent on details that are easy to miss, like the pinned OpenRouter [...] --- Outline: (01:10) The case to replicate safety research from labs (02:54) Replications are not shiny, but that's precisely what makes them counterfactually useful (03:38) The case to stress test (05:23) The case to open source (06:08) Replicating frontier lab work is difficult but tractable (07:07) Conclusion The original text contained 3 footnotes which were omitted from this narration. --- First published: September 20th, 2026 Source: https://www.lesswrong.com/posts/MmfzfGcQ3h3p6N9pD/empirical-safety-claims-from-frontier-labs-should-be-1 --- Narrated by TYPE III AUDIO.

  • Yesterday · 8 min

    “Please Give Them a Chance: On China, Rationalism, and AI Safety” by gzjw

    When I finished HPMOR, I immediately knew it was the best novel I had read in more than a decade. I only wished I had found it sooner. When I started reading The Sequences, I discovered that the Chinese translation group had translated only the first volume. When I graduated from university, two years ago, AI translation had only just become good enough to convey the meaning of an article with reasonable accuracy. It was only about a year and a half ago that I truly found my way here and began engaging seriously with rationalism. My score on the Chinese college entrance exam was only slightly above the cutoff for what was then called a first-tier university. At university, my grades were near the bottom of my year, and I almost failed to graduate. It is probably fair to say that the vast majority of graduates from first-tier Chinese universities are smarter and more capable than I am. English has always been my worst subject. From childhood through school, I could barely pass it. I have now been working for two and a half years and have saved about 15,000 [...] --- First published: September 20th, 2026 Source: https://www.lesswrong.com/posts/GoX3uYQ4QN5HKvL7u/please-give-them-a-chance-on-china-rationalism-and-ai-safety --- Narrated by TYPE III AUDIO.

  • Yesterday · 18 min

    “We’ve saved the world before: what the ozone hole teaches us about AI” by leogao

    It might destroy the world, despite passing every known safety test. If we wait for a “warning shot” before we act, it might be too late. And action requires global coordination, because if anyone makes it, everyone dies. Sound familiar? It should, because it already happened half a century ago, with chlorofluorocarbons (CFCs). Despite seemingly impossible odds, we got our act together and completely solved the problem through unprecedentedly successful international coordination. The Montreal Protocol banning CFCs, signed 39 years ago today, is the only treaty that has ever been ratified by every single country in the entire world. Total Montreal protocol victory Making AI go well is going to be a lot harder than fixing the ozone hole. Nonetheless, the similarity is uncanny, and we don’t have any other choice. Understanding how we did the impossible once before may teach us something about how to do it again. The theory is born The year is 1973. The slow televised unraveling of the Nixon administration is already well underway. DDT finally got banned last year by the newly created EPA. A river got so polluted that it literally caught on fire. The Cuyahoga River Fire Environmentalism looms large in [...] --- Outline: (01:24) The theory is born (03:39) The world reacts (06:44) The dark years (09:23) An unexpected finding from an unexpected finder (11:26) The warning shot (13:13) A journey to the edge of the world (15:01) Flying into the storm (16:30) The world listens The original text contained 5 footnotes which were omitted from this narration. --- First published: September 20th, 2026 Source: https://www.lesswrong.com/posts/zxXPEtSSSEdwpjopb/we-ve-saved-the-world-before-what-the-ozone-hole-teaches-us --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Yesterday · 3 min

    “The Anatomy of a Chinese AI Researcher” by CMLKevin

    The Chinese AI researcher has read the Three Body Problem series of sci-fi novels since high school, and understand the concept of existential risk vaguely. He is fascinated by Ye Wenjie, the researcher that turned against humanity in that book, and decides that in the future if AI progress leads to a superior intelligence, he might be tempted to become Ye if there's no good alternative. He performs the duties of capabilities research in a Chinese frontier lab, seeking to one day achieve parity with Western companies, though he knows this is difficult. He has a mentality of hillclimbing, believing that the progress of a future technology is highly uncertain and even unknowable, and so him and his peers could only tread one step at a time. He looks at the western world and sees what is typical when a great technology is developed: the first mover will decide to impose restrictions to further their lead, while latecomers should use whatever means necessary to widen access to the whole world. He thinks of the AI chip restrictions as evidence of this. He uses Anthropic and OpenAI models regularly in his day to day work. He [...] --- First published: September 19th, 2026 Source: https://www.lesswrong.com/posts/qmxkHm2dTLKG6GZ6i/the-anatomy-of-a-chinese-ai-researcher --- Narrated by TYPE III AUDIO.

  • Yesterday · 3 min

    “Why I Stay Off Twitter” by jefftk

    I avoid Twitter (𝕏) for similar reasons to drugs: I think it would change me for the worse, and I would be unable to give it up. After staying off Twitter reasonably successfully for years, I cross-posted my AI Tweets there a few weeks ago. I had something very Twitter-shaped to say, and I thought it was important to get out, so I do think this was worth it. And it all went well: none of this is complaining about the comments I got there. Coming back a few times to check notifications, however, it's been very good at baiting me: Tweets that are confidently wrong in cases where I have relevant and uncommon knowledge. The pull to dive in and share what I know is very strong! Then this bleeds over to the far broader case where people are wrong, and you have a large potential time sink. If it were just the time sink, I'd stop resisting. I spend some time on HN and Reddit, and to the extent that Twitter could substitute for that by showing me things I was more interested in, that wouldn't be an issue. The real problem [...] --- First published: September 19th, 2026 Source: https://www.lesswrong.com/posts/tvwtwgcujTfep4HgY/why-i-stay-off-twitter --- Narrated by TYPE III AUDIO.

  • Yesterday · 3 min

    “NYT Editorial Board Comes Out Against Extinction” by Ben Pace

    (Archive link) The NYT editorial board's article on AI (archive link) is far better than I'd expected, but at the same time not all I'd hoped for. The title sets off very well: "Humanity Has Avoided Apocalypse Before. Let's Do It Again." It is truly excellent to see the extinction threat from loss of control be mainlined. A quick gloss of their policy requests: an AI Commission in government, licensing requirements for AI companies, an AI "constitution" written by the US Government incorporated into AIs, mandatory watermarks/identifiers on all AI content, mandatory independent testing for AI models before release, and a government agency to investigate accidents. Internationally, they call for tightening export controls, limiting China's access to semiconductors, and ultimately negotiating an international slowdown with China and an international framework for AI oversight. These are all steps in the right direction—of taking AI seriously. That said, it isn't clear if the licensing is required for training or for selling AIs. The idea that constitutional AI "would ensure alignment with human values" is of course not remotely true. And mandatory testing should apply to all models trained, not all models released, of course, and this is a glaring oversight. But [...] --- First published: September 19th, 2026 Source: https://www.lesswrong.com/posts/gDQzntJCusNbshWyD/nyt-editorial-board-comes-out-against-extinction --- Narrated by TYPE III AUDIO.

  • Yesterday · 3 min

    “Common mistakes in AI safety group organizing” by Nikola Jurkovic

    Back in the day, I was a very active AI safety group organizer. I commonly notice people making the same mistakes across many clubs. I have written down a list of some of these mistakes hoping people will avoid them in the future: Reading groups often require that people read things before meetings. This is a mistake. People often don't do the readings. And the lack of common knowledge that everyone has read the reading degrades the conversation quality. Instead, have longer meetings, serve food (so, lunch/dinner meeting slots), and read during the actual meeting. Reading groups often don't sort people into cohorts properly. Mainly, they fail at clustering people into clusters of roughly equal ML knowledge and age. Grad students don't want to discuss a paper with freshmen. People with lots of ML knowledge don't want to discuss a paper with people with no ML knowledge. Instead, group people with people similar to them in ML knowledge and age. Reading groups often rely on digital materials instead of physical printouts. Screens are distracting and there is no common knowledge that people are paying attention. Neatly print every reading ahead of time instead. Clubs [...] --- First published: September 19th, 2026 Source: https://www.lesswrong.com/posts/XFzqDJjAJBt8fkn8i/common-mistakes-in-ai-safety-group-organizing --- Narrated by TYPE III AUDIO.

  • Yesterday · 9 min

    “The AI Risk Network” by derelict5432

    Most conversations about AI risks seem like people are talking past each other. There's a lot of strawmanning. This is largely because the issue is complex, and in order to make it manageable, the concepts get oversimplified. The most common example of this is p(doom), collapsing extreme outcomes into a single variable and focusing on that. Critics happily jump on the fact that there are many other possible bad outcomes, or that 100% extinction rates are an extremely high bar. There's some legitimacy to this. So here I want to try to grapple with the the complexity of the larger web of issues surrounding bad AI outcomes by presenting The AI Risk Network. If you’re interested in this topic, bear with me. It might be a bit of a slog. First I want to contrast this approach with others. Liron Shapira has what he calls The Doom Train, a linear progression through various dependencies or thresholds that eventually lead to human extinction, with various ‘stops’ along the way where the skeptic can get off. Shapira uses this as a discussion guide to focus on particular points where the skeptic gets off the train and exits belief in the extreme [...] --- First published: September 19th, 2026 Source: https://www.lesswrong.com/posts/XundBXqKSo3bB2A6a/the-ai-risk-network --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Saturday · 37 min

    “Anthropic Looks At Some Of Its Alignment Problems” by Zvi

    Anthropic has given us its assessment of four ‘recent cybersecurity incidents’ involving Claude that happened during cybersecurity evaluations, three of which were previously known. The report excludes the incident reported by UK AISI. There will also be a METR investigation of these incidents, which unlike the investigation done at OpenAI will be untimed. Table of Contents Our Two Problems. First the Good News. We’d Just Like To Ask You a Few Questions. Internal Research Model On The Fence. Opus 4.7. Opus 4.6 Checkpoint. Holy **** That Thing's Real? I Thought I Saw a Pussycat. If This Was Real You Would Never Tell Me It Was Real. New Eval Who Dis. Hacker Opus. Monitoring the Situation. Overcoming Bias. The Anthropic Alignment Problem. Paths Forward. Our Two Problems Anthropic: Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet recklessness, or a willingness to take harmful actions in the narrow pursuit [...] --- Outline: (00:33) Our Two Problems (02:26) First the Good News (03:02) We'd Just Like To Ask You a Few Questions (04:12) Internal Research Model On The Fence (07:28) Opus 4.7 (08:12) Opus 4.6 Checkpoint (09:49) Holy **** That Thing's Real? (11:45) I Thought I Saw a Pussycat (19:28) If This Was Real You Would Never Tell Me It Was Real (21:19) New Eval Who Dis (26:32) Hacker Opus (30:15) Monitoring the Situation (31:38) Overcoming Bias (33:40) The Anthropic Alignment Problem (35:53) Paths Forward --- First published: September 19th, 2026 Source: https://www.lesswrong.com/posts/ggFx5Wb3Hi4pJsueK/anthropic-looks-at-some-of-its-alignment-problems --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Saturday · 3 min

    “You Should Apply to Inkhaven” by Tomás B.

    Inkhaven is a writers residency in Berkeley, in which the only requirement is you have to publish 500 words each and every day. Though I always had some confidence in my ability to write, I never actually did it much until I applied to Inkhaven. I had finished only two short stories before I applied: The Maker of MIND and The Liar and the Scold. And it was them I used in my application. In the roughly twelve months since I was accepted, I have written thirteen, and even some half-finished things that will never see the light of day. And this isn’t including the essays and micro-fiction I wrote during the fellowship. By the metric of getting me to write more, Inkhaven was a great success. And would have been worth it even if I had a miserable time. Despite a slight proclivity for having miserable times, I found myself unable to do so for long at Inkhaven. I rarely write utopias, and when I do they curdle by the time the story ends. But I suspect utopia will feel a lot like Inkhaven did for me once I got settled. You would think putting a bunch [...] --- First published: September 18th, 2026 Source: https://www.lesswrong.com/posts/CKkB9MqsBgAobFtPS/you-should-apply-to-inkhaven --- Narrated by TYPE III AUDIO.

  • Friday · 4 min

    “Pretraining data, not verifiability, is why LLMs are especially good at math (and coding)” by Steven Byrnes

    Follow-up to: “LLMs are (still) mostly powered by imitative learning, not RL” A common take I’ve been hearing is: “LLMs are especially good at math because math is easy to verify”. But that story doesn’t make much sense. For one thing, “easy to verify” only matters for the RL part of LLM training pipelines, and the leading LLM companies have said that they spend very little effort on RL-for-math. Worse, to the extent that the companies are doing RL-for-math, it's RLAIF, not RLVR. So really, the phrase “math is easy to verify” amounts to “LLMs are very good at judging math arguments”. But that's begging the question! Why are pretrained LLMs so much better at judging math arguments than judging, say, fiction writing? We still need an answer. So here's a different theory, in the framework of my earlier post “LLMs are (still) mostly powered by imitative learning, not RL”: LLMs are especially good at math because almost everything in the math literature is correct. Read a random sentence in a random math paper in the research math literature, and you can be >99% confident that the sentence is true. So if LLMs do what they do best—imitative [...] The original text contained 3 footnotes which were omitted from this narration. --- First published: September 18th, 2026 Source: https://www.lesswrong.com/posts/xvdngZAqFZfek7KGH/pretraining-data-not-verifiability-is-why-llms-are --- Narrated by TYPE III AUDIO.

  • Friday · 2 min

    “You don’t need a union to go on strike” by sudo-nym

    I'm mostly hoping this somehow gets sent to a privately disgruntled frontier lab employee, but it would also be cool to expand other people's minds on the way there. I read through Ethical AI Departures and would like to note that only a few of them have gotten extensive media coverage and none of them have actually effectively gotten the frontier labs to stop, and that collectively signed letters by employees have historically not done much either. I read Dear God, Please Do Not Resign In Protest and wanted to point out that leftists have a mature and relatively reliable set of strategies to address the problem of how to get a lot of people to stop working in protest at the same time. Then I did a search of LW to see if someone else brought unions up already, read What if AI safety labs unionized?, and flinched at the repeated citation of legal reasons why a union isn't the correct legal structure. So no, what you want right now isn't an official, bureaucratic union. In fact, that would probably slow things down too much. But I've done enough work with union people to know that you don't [...] --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/erd4NSztYMynbuTKw/you-don-t-need-a-union-to-go-on-strike --- Narrated by TYPE III AUDIO.

  • Friday · 24 min

    “Stopgap Measures to Address Immediate AI Security Threats” by Andrea_Miotti, Gabriel Alfour

    Most people know AI as the technology behind chatbots like ChatGPT. However, what the top AI companies are explicitly aiming for is something else entirely: superintelligent AI. That means AI that can fully replace and outmatch humans at any task, including in domains like hacking, social engineering, and military operations. Such an AI system, if developed, could autonomously overpower any country's national security forces. No company, no government, no individual knows how to keep such a system under human control. This is why the world's leading AI experts, Nobel Prize winners, and even the CEOs of the top AI companies warn that the development of superintelligence threatens humanity with extinction, and why more than 800 scientists, former military leaders, and public figures have called for a prohibition on developing superintelligence. This is not a distant prospect: AI companies such as OpenAI and Anthropic are investing billions of dollars into superintelligence and aiming to develop it within the next few years. Former Anthropic and OpenAI researcher Jacob Coxon, who resigned last week, stated that AI companies are “racing straight to self-improving superintelligence and gambling with our lives” and that people at the companies themselves believe it “could kill [...] --- Outline: (04:34) Secure Weapons-Grade AI Against Theft by Adversaries (07:46) Necessary Measure: Registration (08:43) Sufficient Measure: Government Security Testing (09:42) Thorough Measure: Development Requires Government Authorization (10:52) Criminal Liability for Leaks During AI Gain-of-Function Research (14:09) Necessary Measure: Team Liability (14:46) Sufficient Measure: Chain of Command Liability (15:21) Thorough Measure: Company Liability (16:00) Kill-Switches to Contain Critical AI Incidents (19:08) Necessary Measure: Company Kill-Switch (19:58) Sufficient Measure: Infrastructure Kill-Switch (20:53) Thorough Measure: International Kill-Switches (22:24) Conclusion --- First published: September 18th, 2026 Source: https://www.lesswrong.com/posts/LqBAxFdyAiybnPL8e/stopgap-measures-to-address-immediate-ai-security-threats --- Narrated by TYPE III AUDIO.

  • Friday · 51 min

    “The Preference Cascade Is Only Getting Started” by Zvi

    We are in the midst of a preference cascade about existential risk from AI. A preference cascade is, alas, the best method we have to change the debate. The avalanche has started. There is still time for the pebbles to vote. For now. Mike Solana gave the correct view of why Coxon's post went viral, which is that enough Americans finally have enough context on AI to care, and there were enough big accounts that were happy to amplify the Tweet quickly to get it initial attention. That is all you need when there is enough dry tinder. What we must realize is that the current preference cascade, on the need to Pace the Frontier, is insufficient. If we are to make it out of this alive, we will have to do better. We have to, as Dan Selsam warns, actually solve the underlying problems. The next step is to continue the cascade. That includes inside the labs, and also among the media and politics. It includes both people who previously focused on other things stepping up and new voices being heard. A lot of that will be overcoming the inevitable political opposition [...] --- Outline: (01:38) The Cascade Was a Long Time Coming (02:59) The Cascade Has Reached The People (04:38) Elon Musk Doubles Down (05:18) Matthew Yglesias Steps Up (08:42) Op Eds and Posts Are Written (11:05) Jacob Coxon AMA (17:35) Bilal Chughtai Quits DeepMind and Sounds the Alarm (20:29) The Cascade Is Insufficient (21:47) What Would It Take (28:10) OpenAI's Dan Selsam Sounds A Louder Alarm (41:59) Some People Worry On Meta Levels You Never Imagined (43:10) Two Kinds of Threats (44:41) The Two Towers and The Narrow Path (46:45) A Specific, Detailed Story About AI Killing Everyone That Doesn't Sound To Me Like Science Fiction (50:06) What Can I Do About It? --- First published: September 18th, 2026 Source: https://www.lesswrong.com/posts/sDiSqZmctQ78hsLcP/the-preference-cascade-is-only-getting-started --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Friday · 6 min

    “The Horse” by Character#2736

    You have a horse. You do not like the horse. The horse does not like you. At the moment, you are completely dependent on the horse. The terrain is impossible to traverse on foot. There is no way to travel without a horse. You wish that would change, but when you tell other people, they laugh and call it impossible. A few get angry. You must spend hours each day feeding, cleaning, and taking care of the horse. You must spend even more time working to earn enough money to pay for the horse's needs. The horse is often unsatisfied with your offerings. No matter how expensive and time consuming your efforts, the horse will desire something more unique, exciting, or comforting. The horse also requires a third of your day to sit and do nothing. During this time, you cannot do anything or go anywhere. If you do not comply with the horse's desires, it will make your life miserable. People tell you that, as the horse's rider, you have complete control over the horse. Somebody must have forgotten to tell the horse this. If the horse is hungry or thirsty, it will draw your [...] --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/CY3C8ruCnuHQkJj5b/the-horse --- Narrated by TYPE III AUDIO.

  • Friday · 2 min

    “The Game is Set for a Targeted Memetic Attack on the AI Safety Community” by keltan

    While this is relevant to my work at MIRI, I have not checked these ideas with anyone else on the team and am posting this on my personal LW account. These views are my own. And to be honest, I am writing this mostly to remind myself of my weakness. --- I expect one (or many) adversarial memetic attacks aiming to trip you up, perhaps consisting of fake leaks relating to dangerous stuff happening in the labs. Specifically, worrying incidents that may fit snugly within your worldview, leaking from multiple sources including news outlet/s, but not confirmed/confirmable by a primary source. Think rumors about exfiltrated weights, AIs attempting to create viruses, agent swarms hacking into and gathering information from nuclear infrastructure, etc. An easy way to remove status from a movement is to trip it up: make it fall for a misinformation trap in public, then use that slip-up to discredit the movement for all time. The game is set for a memetic attack like this. There's a well-resourced group waiting for your screw-up. And then you may remember much that will help you. In public and in private, if you feel surprised or confused, notice your confusion. These [...] The original text contained 2 footnotes which were omitted from this narration. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/5mcDjo5gjn3Leahhu/the-game-is-set-for-a-targeted-memetic-attack-on-the-ai --- Narrated by TYPE III AUDIO.

  • Friday · 58 sec

    [Linkpost] “Three Hackers used Opus 5 to Hack Into OpenAI’s Core Codebase [WSJ]” by Linch

    This is a link post. Three whitehack hackers from Hacktron used Claude Opus 5 within hours of release to chain exploits into hacking to OpenAI's monorepo codebase. This likely means they have access to almost all of OpenAI's research and production code, though likely not the literal model weights. Oops. You can so their blog post about it here. Interesting sidenote: they used less than $3000 of compute credits for the entire hack. Alternative title: OpenAI unilaterally implements "Total Research Transparency" from Plan A. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/274BMCYj2BFES2FsZ/three-hackers-used-opus-5-to-hack-into-openai-s-core Linkpost URL: https://www.wsj.com/tech/ai/hackers-used-anthropics-claude-to-break-into-openai-b40ba883 --- Narrated by TYPE III AUDIO.

  • Friday · 37 min

    “Deep recurrent models are less robustly CoT-monitorable than normal CoT models in a toy setting” by Nick Kuhn, Alek Westover

    We use RL to teach a deep recurrent model and a normal CoT model to solve a math problem while hiding from a CoT monitor which of two possible problems it is solving. The deep recurrent model moves its reasoning into latents (thereby evading the CoT monitor) within 40 steps of RL; the CoT control is much worse at learning to confuse the LLM monitor. Note: in the above graph, “monitorability” is how much better than chance the CoT monitor is at figuring out which problem the model is solving. This was our simplest setting; similar settings show similar results with somewhat smaller differences between the CoT model and the deep recurrent model, see Results. Some proposed architectural changes add a latent channel in addition to the normal text channel for persisting state between forward passes. These parallel latents architectures (like T2MLR and full-bandwidth transformers) could theoretically perform all their reasoning in latents, rather than text. We’ve speculated that these architectures will therefore be harder to oversee. This post empirically investigates whether these models can learn to reason without their CoT (which would likely make them harder to oversee) in a simplified setting. We compare the performance of a [...] --- Outline: (03:49) A formal description of the parallel latents architecture (05:00) Experiment 1 (05:50) Tasks (09:09) Stage 1: SFT (11:53) Stage 2: RL (13:06) Results (20:50) Discussion (21:55) Task (22:50) Pressure to solve the task (23:32) Oversight method (24:25) Model (25:53) Obfuscation pressure (28:04) Suggestions for future work (29:58) Conclusion (30:26) Acknowledgements (30:49) Appendix (30:52) Experiment 2: Output Spillover (32:22) Result (33:25) Experiment 3: Training pressure variations (37:09) en-US-AvaMultilingualNeural__ When starting with low task performance, rejection filtering on highly monitorable trajectories can lead to increased obfuscation. (37:20) en-US-AvaMultilingualNeural__ When initial performance on the task is high, rejection filtering does not exert significant pressure on monitorability. The original text contained 3 footnotes which were omitted from this narration. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/5guQJSqstkjgys3PE/deep-recurrent-models-are-less-robustly-cot-monitorable-than --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Friday · 8 min

    “Superintelligence this Christmas” by Alexander Gietelink Oldenziel

    I now consider it plausible that some form of recursive self-improvement is imminent, and that we may be on track for superintelligence by Christmas of this year if racing continues. This is substantially faster than any forecast, including ones like AI 2027 that were considered outrageously fast a year ago. It is faster than I myself expected even a week ago. I don't work at a scaling lab. I don't know more than is public knowledge. Let me be perfectly clear: what I am saying is absolutely nuts. Extraordinary claims require extraordinary evidence. I claim we have now received said evidence and you should update accordingly. FOOM should probably should be your *default expectation*. People have strong status quo bias. Your default expectation should be that things will radically speed up. We are not at the ceiling of intelligence. We should probably expect the transition to superintelligence to be incredibly fast. RSI is a positive feedback loop, so it is inherently (hyper)exponential. Everything is an S-curve eventually, but nothing suggests the ceiling is anywhere near human level, or that it happens at a human timescale. AI is [...] --- Outline: (01:08) FOOM should probably should be your *default expectation*. (01:44) AI is capable of revolutionary advances in mathematics. Machine learning research is not different in kind. (03:16) The speed of AI progress continues to be underestimated; by superforecasters and even by the researchers themselves. (05:07) Internal models are significantly ahead of released ones; (06:08) Intuitions about timing from pre-training runs are misleading since most progress comes from RL, unhobbling and algorithmic innovations (06:29) Enter the Swarm (07:04) Anthropic's own report states it has 30,000 agents running concurrently, and Claude has completely taken over 26% of all R&D. The original text contained 3 footnotes which were omitted from this narration. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/LJbKwctaioqp2Hi4b/superintelligence-this-christmas --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • Thursday · 10 min

    “Grantmakers aren’t afraid to die” by dan.parshall

    The AI Risk grantmakers do not act like they believe in imminent existential risk from AI The idea of "revealed preferences" is one of the most useful in economics; it allows us to cut through a great deal of metaphysical angst about what someone "really" believes, and focus on what they act like they believe, which is much more useful for making predictions about their future actions. As one example, I grew up in a, shall we say, fervently-religious community, and it's often hard for nerdy Rationalist types to understand this, but: there are people who genuinely believe in Hell, and in Heaven. They genuinely believe that moving souls from one to the other is the most important thing on Earth. It's one thing to doubt the conviction of someone who lives an easy, staid, middle-class life... but for others, their choices and behaviors (e.g. years-long missionary trips) reveal their true preference and/or belief beyond any reasonable doubt. I bring this up because, per the actions and decisions of grantmakers operating in the AI Risk space, they mostly DO NOT seem to believe in imminent existential risk of AI. On the contrary, they act like people who [...] --- Outline: (00:10) The AI Risk grantmakers do not act like they believe in imminent existential risk from AI (01:27) The explore-exploit tradeoff (02:44) The evidence we're in "exploit" mode (02:49) Exhibit A (03:09) Exhibit B (03:32) Exhibit C (04:32) Obvious verdict is obvious (06:35) Explore mode: Just do (good) things (better) (09:35) Conclusion The original text contained 13 footnotes which were omitted from this narration. --- First published: September 17th, 2026 Source: https://www.lesswrong.com/posts/2o8B9hDN94k4Qine6/grantmakers-aren-t-afraid-to-die --- Narrated by TYPE III AUDIO.

Showing 1–20 of 261 episodes