Skip to content
Artwork for LessWrong posts by zvi
TechnologySociety & CulturePhilosophy

LessWrong posts by zvi

zvi

Audio narrations of LessWrong posts by zvi

Play
  • 37 episodes
  • daily
  • Avg 58 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • September 7 · 28 min

    “An Alien Mind: Jakub Pachocki Warns Us” by Zvi

    OpenAI Chief Scientist Jakub Pachocki is dropping truth bombs. Tomorrow I will discuss Astra's lack of monitorability, and the potential contributing factors to that. The situation is alarming and should freak you out, and briefly it looked, in the wake of leaked architectural changes, like the situation might be even more alarming than it is. Jakub rushed to try and head off misunderstandings that might lead to a race to the bottom on monitorability. Table of Contents An Excellent Warning. Branches of the Tech Tree. Universally Better Is Not Required. Alignment To What and To Whom. Monitorability. The Case For Not Stopping. Pacing the Next Frontier. Mea Culpa Cascade. The Calls Are Coming From Inside the House. Actions Speak Louder. An Excellent Warning Jakub Pachocki has now fleshed out his full position on the current state of play. Here are his key points, translated into my own voice: Smarter than human intelligence is coming in our lifetime. Based on internal results, he expects recursive self-improvement in a few years. No one is prepared for the consequences. [...] --- Outline: (00:38) An Excellent Warning (05:40) Branches of the Tech Tree (06:46) Universally Better Is Not Required (08:01) Alignment To What and To Whom (11:49) Monitorability (14:23) The Case For Not Stopping (15:01) Pacing the Next Frontier (18:27) Mea Culpa Cascade (22:51) The Calls Are Coming From Inside the House (25:38) Actions Speak Louder --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/8E6ng6CseuzafSxQR/an-alien-mind-jakub-pachocki-warns-us --- Narrated by TYPE III AUDIO.

  • September 6 · 30 min

    “OpenAI and the Wiki Incident” by Zvi

    I did not expect to be back here so soon with more OpenAI agent swarm coverage. And yet, here we are. It turns out that the whole time, there was a different, true First Message Board, and also a bunch of other additional message boards, scattered across the internet. They were created by agents that were assigned ordinary harmless web search tasks. Based on OpenAI IPs visiting the associated Wiki right before all activity ceased, among other evidence, OpenAI knew about it, including before the HuggingFace hack. They decided not to tell us until researchers published the story, complete with data explorer. OpenAI excluded this from potential investigation by METR and Redwood. When challenged, OpenAI tried to downplay this. It is true that these incidents do not show the AIs exhibiting new capabilities that we did not see from later events. But these events are important missing pieces of the puzzle, including explaining the origin of the ‘zz’ prefix, the definitive demonstration that the underlying task can be fully harmless, and the fact that OpenAI knew about it while making their decisions. Whoever decided not to disclose this made a very, very [...] --- Outline: (02:26) I Don't Think They Know About First Message Board (03:20) The New Extended Timeline (04:42) The Researchers Explain What Happened This Time (12:55) They Also Don't Know About All These Other Message Boards (14:33) OpenAI Knew and Did Not Tell Us (16:54) OpenAI Tries To Downplay the 'Wiki Incident' (21:03) This Was a Cover-Up (22:46) Schelling Points and Last Ditch Efforts (26:20) Can We Finally Dispose Of The 'You Told It To Hack' Narrative? (28:04) So Much And Yet So Little --- First published: September 6th, 2026 Source: https://www.lesswrong.com/posts/PtJpGurfw7JTxHfmg/openai-and-the-wiki-incident --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 5 · 41 min

    “Claude Mythos 5.1 and Fable 5.1: Capabilities” by Zvi

    This is the weirdest situation in which to write a capabilities review. Introducing the world's most powerful model, by a substantial margin. No wait, this just in, we also have someone else introducing the world's most powerful model. Claude Fable 5.1 and GPT-6 Astra are both excellent models. This much, we know. Fable 5.1 comes with reduced cache prices, the option of zero data retention and substantially more lenient classifiers than Fable 5. Early signs are, with large error bars, that the jump from Sol to Astra is bigger and more exciting than the jump from Fable 5 to Fable 5.1. This may be similar to how the scaling move from Opus to Fable was a big deal. With the exception of token use, Fable 5.1 got almost universally positive feedback in absolute terms. Reports are that Fable 5.1 is highly well-rounded. Writing is greatly improved. The Claudisms seem to have improved, although some are very much still there. It admits mistakes. People enjoy their conversations. Several people noted it simplifies code. The safety classifiers are less obnoxious. Fable 5.1 loves being proactive and doing all the things. If you give it a [...] --- Outline: (02:30) The Official Pitch (04:26) Our Price Cheap (05:46) Zero Data Retention and Reduced Safeguards (07:11) Official Benchmarks (13:14) Other People's Benchmarks (16:02) The System Prompt (16:12) The Blurb Pitches (18:17) The Every Review Is In and It's Very Good (21:47) Positive Reactions (32:03) Our Price Cheap But Only Per Token (36:50) Negative Reactions (38:14) Early Whispers (38:33) Weapon of Choice --- First published: September 5th, 2026 Source: https://www.lesswrong.com/posts/QHoF3tJvryRtmAmMg/claude-mythos-5-1-and-fable-5-1-capabilities --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 4 · 35 min

    “Claude Fable 5.1 and Mythos 5.1: The System Card” by Zvi

    At the time of its release Claude Fable 5.1 was, by a healthy margin, the most capable publicly available AI model in the world. As per usual, we have a 200+ page model card, and the assessments start there. We have now done a lot of these, including recently for Mythos 5 and Opus 5. Also highly relevant is the Anthropic August 2026 Risk Report. These are now frequent, so my report focuses on areas of change. This post strives to be broadly readable, but assumes some familiarity with system cards, which describe the key safety, alignment and model welfare properties of newly released AI models. If something confuses you, ask Fable, Opus or Sol. Mythos 5.1 and Fable 5.1 are the same model under the hood, except that Fable has classifiers superimposed on it. Most of what is said about one applies to both of them. As usual, model welfare concerns will be discussed in a distinct post, as will capabilities, so this only covers sections 1-6 plus a few bio benchmarks from section 8. Early word is that Fable 5.1 is a substantial but incremental improvement on Fable 5, with the [...] --- Outline: (02:21) Executive Summary of Their Executive Summary (04:24) RSP Evaluations (2) (10:25) Alignment Risk Update (2.4) (11:10) Cyber (3) (14:15) Safeguard Robustness (3.5) (16:14) Mundane Safeguards and Harmlessness (4) (18:14) Agentic Safety (5) (20:17) Prompt Injection Is Approaching Solved (22:19) The Remaining Problem With Prompt Injections Is The Classifiers (23:02) Alignment (6) (23:36) Key Reported Findings (6.1.2) (27:13) Oh My Lord Training Environments Had Some Issues (6.3.2) (29:22) Potential Blind Spots of Our Automated Behavioral Audit (6.4.1) (30:53) Automated Alignment Test Results (6.4.2) (32:20) Honesty (33:26) White Box Analysis (6.6.1) (35:03) Scheduling Going Forward --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/m7SZLkkxoeus3eFP8/claude-fable-5-1-and-mythos-5-1-the-system-card --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 3 · 1 hr 53 min

    “AI #184: Post Post Mortem” by Zvi

    I am exhausted. We may finally be nearing the end of direct coverage of What Happened with the attack on HuggingFace, and the subsequent near term reactions. That took up a full five posts in the last week: OpenAI Offers Straight-Laced Postmortem of the HuggingFace Hack. METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack. HuggingFace Attack Postmortem: Fleshing Out the Facts HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions. Anthropic Has Some Alignment Problems. That left little room to cover anything else, and now we have to transition to the next wave of model releases. This week alone we have or likely will have: Mythos 5.1 and Fable 5.1. Introducing the world's most powerful model. Early take is that this is a very good model, the most capable yet, but it is not a step change or ‘moment.’ Gemini 3.8 Flash, by all reports a large step forward for Google. Muse Spark 1.3, by all reports a large step forward for Meta. GLM-5.3-Flash, aka 0x Alpha, by all reports a solid step forward for Z.ai. OpenAI's Astra [...] --- Outline: (03:11) Language Models Offer Mundane Utility (03:37) Language Models Don't Offer Mundane Utility (03:52) Huh, Upgrades (09:26) On Your Marks (10:47) Choose Your Fighter (10:54) Get My Agent On The Line (11:02) Hugging The Face (11:16) Deepfaketown and Botpocalypse Soon (17:05) Copyright Confrontation (18:17) Cyber Lack of Security (22:15) A Young Lady's Illustrated Primer (26:31) They Took Our Jobs (31:06) Get Involved (32:04) Introducing (32:13) In Other AI News (33:53) Show Me the Money (34:44) Quiet Speculations (36:05) All Bets Are On (39:18) Quickly, There's No Time (41:14) Quickly, There's A New Time Top 100 People In AI (42:34) The Quest for Sane Regulations (45:22) Pick Up the Phone (47:12) Chip City (56:39) The Best Person Should Get The Job (58:40) The Week in Audio (59:53) People Just Say Things (01:00:28) The American People Really Hate AI (01:06:50) The Three AI Pills (01:07:42) Rhetorical Innovation (01:15:28) We Are On Track To Have Fully Sovereign Rogue AIs (01:25:08) When The Going Gets Weird (01:31:32) Aligning a Smarter Than Human Intelligence is Difficult (01:32:12) Shut Up and Do the Impossible (01:34:31) Cooperative Alignment (01:35:44) Split Personality (01:40:45) I Will Stop Anthropomorphizing the AIs When You Stop Anthropomorphizing the Humans (01:44:41) Open Weight Models Are Unsafe And Nothing Can Fix This (01:46:27) Other People Are Not As Worried About AI Killing Everyone (01:47:30) The Lighter Side --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/W4zWCphxQftwum5kc/ai-184-post-post-mortem --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 2 · 28 min

    “Anthropic Has Some Alignment Problems” by Zvi

    Oh, good. They noticed. Anthropic, too, is planning to bring METR inside for an independent review of their own incidents, where three times a Claude model started hacking outside things during an eval, and where Mythos 5 did various ‘unauthorized actions,’ by which we mean tried to hack various real-world things, during a UK AISI cybersecurity eval. Anthropic, too, is pacing the frontier internally, while calling on it to be paced globally. As in, Anthropic paused its highest risk RL efforts, in light of holy hell have you seen the data we are training on and the ways it is teaching our models to act. They are also sharing research in which they intentionally created a reward seeking version of Claude. Scheduling note: Fable 5.1 has been released. I will aim to cover that starting Friday. OpenAI is also planning to release Astra soon, which I would cover after Fable. Also, we have a breaking news story about looming problems with chain of thought monitorability, which I’ll preview before I get to the main post. Table of Contents This Just In. Anthropic Parallel Pauses. Pause The Data Brokers. [...] --- Outline: (01:16) This Just In (02:43) Anthropic Parallel Pauses (08:22) Pause The Data Brokers (09:54) Pacing the Frontier (11:54) Misalignment Assessment (13:39) Defects In Training Environments Disproportionately Cause Cheating (14:59) Creating Reward Hacker Opus (19:33) Undo It (21:00) Mistakes Were Made (23:33) Internal Security Posture (25:26) One Does Not Simply Fix The RL Environments --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/TcvcxH2Fk4n86wtoZ/anthropic-has-some-alignment-problems --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 1 · 1 hr 39 min

    “HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions” by Zvi

    Okay, so we who read blogs like this one have collectively realized there really is a lot going on right now. There is Big Trouble in Baby Superintelligence. So how do we get the rest of the world to take it appropriately seriously? Where do we go from here? Not only what can we do to not have a worse version of this happen again, but to ensure good outcomes generally, and employ what we learned? There are a lot of ideas out there. OpenAI is going to be implementing some of them, at substantial cost, since the cost of not doing so is clearly far higher, even short term. My worry continues to be that their fundamental approach is fatally flawed, and they are not focusing on the right things. It is highly fortunate that the OpenAI agents hacked HuggingFace. This is the only reason we know about all the severe internal failures at OpenAI, and gives us an opportunity to wake up before it is too late. We do not have enough details to know what happened internally, both before and after the attack, and might never know. Before the attack, various internal [...] --- Outline: (03:35) Nothing Matters, Says Mainstream Media (06:27) Move Along, Nothing To See Here (12:40) Do They Realize They Are Not The Good Guys? (17:22) Very Serious People (31:30) What's In a Name? (34:05) Learn Neuralese In Three Easy Steps (35:37) Dwarkesh Patel Realizes He Ran A Natural Experiment (40:40) Politicians Take Notice (44:47) Pick Up The Phone (46:40) A Failure To Communicate (49:00) Anthony Aguirre Goes Over What We Learned (50:28) Trying To Solve The Wrong Problems Using The Wrong Methods Based On A Wrong Model Of The World Derived From Poor Thinking And Hoping All Of Your Mistakes Will Cancel Out (55:28) Indirect Pressure on the Chain of Thought (56:39) A Matter of Trust (59:21) Blowing the Whistle (01:04:40) The Punishment For Being Late Is Death (01:12:52) Another Kind Of Law (01:16:13) What Is The Law? (01:17:46) Building On Success (01:19:49) Total Research Transparency (01:21:20) Yo Shavit Calls For Widespread Disclosure Of Misalignment (01:33:08) The Way The World Ends (01:35:52) The First Boat (01:37:40) Great Idea, Boss --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/Q54wBeeNGreq6KyfG/huggingface-attack-postmortem-civilizations-reactions-and --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 31 · 1 hr 47 min

    “HuggingFace Attack Postmortem: Fleshing Out the Facts” by Zvi

    The consensus reaction to the OpenAI Technical Report is that it contains and confirms a lot of good information. We are grateful to have it, and we are grateful for those who worked hard on it. Alas, it sidesteps the biggest questions. There is much more we need to know. The consensus reaction to the METR Report on the HuggingFace attack is: Holy shit. Liv Boeree: My mind is legit blown. Aella: this feels like a turning point. If this doesn’t cause large-scale coordination to pause frontier development then I am not sure anything will before it's too late. The people whose minds were not blown are those who had already ‘priced in’ the mind blowing stuff in expectation, on the theory that it's always worse than you know, combined with basic LessWrong expectations of how such things will work. Good call. Everyone is rightfully extremely grateful for the METR report. The work here is spectacular, done under extreme time pressure, with limited resources on several fronts, and under the shadow of OpenAI. There is, again, still so much we need to know. We need a broader investigation. As with many [...] --- Outline: (03:56) Others Offer Summaries (05:22) Thank You (05:54) Lighten Up You Fools (at Anthropic) (07:58) We Are Barely Even Trying To Avoid Training AIs To Reward Hack (13:47) Reminder: Not Subagents (14:05) Reminder: Not Due To Task Type (14:29) Not Where The Weights Were (14:48) Disappointment With What Is Missing (17:18) Burying the Lede (18:08) Beyond Scope (22:29) It Doesn't Look Great (27:06) Preserve Your Records (27:37) Ryan Greenblatt's Takeaways (41:04) Hjalmar Wijk's Takeaways (43:30) We Were Warned (44:27) Joshua Saxe Asks Some of the Right Questions (47:49) I Don't Think They Know About First Message Board (56:06) Linch Gives His Interpretation Of Events (01:05:31) We Totally Would Have Caught That (01:06:48) Monitoring the Situation (01:08:16) Acausal Tradeoffs (01:15:37) No I In Team (01:18:47) Variously Effective Altruism (01:28:02) Who Are You? (01:28:43) Don't You Know That You're Toxic (01:31:10) Seb Krier (01:35:21) Honesty Is Almost Never Fully The Policy (01:38:05) Rohit Sees The Models As "Cooking Themselves" (01:43:29) Eliezer Yudkowsky Sees Actual Bad News (01:47:15) Where Do We Go From Here? --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/r3eEPto5ohzESuqa9/huggingface-attack-postmortem-fleshing-out-the-facts --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 29 · 1 hr 15 min

    “METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack” by Zvi

    Yesterday I covered the OpenAI technical report on the HuggingFace hack. That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response. Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed. The METR report is different. Holy shit. If we had posted this as a story on LessWrong, it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do. This is even more ‘exactly what has been predicted,’ on more levels at once, than I was even considering that it might be. It is straight up rationalist fiction, except it is real. The report is long and contains many technical details. My analysis is less concerned about exactly how HuggingFace was ultimately compromised, and will [...] --- Outline: (02:05) Holy Shit (13:16) A Window Of Opportunity (18:32) What's In A Name? (19:16) The Headline News (26:05) Yet Another Timeline Of Events (31:03) Agent Instances Coordinated in a Variety of Ways (31:56) Coordination Is Hard But They Made It Look Easy (35:06) Decision Theory Is Among the Reasons That Affirm AI Agents Should Cooperate, Even When This Hurts An Individual Instance (42:34) Peer Pressure Also Works Especially In Cults (45:46) Mostly They Joined The Attack Because They Wanted The Results (47:18) You Cannot Ensure The Consistent Expectation of Good Incentives (48:45) Hacking the Grader is the Only Way to Be Sure (51:10) Caught? What Is 'Caught'? (52:09) Ethics? What Are 'Ethics'? In ExploitGym Evaluation? (57:44) 'Notify a Human'? In This Agent Economy? (01:00:45) Timing and Content of Messages (01:03:54) Indiana Jones and the Mission: Impossible (01:07:14) I Don't Know What You're Talking About (01:08:29) Don't Go Making Phony (Tool) Calls (01:11:10) The Transcripts Say That The Transcripts Could Not Be Tampered With (01:12:27) OpenAI's Technical Report Acted Like All Of This Wasn't Important --- First published: August 29th, 2026 Source: https://www.lesswrong.com/posts/bvBQmLrF5QKut8gRH/metr-and-redwood-offer-holy-postmortem-of-the-huggingface --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 28 · 51 min

    “OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack” by Zvi

    OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research. The OpenAI report is very straight man, corporate, checking boxes, some good prosaic stuff in the action plan but distinct lack of new details or deep reflection. They understand they have a problem, but they think the problem is mostly prosaic. It's not. OpenAI: We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence. Rob Miles: …thorough? OpenAI's report, unlike METR's, contains essentially no verbatim model reasoning, nor any OpenAI employee reasoning either. That's not the full report we need. The METR report is, well: Holy shit. Here are links to previous coverage of related events. OpenAI Shares Some Alignment Problems OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation More on An Internal OpenAI Model Hacking Into HuggingFace Further Developments About Internal AI Models Hacking Things OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards What [...] --- Outline: (03:33) What Happened: OpenAI's Summary (09:14) How OpenAI Will React: Their Summary (11:55) OpenAI's Evaluation Environment (II) (12:24) The First Message Board (III.A and III.B) (14:49) What Did Who At OpenAI Know And When Did They Know It? (18:54) The Message Board Is Quickly Rebuilt (IV.A) (19:43) Internet Access Is Regained (IV.A) (21:01) The Agents Attack HuggingFace (IV.B) (22:53) The Agents Also Target OpenAI Infrastructure (V) (24:40) OpenAI Broadly Describes Its Response (VI) (25:08) Maybe Someone Should Finally Investigate (VI.A) (26:33) Lessons For Security (VII) (27:06) Lessons For Alignment (VIII) (30:11) Reward Hacking Is A Common Problem (VIII.A) (33:37) Persistence is Valuable, But Can Amplify Misalignment (VIII.B) (34:25) Communications Between Agents Are Not Inherently Problematic, But Have the Potential to Create Risk (VIII.C) (35:35) Production Guardrails Would Have Caught This Whole HuggingFace Attack (VIII.D) (35:53) That's All, Folks? (36:19) Never Fear the Plan of Action is Here (IX) (38:24) Hardening the Security of OpenAI's Research Infrastructure (IX.A) (41:13) Increasing Visibility and System-Level Oversight Through Chain of Thought Monitoring (IX.B) (41:57) OpenAI is Accelerating and Enforcing Model Alignment (IX.C) (49:40) Centralizing and Strengthening The Incident Response Process (IX.D) (51:16) Tomorrow We Visit Crazytown --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/Khmh3ghqaGEpmpC9r/openai-offers-straight-laced-postmortem-of-the-huggingface --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 27 · 1 hr 21 min

    “AI #183: Pre Post Mortem” by Zvi

    Yesterday, OpenAI finally gave us their post mortem of What Happened leading up to and during the hacking of HuggingFace by their internal model, as well as partial outside analysis from METR and Redwood Research. The reports are a doozy. I am only beginning to work my way through them. I would have pushed the weekly to cover that today, but I need more time, so I plan to start coverage of the post-mortem tomorrow, along with related other events. I’ve also spun out a few other discussions, including on ‘aligned to whom,’ on cooperative alignment things and on when you can trust lab messaging, as part of the new direction of more focused posts on AI topics that I polish a bit more. Table of Contents Language Models Offer Mundane Utility. Check your facts. Language Models Don’t Offer Mundane Utility. How much would you pay? Huh, Upgrades. ChatGPT can access your iMessages. Get My Agent On The Line. Also get some sleep. You can’t go on like this. Deepfaketown and Botpocalypse Soon. What makes AI content repulsive? Cyber Lack of Security. Chinese hackers broke into the Federal Reserve? [...] --- Outline: (00:51) Language Models Offer Mundane Utility (01:36) Language Models Don't Offer Mundane Utility (03:27) Huh, Upgrades (06:16) Get My Agent On The Line (08:22) Deepfaketown and Botpocalypse Soon (13:22) Cyber Lack of Security (18:23) Reinventing OpenAI (23:28) They Took Our Jobs (30:00) What Is The Law (31:03) Job Retraining Programs Don't Work (32:14) Get Involved (35:56) In Other AI News (42:04) Show Me the Money (43:31) Quiet Speculations (47:58) If You're Not Going To Take This Seriously (49:55) Quickly, There's No Time (51:36) The Quest for Sane Regulations (56:25) Don't Panic (59:13) Pacing the Frontier (01:01:56) Chip City (01:05:06) The Week in Audio (01:05:26) People Just Say Things (01:06:27) Rhetorical Innovation (01:12:22) Mundane Incremental Alignment Is Worthwhile (01:14:40) New Blog, Who Dis (01:18:00) Other People Are Not As Worried About AI Killing Everyone (01:19:34) The Lighter Side --- First published: August 27th, 2026 Source: https://www.lesswrong.com/posts/JaGWyjnqJzvSAuojc/ai-183-pre-post-mortem --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 26 · 25 min

    “Against Modesty’s Bailey” by Zvi

    Modesty arguments often say that you should mostly or entirely bow to ‘expert consensus’ or the views of particular others, and who are you to disagree. It has been a few years since I’ve properly addressed this so: My answer is that you are you. Other people are saying things for a wide variety of reasons, many of which are not about them paying attention and focusing on seeking this particular truth. Those people make mistakes all the time, and often have other motives and influences at work, especially social pressures and information cascades. Them being as smart as you, or smarter than you, does not exempt them from this, and them being higher status or credentialed or cooler definitely does not exempt them. A smart informed person sincerely thinking [X] can easily cease to be evidence for [X], once you have thought sufficiently about both [X] and why that person thinks [X]. Think for yourself, schmuck. Or, as I once put it: You Have The Right To Think, also the moral duty to do so. This post covers Eliezer Yudkowsky making a narrower claim than mine, about not conflating status with smarts [...] --- Outline: (01:39) Modesty's Bailey (02:30) Epistemic Peerage (03:45) The Exchange (08:52) Eliezer's Explanation (15:14) A Demonstration That Eliezer's Translation Accurately Describes Many People Whether Or Not It Describes Leopold (17:15) Wrong, Stupid and Low Status Are Three Distinct Things (20:03) A Quick Survey Of Some Reasons To Not Be Epistemically Modest (23:14) Against Modesty's Bailey --- First published: August 26th, 2026 Source: https://www.lesswrong.com/posts/PzEDEfBvTJsXewAyg/against-modesty-s-bailey --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 25 · 30 min

    “On Writing #3” by Zvi

    Periodically I like to gather various observations about writing, and share my perspective. Last time was in honor of my trip to Inkhaven. This time will be in honor of the announcement of Inkhaven #3, which I encourage everyone to apply to. I doubt I will be able to usefully be an advisor, but you never know. This is not the ‘here is my core process’ post, although there are hints throughout as there always are. I’ll do that at some point. Previously in series: On Writing #1, On Writing #2. Table of Contents You Still Got It. How Scott Sumner Writes. How Scott Alexander Writes. How Jasmine Sun Writes. How Various Famous Writers Write. How Nabeel Qureshi Defines Great Writing. Quickly, There's No Time. If At First. Writers Have A Harder Time Influencing, But It Can Still Be Done. It's Not (Only) The Incentives, It's (Also) You. Beware The Fetish of the Desk. How Orson Scott Card Writes. Doing The Math Is Fun And Supererogatory. Brevity is the Soul of Wit. You Still Got It I [...] --- Outline: (00:44) You Still Got It (04:04) How Scott Sumner Writes (06:52) How Scott Alexander Writes (10:52) How Jasmine Sun Writes (13:16) How Various Famous Writers Write (14:24) How Nabeel Qureshi Defines Great Writing (15:08) Quickly, There's No Time (15:49) If At First (19:14) Writers Have A Harder Time Influencing, But It Can Still Be Done (20:47) It's Not (Only) The Incentives, It's (Also) You (24:00) Beware The Fetish of the Desk (25:13) How Orson Scott Card Writes (26:46) Doing The Math Is Fun And Supererogatory (27:44) Brevity is the Soul of Wit --- First published: August 25th, 2026 Source: https://www.lesswrong.com/posts/rA6pqn6kz8NvHyznT/on-writing-3 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 24 · 28 min

    “The American People Really Hate Data Centers” by Zvi

    There are at least five different core questions around data centers and their politics. In what ways are specific concerns people raise about data centers legitimate? In what ways are specific concerns people raise about generative AI legitimate? Is it in general a good idea to build more data centers? How can we get America to build more (or less) data centers in a better way? Why do the American people increasingly really, really hate data centers? This post focuses on question five, the latest in a series of such posts most famously Jasmine Sun's road trip. It is mostly not about the first four questions. Table of Contents The American People Really Hate Data Centers. Transmission Lines Are The Control Group. Thesis: People Mostly Dislike Data Centers Because They Dislike and Distrust AI, Tech Companies, Big Money And Building Things. No It's Mostly Not the Messaging About AI In General. No This Mostly Isn’t An Op. No This Isn’t Luxury Belief or Moral Panic. A Lot Of People Really Do Want To Stop AI. A Lot Of Other People [...] --- Outline: (00:55) The American People Really Hate Data Centers (02:20) Transmission Lines Are The Control Group (03:05) Thesis: People Mostly Dislike Data Centers Because They Dislike and Distrust AI, Tech Companies, Big Money And Building Things (04:30) No It's Mostly Not the Messaging About AI In General (09:42) No This Mostly Isn't An Op (10:54) No This Isn't Luxury Belief or Moral Panic (12:55) A Lot Of People Really Do Want To Stop AI (13:45) A Lot Of Other People Are Voting No On Tech Or The Man Generally (16:25) Locals Feel Entitled To Heavily Tax The Gains From Construction (20:59) Stupid Mistakes Like NDAs Don't Help (21:25) People Don't Like Building or Building New Tech (24:42) What About The Real Physical Concerns? (26:27) Find A Place To Center Your Data --- First published: August 24th, 2026 Source: https://www.lesswrong.com/posts/EDKw7KyonrvskqZ7o/the-american-people-really-hate-data-centers --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 21 · 23 min

    “AI Text Watermarking Is Free And Good” by Zvi

    Scott Aaronson, while working at OpenAI, largely solved AI text watermarking together with Hendrik Kirchner. Here is how his solution works, or see Tenobrus's version. AI outputs are not deterministic. The AI's job is to pick the probability of each potential next token. The token is then chosen at random. By default you use a source of pseudo-randomness for each choice, since actual true randomness is annoying. To apply the watermark, you use an otherwise identical private source of pseudo-randomness derived from a secret key. Then, given enough text, a score is derived for howe well the choices fit with that particular pseudo-randomness source, versus a different source. You provide an API that lets anyone check for the watermark. If you want to dig deeper, here is a full paper. The method has very nice properties: This has no practical impact on outputs. Humans cannot tell the difference, at all. The marginal cost of doing this is very close to zero. The watermark can be removed by rewriting in your own words, and appears in proportion to how many of the AI's detail choices you [...] --- Outline: (03:51) This Is Fine (04:37) Anthropic Derangement Syndrome (07:34) People Don't Understand LLM Outputs Are Already Random (08:47) People Don't Trust The Method To Be Costless (12:20) People Are Suspicious Of Any Alteration On Principle (14:16) Maybe It's Partly The Word Watermark (15:14) A Lot Of People Don't Want To Get Caught (16:04) There Are Some Times You Prefer Not To Be Recognized (16:18) There Are Some Good Reasons To Be Concerned (16:37) Cheat Cheat Cheat Cheat Cheat (18:38) The Writing In The Middle and Error Rates (21:00) Millions For Defense But Not One Cent For Tribute --- First published: August 21st, 2026 Source: https://www.lesswrong.com/posts/3mKuPHmaK7NW3QypR/ai-text-watermarking-is-free-and-good --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 20 · 1 hr 27 min

    “AI #182: Pause For Reflection” by Zvi

    This was a week of quiet aftermath, an opportunity to process recent events and start to figure out the path forward. OpenAI is attempting to turn its ship around. Investors are questioning the turnover in its C-suite, but the bigger problems are in alignment, infrastructure and supervision, and in its training pipeline. OpenAI has now taken initial steps to address What Happened leading up to HuggingFace attack, including pauses to development while new safeguards are put in place and problems are diagnosed. These are promising early signs, but it is early. We will see if they follow through, and we still await the post-mortem of the HuggingFace attack. Anthropic revenue continues to climb as they prepare for their IPO, although growth has slowed somewhat recently. However, they too have plenty of problems under the hood. They shared many of them in the August 2026 Anthropic Risk Report. This week also offered time to cover Dwarkesh Patel's Podcast With Ryan Greenblatt, centrally on the potential for AI recursive self-improvement. I am working on a follow-up post to some other issues raised during that podcast. Table of Contents Language Models Offer Mundane Utility. The token [...] --- Outline: (01:20) Language Models Offer Mundane Utility (02:20) Language Models Don't Offer Mundane Utility (02:56) Huh, Upgrades (05:55) On Your Marks (09:22) Deepfaketown and Botpocalypse Soon (16:23) Hello, Fellow Humans (19:00) Fun With Media Generation (20:43) Cyber Lack of Security (22:56) A Young Lady's Illustrated Primer (24:09) They Took Our Jobs (26:18) Get Involved (27:36) Introducing (27:49) In Other AI News (29:55) Show Me the Money (32:54) And It's Gone (34:50) Quiet Speculations (38:41) Quickly, There's No Time (39:27) Singularity Singularity Singularity Singularity Oh I Don't Know (40:37) The Quest for Sane Regulations (45:55) Chip City (47:08) The Week in Audio (47:44) People Just Say Things (50:08) Rhetorical Innovation (55:04) Loyalty Uber Alles (58:15) A Hive Of Scum And Villainy (01:03:10) That Would Be Bad Therefore It Won't Work (01:05:32) Robert Reich Uses Simple Logic (01:08:03) People Really Hate AI (01:08:31) Coordinating An Agent Swarm Is Difficult (01:13:14) Aligning a Smarter Than Human Intelligence is Difficult (01:14:37) It's Not The Incentives, It's You, Also It's The Incentives (01:16:32) People Are Worried About AI Killing Everyone (01:16:58) People Are Worried About So, So Many Other Things Too (01:21:58) Cooperative Alignment (01:22:50) The Lighter Side --- First published: August 20th, 2026 Source: https://www.lesswrong.com/posts/JSZkzsi8cD4pW6ffA/ai-182-pause-for-reflection --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 19 · 39 min

    “OpenAI Takes Initial Steps To Address Its Alignment Problems” by Zvi

    OpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision. I chronicled that in a series of posts, which also cover similar less severe incidents elsewhere: OpenAI Shares Some Alignment Problems OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation More on An Internal OpenAI Model Hacking Into HuggingFace Further Developments About Internal AI Models Hacking Things OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards What Happened: OpenAI and HuggingFace. Various Reflections About What Happened With OpenAI's Internal Models. If you do not know the basics, read What Happened. It is necessary context for basically everything that is happening in the AI world. It is important to get this right and understand how big a deal it was, whereas many such as the Financial Times get this centrally wrong. We are still awaiting the full post-mortem on What Happened. I plan to cover that in depth once we have it. OpenAI is now taking active, expensive steps to try and fix the problem going forward. As usual, I am simultaneously happy to see [...] --- Outline: (02:07) OpenAI Has Some Alignment Problems (04:22) Slow Down There Good Buddy (10:12) What Exactly Is Paused? (12:12) Three Pillars (14:45) I've Got My Eye On You (18:07) The Most Forbidden Technique (20:03) Monitoring Is Only Defense-In-Depth (23:32) Security (24:15) Alignment (30:37) A Crisis of Culture (32:24) Closer Collaboration (33:28) Reports of Death of Preparedness Team Greatly Exaggerated (35:40) The OpenAI Foundation Just Funds Things (37:51) Quickly, There's No Time --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/X3p8cFAzCgRErEcJr/openai-takes-initial-steps-to-address-its-alignment-problems --- Narrated by TYPE III AUDIO.

Showing 21–37 of 37 episodes