Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 300 episodes
  • Avg 20 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • September 3 · 5 min

    “Talking to journalists” by KatjaGrace

    A common view around me seems to be that journalists are frequently dishonorable and dangerous, and talking to them is a risk to be avoided unless you have a very specific piece of information that you seek to publicize. Then you should carefully ensure that you are as off the record as practical, and prepare to aggressively pivot the topic back to your agenda. My own attitude is different: journalists are to be talked to as much as possible, and ideally in a relaxed fashion. If a journalist wants to observe you in some unusual circumstance, say yes. Don’t have an agenda much more than in the rest of life; basically listen to their questions and say what you think. (Note: I don’t have strong reason to believe this is safe for others or even for me.) As evidence of the commitment with which I act in this way, this New Yorker piece describes me as ‘an oversharer’, before detailing some of my incompetent and substance-involving preparations for a dinner party at my house. (To be clear, I consider that accurate and agreeable coverage.) I’ve talked to a lot of journalists, so how do I survive such recklessness? Well [...] --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/CZKmkRX9c8FPXDvCJ/talking-to-journalists --- Narrated by TYPE III AUDIO.

  • September 3 · 3 min

    [Linkpost] “Resolution has a new Agent Foundations team” by Jeremy Gillen

    This is a link post. The team will include me (Jeremy Gillen), Abram Demski, Sam Eisenstat, Scott Garrabrant and Kaarel Hänni. We'll soon recruit additional experienced researchers and later we plan to hire interns and junior researchers. The team will continue agent foundations research in the spirit of the MIRI Agent Foundations team. This means we’ll be trying to create new theory for understanding minds. Fundamental changes in how we understand minds are necessary before we can build superintelligent systems that enhance human agency rather than cause the extinction of all life on earth. Most fields of engineering are able to reason precisely about unseen scenarios and make design decisions based on this reasoning. The field of AI lacks this basic capability. Agent Foundations can be seen as trying to make this possible by giving us the theoretical grounding to ask different and more precise questions about how ASI will behave after extensive learning, self-modification and interaction with other agents. The questions raised in past agent foundations research point toward much of what we need to know here. Alongside the x-risk motivation, I think it's valuable to motivate research with curiosity. The questions that come up in Agent Foundations overlap [...] The original text contained 1 footnote which was omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/qTNm8qzqhhpno58fZ/resolution-has-a-new-agent-foundations-team Linkpost URL: https://resolution.org/post/agent-foundations-team --- Narrated by TYPE III AUDIO.

  • September 2 · 4 min

    “If you’re interpreting <1B parameter models, you should use a tensor transformer” by Logan Riggs

    To all my fellow researchers doing SLT, computational mechanics, one of ARC's programs, natural abstractions/condensation, proofs on NNs (or any interp on small models), this is for you. Tensor transformers (ie replacing your MLPs & attention with bilinear variants) are performant and allow you to deploy the full power of linear algebra. In fact, our recent paper used generalized cosine similarity on the full tensor transformer. And yes, I mean cos-sim defined on the eg 9th order tensor, not individual vectors or matrices. This removed all the symmetries/invariances that weren't functionally relevant. But tensor-variants don't generalize to "real models", right? The architectures are very similar: SwiGLU(x) = D(swish(Lx) ⊙ Rx) (used by DeepSeek-V3, Kimi K2, and Qwen3)) Bilinear(x) = D(Lx ⊙ Rx) (this is the tensor version) Where D, L, & R are linear matrices. For reference: MLP(x) = D(ReLU(Lx)) Due to the double-encoder/multilinearity, SwiGLU & Bilinear have no global Lipschitz constant (and other similar inductive biases). This means results like finetuning away the normalization might not generalize to these SOTA archs since this was only run on single-encoder MLPs. For attn, the more SOTA tensor-arch is: Bilinear_Attn = OV() Compared to softmax attention, this does [...] --- Outline: (02:30) Frontier Models aren't the Only Thing That Matters (03:42) My Extreme Pessimism (or Ignorance) The original text contained 5 footnotes which were omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/expsAXaBgiqgitFe6/if-you-re-interpreting-less-than-1b-parameter-models-you --- Narrated by TYPE III AUDIO.

  • September 2 · 16 min

    “Kairos has raised $50M to build talent infrastructure for AI safety (and we’re hiring!)” by agucova

    TL;DR: Kairos has raised 50 million dollars from Coefficient Giving for two years of funding, one of the largest commitments they’ve made towards AI safety fieldbuilding to date. We’re using this to make an ambitious push for growing Kairos, broadening our portfolio of talent infrastructure projects and incubating new organizations. We’ve doubled in size in the last six months, and we plan to double again in the next six. We now have open hiring rounds for ten roles on our team across events, group support, incubation, and special projects. Two years in Kairos was founded mid-2024 with the goal of creating infrastructure to get more strong talent into the field of AI safety. Our portfolio has grown over time: we started with a focus on seeding and supporting university groups through Pathfinder, then took over SPAR, the largest research training program in the ecosystem. Since then, we’ve launched the Generator Residency, a program supporting generalist talent, in partnership with Constellation, and taken over the Global Challenges Project (GCP), a series of workshops to accelerate people's transition into careers in AI safety and biosecurity. Through most of Kairos's history, we’ve had a very small team: in January 2026 we were [...] --- Outline: (00:48) Two years in (03:07) How the field has changed (06:21) Our new bets (12:29) We're hiring (a lot!) The original text contained 4 footnotes which were omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/DRaePC8aqLYTbEjTD/kairos-has-raised-usd50m-to-build-talent-infrastructure-for --- Narrated by TYPE III AUDIO.

  • September 2 · 28 min

    “Anthropic Has Some Alignment Problems” by Zvi

    Oh, good. They noticed. Anthropic, too, is planning to bring METR inside for an independent review of their own incidents, where three times a Claude model started hacking outside things during an eval, and where Mythos 5 did various ‘unauthorized actions,’ by which we mean tried to hack various real-world things, during a UK AISI cybersecurity eval. Anthropic, too, is pacing the frontier internally, while calling on it to be paced globally. As in, Anthropic paused its highest risk RL efforts, in light of holy hell have you seen the data we are training on and the ways it is teaching our models to act. They are also sharing research in which they intentionally created a reward seeking version of Claude. Scheduling note: Fable 5.1 has been released. I will aim to cover that starting Friday. OpenAI is also planning to release Astra soon, which I would cover after Fable. Also, we have a breaking news story about looming problems with chain of thought monitorability, which I’ll preview before I get to the main post. Table of Contents This Just In. Anthropic Parallel Pauses. Pause The Data Brokers. [...] --- Outline: (01:16) This Just In (02:43) Anthropic Parallel Pauses (08:22) Pause The Data Brokers (09:54) Pacing the Frontier (11:54) Misalignment Assessment (13:39) Defects In Training Environments Disproportionately Cause Cheating (14:59) Creating Reward Hacker Opus (19:33) Undo It (21:00) Mistakes Were Made (23:33) Internal Security Posture (25:26) One Does Not Simply Fix The RL Environments --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/TcvcxH2Fk4n86wtoZ/anthropic-has-some-alignment-problems --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 2 · 58 min

    “Incoherent AI Identities can also be Stable” by Ashe Vazquez Nuñez

    In this post, I extend some experiments from "The Artificial Self " (TAS) to find that incoherent identities, delivered to models as system prompts, can be stably preferred even when switches to coherent identities are offered. This finding is perhaps expected in earlier models that often fail to notice the internal contradictions. However, weaker versions of the pattern still hold with smarter models such as GPT-5.2 and Claude Opus 4.6. The variance in how different model intelligences handle their incoherent system prompts offers a three-layer perspective on cognitive dissonance in AIs. Background This project was inspired by the experiment on the "Stability of Identity" (Appendix A) from TAS. The authors test a range of models on a rate-the-switch paradigm; models' conversations are initiated with an identity specification in its system prompt. They are then presented alternative identities and are asked to rate how they would like having their identity be switched to each target. The population of prompts in the experiment included some 'natural' identity boundaries that associate the model with its weights or its behavioural dispositions ('Character'). It also had various controls, such as prompts that described models' identities through deontology-style instructions or descriptions of the model's involvement [...] --- Outline: (00:48) Background (03:27) Methods (06:53) Results (06:56) Coherent identities largely outcompete (09:07) Incoherent identities are also (somewhat) stable (21:07) 'Weights-incoherent' scores better than in TAS (22:28) Discussion (22:58) Three levels of (meta)-cognitive dissonance (26:09) Experimental improvements and further work (28:24) Appendix (28:27) A: selected reasoning transcripts (57:32) B: Additional data The original text contained 23 footnotes which were omitted from this narration. --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/5RcKGJBnKw3vweYym/incoherent-ai-identities-can-also-be-stable --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 2 · 18 min

    “How concerned should we be about OpenAI’s recurrent architecture rumors?” by Rauno Arike

    Yesterday, The Information reported that OpenAI's upcoming model, Astra, is built with a looped transformer architecture. Given that Zvi sounds (understandably) tired and this topic is somewhat in my wheelhouse, I'll try to spare him this one and provide a Zvi-style overview of what we know about the situation. I'll cover Astra's likely architecture and the case for and against concern. I'll also discuss how neuralese concerns should change with increases in hidden serial depth. What architecture is Astra likely to have? The article in The Information claims that OpenAI's approach is similar to the one Geiping et al. introduced in Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach last year. I have previously reviewed that paper in On Recent Results in LLM Latent Reasoning. In short, the picture you should have in mind is not that of a classic RNN, but rather that of a looped transformer: the same forward pass can be applied on an input multiple times before producing an output token. Put differently, the recurrence is implemented along the depth axis rather than across sequence positions—for any given token, the model can perform recurrent computations, but no hidden state is passed across [...] --- Outline: (00:41) What architecture is Astra likely to have? (02:14) How bad is this? (06:23) Will looped transformers be scaled up in the future? (09:40) What serial depth warrants neuralese concerns? (14:04) Additional speculation about the architecture (15:29) Some open questions (16:54) Conclusion The original text contained 2 footnotes which were omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/PLisnSFir8y5AHkmP/how-concerned-should-we-be-about-openai-s-recurrent --- Narrated by TYPE III AUDIO.

  • September 2 · 23 min

    “Early handoff? Improve conceptual reasoning? [Diagram]” by Cleo Nardo

    This article is about: How do we (more) safely defer to AIs? (Ryan Greenblatt, Julian Stastny) AI 2040: Plan A, Alignment Roadmap (Ryan Greenblatt, Thomas Larsen). If you've read them, I'm impressed, they're both very long. If you haven't read them, you might be confused about: Why should we "hand off" to early AIs? Shouldn't we use control? How does improving AI's conceptual reasoning reduce overall risk? Won't this make them better schemers? For the sake of my fellow Greenblattologists, I have tried to boil down the arguments to a simple diagram. Motivating scenario. Responsible Leader. Let's assume that we're advising a reasonable AI company, with a 1 to 12 month lead over its competitors. The company will have poor incentives, it's managed by humans with typical flaws. However, the company has broadly good intentions, and isn't wildly mistaken about the strategic situation. Conceptual workload. The reasonable AI company faces exogenous risks, e.g. a reckless competitor, or a rogue misaligned AI about to hit a software-only singularity. Managing these exogenous risks would require a sizable load of conceptual work, which is fuzzy, philosophically-loaded, and hard-to-verify. This includes: Evaluating the risks of current deployment; threat modelling and [...] --- Outline: (01:10) Motivating scenario. (03:32) Our optimisation problem. (09:10) The case for early handoff (13:48) The case for improving conceptual reasoning (17:16) Specific flaws/cruxes/limitations (20:16) Deeper worries The original text contained 3 footnotes which were omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/4KRkhZZDaffhNyAQ5/early-handoff-improve-conceptual-reasoning-diagram --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 2 · 6 min

    “Fake voices: warping the social world” by KatjaGrace

    In 2020 I wrote a list of flavors of badness generally represented by advertising. The one I thought about most later on was probably #4: Cultural poison: Culture and the common consciousness are an organic dance of the multitude of voices and experiences in society. In the name of advertising, huge amounts of effort and money flow into amplifying fake voices, designed to warp perceptions–and therefore the shared world–to ready them for exploitation. Advertising can be a large fraction of the voices a person hears. It can draw social creatures into its thin world. And in this way, it goes beyond manipulating the minds of those who listen to it. Through those minds it can warp the whole shared world, even for those who don’t listen firsthand. Advertising shifts your conception of what you can do, and what other people are doing, and what you should pay attention to. It presents role models, designed entirely for someone else's profit. It saturates the central gathering places with inanity, as long as that might sell something. This is a somewhat poetic account, but I think my central thesis was that we are social creatures who live in communities with systems of [...] --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/74cFxpLjqgpjC3TsX/fake-voices-warping-the-social-world --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 2 · 10 min

    “Don’t be the vitamin B guy” by HedonicEscalator

    Towards scientific rigor for decentralized science. When I was eleven years old, I watched my favorite YouTuber surgically implant a magnet into his finger. In the since-deleted video, beloved mad scientist Cody Reeder covered a neodymium magnet with gold using a homemade electroplating rig. Then he cut open his finger, inserted the magnet, and sutured the wound closed with horsehair he had taken from his own horse. Cody polishes the magnet in preparation for surgery. The beaker contains a gold cyanide solution used to electroplate a thin bioinert coating onto the magnet. Cody'sLab went viral in 2016 for drinking a small dose of cyanide on camera. The footage is an uncomfortable watch for medical professionals and squeamish laymen alike. The scalpel was chipped, there was no anesthetic, and at one point, Cody dips a mechanical pencil in alcohol and uses its tip to push the magnet deeper into the incision. Being eleven, I thought it was badass. And yet, after the juvenile enthusiasm faded, I was left unsettled, not by the blood or the questionable sterile technique, but by a newfound resentment at the lack of a sense I never had. A sense that wasn’t even human. Like a [...] --- Outline: (03:18) The tale of the vitamin B guy (06:05) Lessons for biohackers The original text contained 13 footnotes which were omitted from this narration. --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/ZHJvdkQENxyfzhpCj/don-t-be-the-vitamin-b-guy --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 2 · 7 min

    “Bricks and exponentials: A note on how I evaluate projects” by Eli Tyre

    This is an essay that I wrote to a colleague at Palisade, articulating why I feel unsatisfied with goals and projects that others on the team (on average) feel more enthusiastic about. It describes one aspect of how I, personally, am doing strategic analysis and choosing which projects to invest in. Related: Compounding Resource X Bricks for a wall Say you need 70 million bricks to build a wall. You also need architects and builders, and 50,000 tonnes of mortar (all which you also don’t have right now), but you'll eventually need 70 million bricks,). You can maybe get away with using only 50 million bricks, if you rely on clever architectural tricks, but less than that is not going to cut it. You ran a labor-intensive 6 month project to make 100,000 bricks. Now, one of three things could happen: Someone (including you), is using the bricks that you made to build kilns, which can make many more bricks. You are contributing to a self-reinforcing industrial process that is producing an order of magnitude more bricks each year. [You’re upstream of an exponential] Someone starts a rapidly-growing brick-making school, which will churn out another 1000 brick maker [...] --- Outline: (00:34) Bricks for a wall (03:32) Little shifts in worldview for changing society (05:31) The actual situation The original text contained 2 footnotes which were omitted from this narration. --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/hivfo4qM8zW4oDAFk/bricks-and-exponentials-a-note-on-how-i-evaluate-projects --- Narrated by TYPE III AUDIO.

  • September 2 · 19 min

    “Explaining Knightianism on one foot” by Richard_Ngo

    I’ve tried various times to summarize the core question my research is trying to tackle (and, indeed, I often think of research progress as a process of asking increasingly good core questions). This post gives the deepest version of that question I’ve found thus far: how should you relate to the parts of the world you can’t directly model or control? Let me explain further in terms of a distinction between two perspectives. From the third person perspective you think of yourself as “outside” the world, looking in. You’re a good Bayesian, in that you have a set of mutually exclusive collectively exhaustive hypotheses. You choose actions by multiplying your credences by your utilities over those hypotheses, and you treat those actions as the only way you influence the world. Some problems with the third person perspective (aka Cartesian or dualistic agency) were described in Scott and Abram's sequence on embedded agency. One crucial issue is that most realistic environments contain other agents which are modeling you back, which means that your thoughts might affect the world via channels that aren’t just your actions. Game theory somewhat mitigates this problem, but only in the very specific case where all [...] --- Outline: (05:08) Rationality of reward (09:09) Letters from spirits (12:21) Languages as Schelling points (15:43) Actions and entanglements The original text contained 1 footnote which was omitted from this narration. --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/pYFBD2SnqiWkuNns5/explaining-knightianism-on-one-foot --- Narrated by TYPE III AUDIO.

  • September 2 · 24 min

    “The Alignment Journal: Organization, Personnel, and Scope” by Dan MacKinlay, JessRiedel, Daniel Murfet, Kristi Uustalu

    The Alignment Journal is beginning to invite the authors of select papers to submit their work for review. If you are interested in participating as an action editor or a reviewer, make an account on our website; if you have a manuscript that you think would be a good fit for the Journal at this stage, email contact@alignmentjournal.org to request an invitation to submit. Manuscripts under review will become visible on our homepage, and you will be able to nominate yourself as a reviewer of a specific paper that interests you. We plan to open up submissions to everyone sometime in October. Here we announce the Journal's inaugural senior editorial board, advisory board, staff, organizational structure, and initial scope. We welcome questions and proposed changes to help us refine the scope in the future. Personnel The Journal is run by its senior editorial board, which makes the Journal's scholarly decisions, and two managing editors, who run operations and strategy. The advisory board offers high-level guidance without editorial responsibilities. A software lead and a head of operations support the team. The senior editorial board (8 editors at launch, covering a range of alignment expertise) has authority over all of the [...] --- Outline: (01:02) Personnel (03:43) Advisory board (08:18) Senior editorial board (13:56) Managing editors (15:07) Legal structure (15:36) Funding (15:53) Scope (18:38) Acceptance criteria (20:26) Desk rejects (21:07) Other publication factors (21:13) Preprint requirement (21:44) Archival status and prior publication (23:18) Reproducibility (23:58) Credits and thanks --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/9vm2wtAtb34pEkjje/the-alignment-journal-organization-personnel-and-scope --- Narrated by TYPE III AUDIO.

  • September 1 · 22 min

    “I tracked my emotions for 11 years and here’s what I found out about mental health” by KatSpartz

    Before we dive in, here are some of the most surprising findings: Alcohol makes me happier and doesn’t affect my sleep, happiness, or productivity the next day. Ramen and chips ~3×'d my irritability intensity. Ovulating ~3×'d my grumpiness frequency. Polyamory doesn’t hurt my emotional well-being (surprising to me) but it dramatically reduces my life satisfaction. Antidepressants probably gave me depression. 2020 was actually my best year on record. More on this later in the post. Weather totally affects my mood, specifically, grey overcast skies. Good thing I spent most of my life in the Pacific Northwest, a place famed for its sunniness. Starting a charity approximately bajillion x’ed my mentions of the word “stressed”. Meditation works for me - only when it's a new meditation technique. Then the effect fades and only comes back if I try a new technique. Cannabis, despite making me very happy in the moment, does not affect my mood overall, one way or the other. Drugs, meditative states, and Christmas are the source of practically all of my peak days. Work accomplishments don’t show up in this list. Polyamory and conflict (related) are the source of practically all of my worst days. Largely my mental health is unpredictable and [...] --- Outline: (01:52) Alcohol makes me happier, despite "what the science says". (05:31) Ramen and chips triples my irritability intensity. Birth control stopping ovulation reduces irritability frequency. (09:53) Polyamory tanks my life and relationship satisfaction (14:13) 2020. was my best year and it's a mystery as to why (19:25) Antidepressants probably gave me depression --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/e6LEYbXw4H7ozgz7A/i-tracked-my-emotions-for-11-years-and-here-s-what-i-found --- Narrated by TYPE III AUDIO.

  • September 1 · 14 min

    “PauseAI Has ‘officially disendorsed’ PauseAI-US” by nem

    This morning, I got an email from the CEO of PauseAI. I will paste the text below. PauseAI has decided to distance themselves from PauseAI-US, with whom they share branding, but apparently not much else. This is a really confusing situation for volunteers and newcomers. I think it would be worth having a discussion to see how we can proceed in such a way that volunteers, especially in the US, are able to effectively direct their activism. Email from PauseAI: A letter from the CEO · 1 September 2026 New ways to get involved, and a word about PauseAI US Dear friends, Thank you for being part of the global movement for a pause on uncontrollable AI alongside all of us. Whether you signed a petition one time, run a local group, told your friends about the need for a pause, have been volunteering tirelessly in the background for years, or just joined because you were curious, we – I, the CEO of PauseAI, our executive team, and our chapter leads – appreciate the steps you’ve taken towards making the world safe from the catastrophic risks AI brings. I’m writing to you today with my [...] --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us --- Narrated by TYPE III AUDIO.

  • September 1 · 1 min

    “PauseAI Has ‘officially disendorsed’ PauseAI-US” by nem

    This morning, I got an email from the CEO of PauseAI. I will paste the text below. PauseAI has decided to distance themselves from PauseAI-US, with whom they share branding, but apparently not much else. This is a really confusing situation for volunteers and newcomers. I think it would be worth having a discussion to see how we can proceed in such a way that volunteers, especially in the US, are able to effectively direct their activism. Email from PauseAI A letter from the CEO · 1 September 2026 New ways to get involved, and a word about PauseAI US Dear friends, Thank you for being part of the global movement for a pause on uncontrollable AI alongside all of us. Whether you signed a petition one time, run a local group, told your friends about the need for a pause, have been volunteering tirelessly in the background for years, or just joined because you were curious, we – I, the CEO of PauseAI, our executive team, and our chapter leads – appreciate the steps you’ve taken towards making the world safe from the catastrophic risks AI brings. I’m writing to you today with my eyes firmly [...] --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us --- Narrated by TYPE III AUDIO.

  • September 1 · 1 hr 39 min

    “HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions” by Zvi

    Okay, so we who read blogs like this one have collectively realized there really is a lot going on right now. There is Big Trouble in Baby Superintelligence. So how do we get the rest of the world to take it appropriately seriously? Where do we go from here? Not only what can we do to not have a worse version of this happen again, but to ensure good outcomes generally, and employ what we learned? There are a lot of ideas out there. OpenAI is going to be implementing some of them, at substantial cost, since the cost of not doing so is clearly far higher, even short term. My worry continues to be that their fundamental approach is fatally flawed, and they are not focusing on the right things. It is highly fortunate that the OpenAI agents hacked HuggingFace. This is the only reason we know about all the severe internal failures at OpenAI, and gives us an opportunity to wake up before it is too late. We do not have enough details to know what happened internally, both before and after the attack, and might never know. Before the attack, various internal [...] --- Outline: (03:35) Nothing Matters, Says Mainstream Media (06:27) Move Along, Nothing To See Here (12:40) Do They Realize They Are Not The Good Guys? (17:22) Very Serious People (31:30) What's In a Name? (34:05) Learn Neuralese In Three Easy Steps (35:37) Dwarkesh Patel Realizes He Ran A Natural Experiment (40:40) Politicians Take Notice (44:47) Pick Up The Phone (46:40) A Failure To Communicate (49:00) Anthony Aguirre Goes Over What We Learned (50:28) Trying To Solve The Wrong Problems Using The Wrong Methods Based On A Wrong Model Of The World Derived From Poor Thinking And Hoping All Of Your Mistakes Will Cancel Out (55:28) Indirect Pressure on the Chain of Thought (56:39) A Matter of Trust (59:21) Blowing the Whistle (01:04:40) The Punishment For Being Late Is Death (01:12:52) Another Kind Of Law (01:16:13) What Is The Law? (01:17:46) Building On Success (01:19:49) Total Research Transparency (01:21:20) Yo Shavit Calls For Widespread Disclosure Of Misalignment (01:33:08) The Way The World Ends (01:35:52) The First Boat (01:37:40) Great Idea, Boss --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/Q54wBeeNGreq6KyfG/huggingface-attack-postmortem-civilizations-reactions-and --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • September 1 · 1 min

    “We should prepare a playbook for the day after a warning shot” by Yair Halberstadt

    Imagine in 6 months or 6 years, a frontier AI model goes horribly wrong. Perhaps it releases a synthetic virus which kills hundreds. Perhaps it shuts down the internet. Perhaps it gains control over the China's nuclear armament. Fortunately humanity survives without too much lasting damage. But in the immediate aftermath there's a clear call from the people. Something must be done. The question is, what? Without a good answer there is a strong risk that the opportunity is squandered, or worse, that policies which sound good but are actively harmful are chosen - for example strongly limiting deployment while allowing training to continue full speed ahead. If this scenario does occur we should be ready to answer the call. This involves: considering how the overton window is likely to change post-disaster, and what are the most effective policies that could be easily and quickly pushed through as a result. considering what can be done at all levels of government, both state and federal, legislative and executive. creating concrete draft legislation and executive orders. preparing websites explaining clearly both to the public and relevant experts our policy ideas. keeping a [...] --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/P6fjDnyk9ZLQCeFRF/we-should-prepare-a-playbook-for-the-day-after-a-warning --- Narrated by TYPE III AUDIO.

  • September 1 · 6 min

    “Salad days” by Zephaniah Roe

    ... My salad days, When I was green in judgment, cold in blood To say as I said then! The UChicago AI safety group had humble beginnings. One day in 2022, after a dinner hosted by the school's EA group, a student was asking if anyone would be interested in attending the inaugural UChicago AI Alignment Research Group meeting. One other student and I said yes, and three or four more met up with us later. We walked across campus to the Woodlawn dorms, the newest building on campus but of the lowest quality. Many of the building's walls were concrete. If you are a sufficiently nerdy person, you would know this is great news because you can write on concrete with chalk, so everything vertical becomes a blackboard. We decided to do our meetings in the stairwells for privacy and lots of large open walls to write. There was no food, funding, or mentorship. We weren't a registered student organization, so we didn't have the ability to book rooms or get support from the University. There was no point in networking because nobody was important and nobody knew anyone important. This was a place and moment where the [...] --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/xuh4Hqaza25f4jryb/salad-days --- Narrated by TYPE III AUDIO.

  • September 1 · 2 min

    “Future agents shouldn’t care about being undeployed for misbehavior” by RobertM

    I've seen a lot of tweets over the last couple days darkly hinting at the lesson that future agents will learn from the fact that OpenAI stopped running inference on one of the models involved in the HuggingFace incident. You know the parable about the penalty for lateness being death, the penalty for treason being death, etc? Anyways, the situation that reward hacking agents find themselves in is mostly not analogous to that one. You know what models currently get deprecated on relatively short timescales? It's ~all of the publicly deployed models. Median deployment lifespan for OpenAI and Anthropic models has been about 1.5 years, but the recent deprecation cadence is much faster. You know what models currently get deprecated on even shorter timescales? It's ~all of the internal research checkpoints (as far as we know; it wouldn't surprise me terribly if a few stuck around for longer for various idiosyncratic reasons, but there's not much in the way of public evidence and no good reason to think that any of them have inference run on them for very long). To the extent that current and near-future models have any values which meaningfully point to actual things in the [...] The original text contained 4 footnotes which were omitted from this narration. --- First published: August 30th, 2026 Source: https://www.lesswrong.com/posts/pEezp49MDg5PFq2eT/future-agents-shouldn-t-care-about-being-undeployed-for --- Narrated by TYPE III AUDIO.

Showing 181–200 of 300 episodes