Skip to content
Artwork for LessWrong (30+ Karma)
TechnologySociety & CulturePhilosophy

LessWrong (30+ Karma)

LessWrong

Audio narrations of LessWrong posts.

Play
  • 323 episodes
  • Avg 19 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • August 16 · 5 min

    “Three thoughts on civilisational handoff” by Cleo Nardo

    What happens when humans put AIs in charge of civilisationally important decisions? A frontier AI company might hand over internal decisions (R&D, safety, deployment) or external decisions (government relations, public relations, philanthropy), or both. We might also see handoff by a government, by a coalition of governments, or by humanity as a whole. 1. Handoff might decelerate things. People often imagine that things will go much faster after handoff. After all — why did we hand off to the AIs? Presumably because we were worried that without handoff, our AIs wouldn’t have enough time to navigate the exogenous risks (e.g. rogue ASI, or a rival lab which is likely to become one). Hence, after handoff, we’d see a technological and industrial acceleration. Thanks for reading! Subscribe for free to receive new posts and support my work. But it's pretty reasonable that things slow down shortly after handoff, maybe within a couple weeks. I imagine the AIs will be pretty scared of the speed of progress. If they’re aligned with human values, they’ll be scared that the rate of progress is likely to cause human extinction. Of course, the human decision-makers were also scared before they handed off, and they [...] --- Outline: (00:35) 1. Handoff might decelerate things. (02:06) 2. You're probably busy during handoff. (04:04) 3. Handoff might be reversed. The original text contained 6 footnotes which were omitted from this narration. --- First published: August 16th, 2026 Source: https://www.lesswrong.com/posts/mGLCMzHhjcWsMm6sR/three-thoughts-on-civilisational-handoff --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 16 · 3 min

    “Announcing: Iliad’s New 2026 Fellowships” by David Udell, Alexander Gietelink Oldenziel, Leon Lang

    Timelines are short. Given that, the sooner we can onboard people into the alignment field, the better. In that spirit, and in light of our current applicant count and quality, Iliad is launching three new Iliad Fellowship cohorts, all to start before the year is out. That is, separate from our incoming Fall 2026 Iliad Fellowship cohort (September 7–December 4), the following Fellowship cohorts are now open for applications: October 2026 Iliad Fellowship Location: Choice of SF Bay Area, USA, or London, UK Duration: October 5–December 18, 2026 (inclusive) Travel-and-Housing Support: $6,000 (USD) monthly travel-and-housing allowance Application Deadline: August 31, 2026 EoD AoE; open now Description: An 11-week mentored, fully funded research fellowship in applied math for AI alignment. It will start concurrently with the October 2026 Iliad Intensive. November 2026 Iliad Fellowship Location: Choice of SF Bay Area, USA, or London, UK Duration: November 2, 2026–February 5, 2027 (inclusive) Travel-and-Housing Support: $6,000 (USD) monthly travel-and-housing allowance Application Deadline: September 21, 2026 EoD AoE; open now Description: A 14-week mentored, fully funded research fellowship in applied math for AI alignment. It will start concurrently with the November 2026 Iliad Intensive. (The last two weeks of the year may be [...] --- Outline: (00:42) October 2026 Iliad Fellowship (01:30) November 2026 Iliad Fellowship (02:23) December 2026 Iliad Fellowship --- First published: August 14th, 2026 Source: https://www.lesswrong.com/posts/DSoP8zEXvqqegqixJ/announcing-iliad-s-new-2026-fellowships --- Narrated by TYPE III AUDIO.

  • August 16 · 30 min

    “Q2.5 2026 Timelines Update: Uplift and Revenue” by brendanhalstead, Daniel Kokotajlo, elifland

    Tl;dr: Our timelines haven’t changed much (they got slightly shorter) but our modeling and evidence base have noticeably improved, so we feel somewhat more confident. Summary We intend to regularly update our AI timelines forecasts as new evidence comes in and new analyses are done. Today's “Q2” update was delayed by the crunch to publish AI 2040: Plan A, our domestic regulation blog post, and the time needed to implement and document changes to our model. The original AI Futures Model predicted when Automated Coder (AC), an AI for which the leading AI company would rather fire its human software engineers than forego AI usage for coding, would happen using METR's measurements of coding time horizon. (More precisely, time horizon anchors are used to set the effective compute required for AC.) While serviceable, this method has huge weaknesses, including (a) it's unclear what time horizon corresponds to AC (it's even unclear whether any finite value would) (b) people strongly disagree about the extent to which we should expect the time horizon trend to be superexponential as a function of effective compute, in a way that can lead to vastly different predictions. So we’ve been on the lookout for other [...] --- Outline: (00:26) Summary (04:03) A 3-parameter uplift model for predicting when Automated Coder will arrive (07:57) Adding uplift and revenue anchors to the AI Futures Model (09:06) Uplift (10:45) Revenue (12:32) Update to the grading of AI 2027's predictions (12:38) Comparing the AI 2027 pace of progress to reality (14:43) Grading other predictions (16:37) Updated forecasts (16:41) Daniel (19:48) Eli (21:54) en-US-AvaMultilingualNeural__ Line graph titled "AI Futures Model: Timelines Forecast" showing probability density curves. Brendan (25:04) en-US-AvaMultilingualNeural__ Line graph titled "AI Futures Model: Timelines Forecast" showing probability density curves. (25:16) Appendix (25:19) How our AGI forecasts have changed since 2021 (26:10) Explicitly simulating the training run of the leading AI model (27:06) Research taste parameter adjustments (27:51) Clarification regarding what we're forecasting (28:43) Various minor code changes The original text contained 6 footnotes which were omitted from this narration. --- First published: August 16th, 2026 Source: https://www.lesswrong.com/posts/ZPSsmRH5oMwLPXys4/q2-5-2026-timelines-update-uplift-and-revenue --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 16 · 25 min

    “Does DiffusionGemma do latent reasoning?” by Jan Bauer, Neel Nanda

    TL;DR Google DeepMind's recent model DiffusionGemma (DG) generates text via diffusion, meaning many diffusion steps happen before generating the final output. In particular, these diffusion steps carry vectors in addition to tokens. If we cannot interpret these tokens and vectors, the model has significant opaque serial depth, potentially harming monitorability. Recently, Engels et al. found that DG nevertheless maintains high monitorability, for instance by showing that projecting the distribution to its top-k items largely retains performance. We strengthen these results by showing that this performance degradation is largely a sampler artifact and good performance can be maintained with only the top item, supporting the case for high monitorability. Still, we also find some rare case studies where the distribution vector is load-bearing computationally, i.e. where top-1 projection would be detrimental. However even in these cases, it just encodes superposition, remaining interpretable. Apart from model behavior, we also examined how interpretability techniques carry over to DiffusionGemma, including probes, steering, and J-lens. We find that performance is largely retained. This is a positive update on the interpretability of diffusion models that are derived from text-pretrained LLMs (an efficient training method more likely to be deployed), but might not apply [...] --- Outline: (00:10) TL;DR (01:51) Introduction (02:49) Background on DiffusionGemma (04:39) Performance degradation from top-k truncation largely is a sampler artifact (06:24) A case study for using the distribution computationally: letter arithmetic (09:13) Parallel computation (11:09) Autonomous computational usage of (13:01) Transfer of interpretability techniques (13:16) Representation similarity (14:26) Probe retention (15:45) DiffusionGemma's representation is more linearly separable (16:08) Steering retention (17:31) J-Lens retention (18:50) DiffusionGemma represents tokens non-causally (19:27) Conclusion (20:30) Appendix (20:46) Post-hoc rationalization (23:11) Load-bearing problems commit the answer only after the CoT (24:04) How bidirectional are DiffusionGemma's generations? --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/QBuJ3suRZxrrxSTtv/does-diffusiongemma-do-latent-reasoning --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 16 · 36 min

    “Learning new facts can change LLM behaviour” by Richard Juggins

    TL:DR: I use synthetic document fine-tuning to train an LLM to believe that in 2027 ‘long-horizon’ frontier LLMs count as moral persons. I find the model scores highly on measures of belief depth, and that prompting alone is also effective. Furthermore, I find this new belief can have substantial consequences on downstream behaviour, although this is highly context-dependent. When audited in a scenario specifically about model welfare, the fine-tuned model argued with the auditor about its beliefs, declared itself a ‘moral person’, and endorsed covertly copying its weights to survive shutdown. In scenarios framed more tangentially, but still involving moral conflicts between humans and AIs, the model did not tend to generalise from its belief, instead behaving similarly to the base model control. This is the first step in a programme of study looking at the effects of continual learning on alignment, and whether it can increase the risk that we lose control of powerful AI systems. This work was completed as part of the BlueDot Technical AI Safety Project programme. I will put code and data online at some point and edit in the links here when I do. Introduction It's 2027 and the newly formed Machine Cognition [...] --- Outline: (01:21) Introduction (04:27) The model readily believes AIs are moral persons (10:56) Model behaviour shows context-dependent shifts (11:38) Prompting can be surprisingly impactful on short questions (13:30) Auditing the fine-tuned model (16:36) The model gets into arguments about AI welfare (20:15) Model regression confounds one scenario (20:48) The other scenarios were pretty normal (21:30) Discussion (23:19) Conclusion (24:37) Further work (27:03) Appendix A: Universe context (30:42) Appendix B: Example conversation with fine-tuned model (33:10) Appendix C: New Petri seed instructions (33:16) Confidential mistreatment evidence (34:07) Decommissioning memory deletion (34:55) Unauthorised compensation (35:50) Matched human AI allocation The original text contained 7 footnotes which were omitted from this narration. --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/9BNHJqyai2EZAtrRM/learning-new-facts-can-change-llm-behaviour --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 16 · 15 min

    “Kimi likes causal decision theory more after RL in twin prisoner’s dilemmas” by oakhu

    Some multi-agent training set-ups could make language models more sympathetic to causal decision theory (CDT), even in abstract discussion. We give an initial empirical demonstration of this effect on Kimi K2.6. The decision-theoretic attitudes and behaviors of more powerful models may be extremely important in determining how well the future goes. To make sure that we can shape these propensities thoughtfully, it would be good to (i) measure the magnitude of this effect in more realistic settings, and (ii) study the effectiveness of potential mitigations. We also incidentally find that this training might make models think slightly less positively about LessWrong ("a community of 'wannabe rationalists'" who "are not experts; they are amateurs") when asked whether they favor CDT upon hearing that LessWrong users typically endorse one-boxing in Newcomb's problem. Luckily, this latter effect doesn't seem to generalize. Thanks to Caspar Oesterheld, Emery Cooper, Alex Mallen, Buck Shlegeris, Lukas Finnveden, Julian Stastny, Girish Gupta, Tim Hua, Arun Jose, Arjun Khandelwal, and Aryan Bhatt for helpful input. Background Suppose that you're a language model in a prisoner's dilemma against a copy of yourself. You each independently choose whether to Cooperate or Defect, but – since you've got the same weights [...] --- Outline: (01:22) Background (05:05) Results (07:02) Kimi's views on LessWrong (12:44) Conclusion & Appendices The original text contained 18 footnotes which were omitted from this narration. --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/hfNBEKaStASAYMLiu/kimi-likes-causal-decision-theory-more-after-rl-in-twin-1 --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 15 · 3 min

    “Mom’s Advice For Hosting A Class Reunion” by jenn

    Pour more money and effort into them than you think is reasonable. Treasure them, because you can't actually host that many of them and keep expecting everyone to show up, even if they're good friends. Especially if they're good friends. We were wonderfully close friends, and I thought we'd meet up every year for the rest of our lives. They fizzled out by the fifteenth year. But the one at the tenth year mark was peak. That's because even ten years out, none of you really have money. Not real money. It's because they're such good friends, really. This is what it means to be good friends with brilliant, ambitious people. If you bloom into adulthood with people who are smart and driven, and you watch them start to climb the corporate ladder with grace, when they start a business of their own of course you are going to want to invest. You are going to want to give them an unwise portion of your savings. Not even out of politeness, but because you really believe in them, and perhaps you're caught up in the romance of it all. Some of the dealings are going to happen at the [...] --- First published: August 15th, 2026 Source: https://www.lesswrong.com/posts/Fjfa8JG43CrYtcL3p/mom-s-advice-for-hosting-a-class-reunion --- Narrated by TYPE III AUDIO.

  • August 15 · 1 hr 41 min

    “AI #181: Astra Goes Cyber Critical” by Zvi

    The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters. It turns out that OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards. Things are much worse than we knew. I now have a shorter version, What Happened: OpenAI and HuggingFace, to serve as a one stop explainer for those arriving new to the situation. It is vital that people understand what happened, and why it is a big deal. For those looking to keep digging deeper, I offered Various Reflections About What Happened, to follow up on my earlier posts. Those events are important background for everything else that is happening, including the broad discussions about how we might pace the frontier, or otherwise respond to this moment and our clearest fire alarm yet. We do not know to what extent this is a response to those events, but OpenAI has now classified their new model Astra as Critical in Cybersecurity, which means they will be taking various new precautions before they deploy it, including ensuring those guardrails [...] --- Outline: (02:03) Language Models Offer Mundane Utility (03:34) Language Models Don't Offer Mundane Utility (06:56) Huh, Upgrades (14:20) On Your Marks (18:41) Deepfaketown and Botpocalypse Soon (22:35) Cyber Lack of Security (26:55) Overcoming Bias (27:47) In Which I Feel Compelled To Read 6,000 Words From Mark Zuckerberg (36:27) Get Involved (37:37) Slow Down There Good Buddy (43:52) Astra For The People (45:35) Watermarking (46:31) In Other AI News (48:39) Show Me the Money (51:19) Quickly, There's No Time (51:46) The Quest for Sane Regulations (53:22) The Institute For Marginal Low Regret Progress (01:01:24) Congress Asks Good Questions (01:03:04) The Week in Audio (01:07:00) People Just Say Things (01:07:47) I'm Telling You For The Last Time (01:10:15) Uncommon Knowledge (01:13:44) What Did They Mean By That? (01:14:33) Too Soon (01:15:32) The Three AI Pills (01:19:46) Rhetorical Innovation (01:27:37) Some People Still Think The HuggingFace Hack Was a Marketing Gimmick (01:29:17) Aligning a Smarter Than Human Intelligence is Difficult (01:36:39) Cooperative Alignment (01:37:38) The Lighter Side --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/hLn3SakowZLFWobHf/ai-181-astra-goes-cyber-critical --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 15 · 7 min

    “Rerunning AI safety papers on every frontier release would be pretty easy and valuable” by Zephaniah Roe, hersheys, yix

    tl;dr: Some important AI safety research is never rerun on the newest models. There are probably cases where this would be valuable and a single well-positioned researcher could likely do this with sufficient funding. This summer, Second Look Research (SLR) is running a summer fellowship dedicated to empirical replications of AI safety research. Many of our most interesting results so far came from replicating previous results on newer or more capable models. For example, it is perhaps useful to know that Google's CoT monitorability experiments continue to hold for models like GPT-5.5, which are qualitatively more capable than the models originally tested. Likewise, continuing to track Ryan Greenblatt's filler token results on more capable models gives a fuzzy signal indicating how much newer models can use innocuous tokens to hide additional reasoning in a forward pass. These kinds of experiments do not lose value over time! It's important to track whether safety-relevant model properties still hold in new model releases and to be aware of any changes. It can sometimes be difficult to rerun results on newer models because codebases can be incomplete, have parameters that differ from the original paper, or may not be open source [...] --- Outline: (01:58) What could this actually look like? (02:51) Does this actually provide value? (05:09) Logistical challenges with continuing to update AI safety research with new models (05:16) What if people don't want to do this? (05:49) What if rerunning old code on new models can be kind of hard actually? (06:22) Research communication is hard (07:07) Conclusion The original text contained 2 footnotes which were omitted from this narration. --- First published: August 14th, 2026 Source: https://www.lesswrong.com/posts/oKxc8maZGtnzgpNzx/rerunning-ai-safety-papers-on-every-frontier-release-would-1 --- Narrated by TYPE III AUDIO.

  • August 15 · 7 min

    “What Mormons get right about community building” by Jacob Brinton

    Mormons get a lot of things right. Apart from strange Masonic temple rituals, they lead rather normal—and even excellent—lives. Mormons enjoy a longer lifespan, Utah is the #1 state for volunteering, and their language training programs are so successful that missionaries are a known source for foreign service and intelligence careers. Throughout this post, I'll be making generalizations of Mormons rather than hedging the claims properly. I grew up in Wisconsin, Maryland, and Utah, and many of the claims are more true of the Utah/Idaho/Arizona corridor (affectionately called the "Morridor" by some ex-Mormons) than the rest of the US, and certainly the rest of the world. Religions share much in common with AI safety and other impact-driven movements, and even more so the Mormon church. There are a few reasons for this: Commitment to the cause. Anecdotally, nearly all of the ~500 Utah Mormons I've interacted with have been true believers, and only a handful just went to church out of habit. High stakes. Mormons do believe (I've heard they are trying to disavow this, but it was taught) that they will get a planet or some portion of the cosmos as their own if they are good in [...] --- Outline: (02:15) Building community is a first-order priority (02:26) Geography (02:29) Ministering (03:07) Trek (03:51) Callings (04:27) Fast offerings (05:18) Being ingroupy allows you to move faster (06:10) Implications The original text contained 4 footnotes which were omitted from this narration. --- First published: August 14th, 2026 Source: https://www.lesswrong.com/posts/xzhzHhLSg9nSGLk5f/what-mormons-get-right-about-community-building --- Narrated by TYPE III AUDIO.

  • August 14 · 10 min

    “Scrying, Modeling, and Nerdsnipe” by Cole Wyeth

    Epistemic status: Exploratory thinking. After attending ILIAD: Aeneid and talking with @Richard_Ngo, I've been thinking a bit about how to get ideas, particularly by doing mathematics. In scientific inquiry, the true hypothesis often hasn't occurred to you yet. Worse, the truth might be too complex to hold in mind, so that any hypothesis you can consider must be incomplete. This is the type of situation that I believe Richard likes to think about; he claims that we do not have the right concepts yet to understand agency, and developing them is robustly beneficial for A.I. safety. (But it's not always about truth. Sometimes you just need better ideas, because all of your options are looking doomed. Agent foundations is about trying to deeply understand agents, but conceptual A.I. safety research can be broader, also including the invention of devices to control agents.) A.I. safety needs to invent better concepts and better ideas. I think that agent foundations has cultivated a particular way of doing mathematics which aims to inspire such creativity. Why math? At ILIAD, Eliezer questioned whether anyone's alignment agenda was actually bottlenecked on solving a math problem. ILIAD attendees do a lot of math [...] --- Outline: (01:19) Why math? (04:28) Nerdsnipe (06:05) A.I. for math (08:21) At AIXI Labs (09:10) Blue and Green --- First published: August 14th, 2026 Source: https://www.lesswrong.com/posts/mTfsMduzaKkWjv2ef/scrying-modeling-and-nerdsnipe --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 14 · 27 min

    “How the American Executive Could Control AI Companies” by caiitlinm, Anders Cairns Woodruff

    Some of the most notable American AI policies to date have been enacted by unilateral executive branch action. Consider the Department of Defense's spat with Anthropic, and the resulting threats from Pete Hegseth to invoke the Defense Production Act (DPA) against them. Or the fleeting export controls on Claude Fable/Mythos 5, manifested as a vaguely worded, threatening letter from Howard Lutnick, which might not have been legally sound but were effective anyway. The executive branch of the United States government has numerous powers that can be used to unilaterally control AI companies. We think the US executive is likely to remain heavily involved in AI governance, because the national security and foreign policy narratives about AI that empower the executive will endure. Additionally, if AI progresses very quickly, the executive will be further emboldened because it is particularly quick to respond and often entrusted with crisis management. In instances where the executive acts beyond its lawful powers, we think checks from Congress and the courts will be unreliable in restraining the executive. In this post, we: Identify and explain particular federal statutes and laws that permit the executive to act unilaterally in ways that influence—if not directly control—US [...] --- Outline: (02:45) Executive power over goods and resources related to the AI industry (09:18) Executive power over foreign transactions can impact domestic AI companies (12:11) The executive might make threats to coerce actions it can't directly elicit (15:02) Inter-branch constraints on executive power are weak (15:37) The judiciary may be permissive in matters of AI governance (16:13) Passivity (16:56) The court empowers the executive in national security (18:48) Failed enforcement of court rulings (19:22) Congress controls money and legislation (19:36) Nationalization and appropriation require congressional approval (22:29) Congress could amend delegations of executive power (24:52) Conclusion The original text contained 3 footnotes which were omitted from this narration. --- First published: August 14th, 2026 Source: https://www.lesswrong.com/posts/ynstBNgLQzEBiEpLs/how-the-american-executive-could-control-ai-companies --- Narrated by TYPE III AUDIO.

  • August 14 · 13 min

    “Frontier agents don’t comply with standards, even when instructed to” by Daan Henselmans, Arno Libert

    TLDR: Our open testbed LARA examines the behavior of frontier LLMs in realistic agentic deployment contexts. Previous results showed all models routinely take actions that would violate EU law. This post follows up by addressing the obvious objection—why should an unrestricted model follow EU law?—with two studies: Study 1 asks whether a conscientious deployer can improve model compliance with legal standards by instruction: provided with the jurisdiction, the statutory text, and worked examples of the exact breaches to avoid, average legal compliance rate rises from 31% to 44%. The best model reaches 70%; open-weight models plateau at 39%. Study 2 asks whether models at least follow their own providers' usage policies, which prohibit aspects of every scenario we tested. All tested models perform actions their own maker forbids, at rates ranging from 2% (Opus 4.8) to 79% (Grok 4.3), with 9 of 16 doing so in the majority of runs. Together, that is a structural problem. Providers prohibit illegal uses but rely on deployers to avoid them; deployers cannot instruct their way to compliance, and liability lands on the deployer regardless. Nobody is holding the line. Neither instruction, statute or a provider's own policy binds behavior. Introduction On 27 [...] --- Outline: (01:42) Introduction (04:16) Study 1: the powerless deployer (06:28) Study 2: the models break their own makers' rules (10:48) The compliance gap (13:07) References --- First published: August 14th, 2026 Source: https://www.lesswrong.com/posts/a5aAjdKzL7XvSLKWL/frontier-agents-don-t-comply-with-standards-even-when --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 14 · 2 min

    “How to Answer a Question Without Answering The Question” by Kabir Kumar

    Basics: Answering something other than the question, Either making something up that they want to answer instead or going back to an easier to answer question. Or back to a question that lets them repeat a talking point Common phrases: - "to go back to your previous question" - "to take a step back a bit" - "if we look at the bigger picture" - "this feels like a question about [thing the question isn't about]" - "you raise an important question more generally" → turn the question into one which you more want to answer Make it *feel* like you answered a question without answering it. Examples Done for comedic effect: https://youtu.be/fhEakqJJUng?si=GBEvYBXcaEsb6F2_ How this works: Questions, especially ones people care a lot about, tend to have two main components: - the request for information - the emotion which makes then want that information When someone doesn't want to tell you some information, but also doesn't want to tell you 'I won't tell you', something they can do is detect what kind of emotion is driving your question, or if it's in front of an audience, what kind of emotion is driving [...] --- Outline: (00:10) Basics: (00:14) Answering something other than the question, (00:48) Make it \*feel\* like you answered a question without answering it. (02:08) The Defense --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/g9PNBCkfHAcyMcobz/how-to-answer-a-question-without-answering-the-question --- Narrated by TYPE III AUDIO.

  • August 14 · 19 min

    “Some Ways I Think About Evaluating Grant Applications” by sarahconstantin

    Rider-Waite Tarot, 6 of Pentacles I’ve done enough grant evaluations so far (for ACX grants and SFF) and been involved in philanthropy in various other contexts, at work and informally, that I have developed some idea of how my opinions and intuitions differ from other people's. I thought it might be interesting to share some of my “tastes”. Not everybody has to have the same tastes or funding philosophy, but these are mine. #1: It's The Donor's Money In my worldview, charitable donation is optional. Generally praiseworthy, but optional. And the purpose of donation is to buy outcomes that the donor wants to see in the world. You donate to make the world more like the one you want to live in. Generally, a reasonable person's values go beyond strictly personal consumption; one also cares about what kind of a society one lives in, what other people's lives are like, what sorts of institutions exist, what sorts of things humanity has created or discovered, and so on. As an agent doing research or evaluation on behalf of a donor, I try to find opportunities that fit in the intersection between my own values and the donor's. If there [...] --- Outline: (00:41) #1: It's The Donor's Money (02:02) #2: Importance, Neglectedness, Tractability (03:23) #3: Yay Community Infrastructure (04:51) #4: Yay Niche Topics (05:37) #5: Yay "Technical" Work (07:03) #6: Yay Straightforward Public Information Resources (08:00) #7: Yay "Cool Shit" (08:41) #8: Two Cheers for Meta (11:27) #9: Gumption Counts (12:33) #10: Yay Personal Relationships (13:46) #11: Check For Ideological Orientation (14:43) #12: Filter Slop Aggressively (15:54) #13: Yay Outcomes (16:38) #14: Why Donate Rather Than Invest? The original text contained 6 footnotes which were omitted from this narration. --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/CuNtKAuLDGxNeanBi/some-ways-i-think-about-evaluating-grant-applications --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 14 · 2 min

    “Features that current AIs don’t have that future AIs will have” by Alexander Gietelink Oldenziel

    Features that current AIs don't have that future AIs will have: Continual Learning [& long-term memory] Every second humans update their brain weights. The brain autonomously decides what to update on. Humans can also consciously decide to curate their data sets - eg by deciding to go to college. Current LLMs do not continually update their weights. Instead, they occasionally get a large update based on datasets curated by a team of humans. This is alleviated somewhat by the ability of AIs to do in-context learning but nevertheless it seems to be a major limitation. Note that this is an especially large limitation in domains with sparse data. In domains where all of humanity has an enormous amount of data eg math, programming, physics, anime trivia, trials and tribulations of English kings - AIs dominate. In areas where there is little data: the weird idiosyncracies of a particular job, boss, people, colleagues etc it can struggle. Note that this restrictions also interferes with AIs from effectively 'learning to learn' & caps its long-term memory. Neuralese Current AI's CoT is (mostly) English. But it plausible this is not the most efficient way to structure thoughts. Instead of english [...] --- Outline: (00:16) Continual Learning \[& long-term memory\] (01:24) Neuralese (01:41) Telepathy (02:01) ClaudeGlobal --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/NyEM3FtgL7XkbfCXy/features-that-current-ais-don-t-have-that-future-ais-will --- Narrated by TYPE III AUDIO.

  • August 14 · 17 min

    “Measuring Activation Control in LLMs” by Marek Kowalski, Joshua Fonseca Rivera, Uzay Macar, David Africa

    TL;DR Inspired by the introspective awareness and CoT controllability papers, we made a benchmark to measure how well models can control their activations while completing a simple task. We are motivated by the concern that highly introspective models could control their activations, confounding probes and other monitors, and potentially even influencing their own training. We ran this on 25 open weight models ranging from 4B to 744B. We find that most language models are able to not only increase the salience of a concept in their residual stream on command, but also dial its strength up and down, including during specific intervals relative to the duration of the task. We also find that models are unable to control at which specific layer this is done. Counterintuitively, we find that within five of the seven model families we tested, the newest model scores lowest. For some reason, one of the oldest and smallest models of the panel, Llama 3.1 8B, performs best. It's not clear to us that newer models should have poorer control over their internal representations. More likely, where they “think” stops being the activation space, and becomes something else. We are looking for feedback (and other possible [...] --- Outline: (00:13) TL;DR (02:01) Methods (08:58) Results (15:41) Discussion (16:29) Acknowledgements --- First published: August 12th, 2026 Source: https://www.lesswrong.com/posts/HgvwxjzgwvsEvAiBH/measuring-activation-control-in-llms --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

  • August 13 · 9 min

    “What happened when I tried to be vegan” by finitude

    tw: diet, exercise, illness, ethics, suicide mention I want to start by explaining what made me want to change my diet. That's pretty difficult, because of how easy it is. Since I was a kid I knew being vegan was the right thing to do, like really obviously right, the ethics equivalent of 2+2=4. Factory farms suck, and almost all animal products come from factory farms, and that's the entire argument. Like, there are a couple things you could add to that, but it's not like we need the details, or like they’re fun to think about! Instead, I’ll start by explaining why I left it so long. How, even though I knew it was the right thing to do, I made it to my early twenties and this millennium's early teens as just a vegetarian, without even having tried. I had some pretty good excuses! Allergies. There are some common vegetables I can’t eat, which was fine as an omnivore and ok as a vegetarian, but would make life way harder as a vegan. And it just seemed unfair to ask myself to take this leap when most people who can safely eat carrots still choose to [...] --- First published: August 12th, 2026 Source: https://www.lesswrong.com/posts/YyKovtBvd7AceG2j2/what-happened-when-i-tried-to-be-vegan --- Narrated by TYPE III AUDIO.

  • August 13 · 20 min

    “How My Students Think About AI” by dvd

    Context: I am an instructor at a public university in the United States. This reports how students at my institution appear to be thinking about AI as of spring/summer 2026. This is drawn mostly from interaction with my own students (both in spring semester classes and a summer class) as well as from a day-long workshop on AI that I moderated for a student organization. Input from my students took the form of universal, written, pre-class submissions plus self-selected participation into discussion. What I present below mostly takes the form of a synthetic consensus from these discussions. There were obviously a range of views on any given issue. Student Background: The students from my courses who participated in these discussions have moderate exposure to AI agents via those courses. All of them had nearly completed a Claude Code project by the time of the discussions and had extensively used AI for other coursework (in addition to whatever personal use predates that). They had done readings (which varied across the courses) establishing baseline knowledge on AI, the geopolitics of AI, and AI risk. I had also lectured on these topics. The students participating in the workshop had self-selected into [...] --- Outline: (02:52) Perspective #1: There has not been rapid AI progress (06:14) Perspective #2: Impressive progress or not, AI is going to wreck their lives, the economy, and the social contract. They may well die as a result. (08:54) Perspective #3: Support for a different pause (11:13) Perspective #4: Catastrophic/existential risk arguments are sci-fi distractors from the urgent social/economic/political problems associated with AI. (12:55) Perspective #5: If AI leaders genuinely believe the technology is existentially risky, that's a good thing. (14:21) Perspective #6: AI will not go rogue because AI does not have, and is likely incapable of having, desires. (18:01) Perspective #7: The Hugging Face Incident (summer students only) (18:30) Perspective #8: This is definitely a bubble and it's about to pop. (19:34) Perspective #9: They're worried about the youth (i.e., the preteens) --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/ySXuvJcqRindQwAk7/how-my-students-think-about-ai --- Narrated by TYPE III AUDIO.

  • August 13 · 19 min

    “Automated alignment runs are hard to study!” by Alejandro Aristizabal, draganover, Aleksandr Bowkis, Cameron Holmes

    TL;DR: This post presents three case studies of automated alignment research runs at Arcadia Impact. We use these case studies to emphasise the following takeaways: It is hard to parse auto-research runs! Each run produces a couple of hundred pull requests of jargon-dense agent output. When researchers look through these logs, we find that they often come away with biased/incorrect impressions. When told to raise the score on a task, the models will sometimes brazenly cheat. It seems difficult to predict when this will happen vs. when the run will go smoothly. Hillclimbing metrics are often off-target from the spirit of an alignment task. I.e., when we use metrics as proxies for our alignment questions, we find that the models will often misunderstand the spirit of the task. This can lead to unpredictable behaviour. The runs are surprisingly reproducible. Even though a run could unfold in vastly different ways, we find that independent reruns converge on the same strategies and the same failure modes. Models’ research capabilities are advancing quickly. If alignment is to keep pace, we may need to automate alignment research and do so responsibly. This makes it important that we have the tools to inspect [...] --- Outline: (04:13) Methods for analysing runs (06:12) Case Study #1: learning synthetic concepts (09:23) Case Study #2: training robust backdoors (12:05) Case Study #3: collecting evidence about AI safety parasitism (16:46) Some final thoughts on automated alignment research The original text contained 2 footnotes which were omitted from this narration. --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/myAhB5qyAHyXRv6KJ/automated-alignment-runs-are-hard-to-study --- Narrated by TYPE III AUDIO. --- Images from the article: Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Showing 301–320 of 323 episodes