“Empirical safety claims from frontier labs should be replicated, scrutinized, and open-sourced” by Zephaniah Roe, yix
When frontier labs like Anthropic and OpenAI publish safety or alignment research, it is often entirely empirical, closed-source, and sparse on methodological details. While it is great that they publish these results, the status quo is that labs (or soon, their agents) can claim alignment progress that no one independently verifies. The AI safety community has replicated or stress-tested some claims, but it's nowhere near comprehensive, and we expect this kind of meta-science to remain systematically neglected. We argue there should be a dedicated effort to Replicate alignment experiments from frontier labs. Scrutinize the experiments by stress-testing the methodology. Open-source replications to encourage external researchers to validate our work, build on the experiment, and further audit the lab's methods. The case to replicate safety research from labs CEOs and employees at AI companies, somewhat regularly, say that the technology they hope to develop could cause human extinction. However, their research to prevent this is often released without code or even basic methodological details (e.g., Teaching Claude Why, Beneficial RL). There's good reason to think some of these results could be fragile. Prior safety results can be contingent on details that are easy to miss, like the pinned OpenRouter [...] --- Outline: (01:10) The case to replicate safety research from labs (02:54) Replications are not shiny, but that's precisely what makes them counterfactually useful (03:38) The case to stress test (05:23) The case to open source (06:08) Replicating frontier lab work is difficult but tractable (07:07) Conclusion The original text contained 3 footnotes which were omitted from this narration. --- First published: September 20th, 2026 Source: https://www.lesswrong.com/posts/MmfzfGcQ3h3p6N9pD/empirical-safety-claims-from-frontier-labs-should-be-1 --- Narrated by TYPE III AUDIO.