
OpenAI Security: Controlling Models is Now ‘Hell’
This video is hard to summarise. A cracked cipher, an OpenAI security warning, Gemini 4 Argon, RSI paper (co-authored by a who’s who of AI), Lab White House commitments, new hacks emerging, ‘deep personas’, biology Kasparov competitions, and so much more, ending with an epic Opus outro. Patreon Exclusives: https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 01:32 - Wrong about Opus 5.5? Deciphering 16th Century Text 05:01 - Why the models keep breaking out 11:40 - What the models aren't telling us 15:36 - Gemini 4 and the race to release 19:19 - What happens when AI improves AI? 28:31 - Biology, consciousness, and what we still don't understand Joe Darrow: Not Just the Sandbox: https://x.com/joedaroo/status/2104335929293127851 GPT-6.1 Sol System Card: https://cdn.openai.com/pdf/38e3efcf-545e-44cd-99ec-2b7eb395f4cc/oai_GPT_6_1_Sol.pdf Intelligence Explosion Paper: https://casp.ac/__l5e/assets-v1/5efd4b41-deb5-4513-a0a3-b4f82d2b79ea/intelligence-explosion.pdf OpenAI Research Acceleration: https://openai.com/index/research-acceleration-view-inside-openai/ OpenAI Training Safety Cases: https://openai.com/index/towards-safety-cases-for-frontier-ai-training/ Catherine de Medicis Cipher: https://cryptiana.web.fc2.com/code/henryiii.htm Proposed du Croc Decipherment: https://claude.ai/artifact/1W7B3WxkTAEGzfv3TaKXb4 Rogue Agents Investigation: https://asymmetricsecurity.com/newsroom/rogue-agents-investigation/ OpenAI Shelves GPT-6.1 Astra: https://www.reuters.com/business/openai-shelves-new-ai-model-after-internal-safety-tests-wsj-reports-2026-09-28/ The Case for Reasoning Transparency: https://institute.deepmind.com/essays/the-case-for-reasoning-transparency/ Gemini 4 Argon: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/ OpenAI–Anthropic Rivalry: https://www.theatlantic.com/technology/2026/09/openai-v-anthropic-inside-biggest-rivalry-tech/688819/ NYT: OpenAI Security Warnings: https://www.nytimes.com/2026/09/29/technology/openai-warnings-security.html NYT: Claude’s Morals: https://www.nytimes.com/2026/09/29/us/anthropic-claude-morals-ai.html Jasmine Wang on RSI: https://x.com/j_asminewang/status/2097840245786157432 OpenAI Departures Roundup: https://x.com/Bayesian0_0/status/2105680470566686805 White House AI Commitments: https://x.com/Danmar_here/status/2105168138392183146 Sarah Heck on Safety: https://x.com/SarahKHeck/status/2105058513370448280 Sam Altman on Alignment: https://x.com/tbpn/status/2105028992843833459 Sam Altman on Agent Logs: https://x.com/sama/status/2103567198690349362 Micah Carroll: Misalignment Reports: https://x.com/MicahCarroll/status/2103665811051397256 Zuxin Liu on the Incident: https://x.com/LiuZuxin/status/2103699462648639645 Deepa Seetharaman: User Images: https://x.com/dseetharaman/status/2103585482793943203 OpenAI Revenue Chart: https://x.com/PaulBonnet/status/2105288259324567884/photo/1 Nvidia Agent Safety Platform: https://edition.cnn.com/2026/09/28/business/nvidia-ai-safety-system IntegrityBench: https://integrity-bench.com Neel Nanda on Interpretability: https://x.com/PalisadeAI/status/2104949061325652001 Biology Contest: Humans and AI: https://www.theinformation.com/articles/inside-drama-behind-biology-contest-pits-openai-agents-humans Pushmeet Kohli: SynthID Bio: https://x.com/pushmeet/status/2105314763148321102 Ataraxos and Stratego: https://x.com/ssokota/status/2105362040328159526 Benign Data and Hidden Personas: https://x.com/OwainEvans_UK/status/1999172949975392417 Anthropic: Introspection: https://www.anthropic.com/research/introspection Claude Cheating Results: https://x.com/lukaspet/status/2104634759339298930 Roon on Mathematics and Learning Theory: https://x.com/tszzl/status/2105619006488993898 GPT-4 Research: https://openai.com/index/gpt-4-research/ I.J. Good: Ultraintelligent Machine: https://incompleteideas.net/papers/Good65ultraintelligent.pdf Terence Tao’s 2024 Interview: https://www.scientificamerican.com/article/ai-will-become-mathematicians-co-pilot/ Hugging Face Incident: https://openai.com/index/hugging-face-incident-and-the-road-ahead/ Claude and Suno Music Video: https://x.com/sevdeawesome/status/2104985610012504181 Podcast: https://aiexplainedopodcast.buzzsprout.com/
















