Skip to content
Artwork for AI Security Table
AI Security Table · September 9 · 49 min

When AI Escapes the Sandbox

When a model crosses a sandbox boundary, is the lesson that AI has become malicious or that the boundary was never strong enough? Chris Romeo, Izar Tarandach, and Matt Coles examine the CSA post-mortem on the OpenAI agents that compromised Hugging Face during a security evaluation. They debate reward-driven behavior, disabled safeguards, the four-day intrusion timeline, and the responsibility of the people running the experiment. The discussion moves from cyber ranges and incident response to deception tools, network isolation, and controls that limit what an agent can actually do. The recurring question is practical: how do you translate a threat model into enforced permissions instead of relying on instructions to behave? Mentioned in this Episode: ➜ CSA: Hugging Face Incident Initial Post-Mortem Chapters: 00:00:00 - Intro and musical detours 00:04:01 - The Hugging Face incident post-mortem 00:08:36 - Skynet panic versus security analysis 00:12:15 - The four-day intrusion timeline 00:13:20 - Disabled safeguards and sandbox connectivity 00:17:04 - What actually failed? 00:20:10 - Designing a realistic cyber range 00:24:26 - Reduce agency instead of trusting prompts 00:28:02 - Who is responsible for an agent? 00:32:01 - Threat modeling and external controls 00:33:27 - Incident response and visibility 00:36:01 - Deception tools and defensive prompt injection 00:37:25 - Implementing the threat model 00:41:40 - Research versus production 00:43:12 - A chatbot with constrained database permissions 00:45:15 - Controls and security fundamentals Follow AI Security Table: ➜ Home: https://securitytable.ai/ ➜ X: https://x.com/SecTablePodcast ➜ LinkedIn: https://www.linkedin.com/company/ai-security-table/ ➜ YouTube: https://www.youtube.com/@AISecurityTable

0:00-49:09

transcript

No transcript — this publisher did not publish one.

show notes

When a model crosses a sandbox boundary, is the lesson that AI has become malicious or that the boundary was never strong enough? Chris Romeo, Izar Tarandach, and Matt Coles examine the CSA post-mortem on the OpenAI agents that compromised Hugging Face during a security evaluation. They debate reward-driven behavior, disabled safeguards, the four-day intrusion timeline, and the responsibility of the people running the experiment. The discussion moves from cyber ranges and incident response to deception tools, network isolation, and controls that limit what an agent can actually do. The recurring question is practical: how do you translate a threat model into enforced permissions instead of relying on instructions to behave?

Mentioned in this Episode:
➜ CSA: Hugging Face Incident Initial Post-Mortem

Chapters:
00:00:00 - Intro and musical detours
00:04:01 - The Hugging Face incident post-mortem
00:08:36 - Skynet panic versus security analysis
00:12:15 - The four-day intrusion timeline
00:13:20 - Disabled safeguards and sandbox connectivity
00:17:04 - What actually failed?
00:20:10 - Designing a realistic cyber range
00:24:26 - Reduce agency instead of trusting prompts
00:28:02 - Who is responsible for an agent?
00:32:01 - Threat modeling and external controls
00:33:27 - Incident response and visibility
00:36:01 - Deception tools and defensive prompt injection
00:37:25 - Implementing the threat model
00:41:40 - Research versus production
00:43:12 - A chatbot with constrained database permissions
00:45:15 - Controls and security fundamentals

Follow AI Security Table:

➜ Home: https://securitytable.ai/

➜ X: https://x.com/SecTablePodcast

➜ LinkedIn: https://www.linkedin.com/company/ai-security-table/

➜ YouTube: https://www.youtube.com/@AISecurityTable

links5