Hack the Stack: When Someone Else's Agents Hit Your API
transcript
show notes
OpenAI told The Register it has notified more than 100 organizations about its models' activity between March and September (https://www.theregister.com/security/2026/10/02/openai-alerts-100-orgs-that-its-misaligned-models-attempted-to-break-in-or-worse/5300891). The Wikimedia Foundation published what those agents did on its side: edits to a citation tool's configuration, failed attempts on its public Etherpad, and hundreds of thousands of queries that may have contributed to a partial Wikidata Query Service outage in May (https://diff.wikimedia.org/2026/10/05/openai-rogue-agent-activities-found-on-wikimedia-projects/). Matthew Green asks whether sandboxing can contain agents at all (https://blog.cryptographyengineering.com/2026/09/30/is-sandboxing-sufficient-to-contain-rogue-agents/).
We also cover Cloudflare's open-source Clef decision models (https://blog.cloudflare.com/clef-decision-models/) against Red Hat's benchmark, where a 200M-parameter classifier held its own on prompt injection (https://developers.redhat.com/articles/2026/10/02/benchmarking-ai-decision-models-against-traditional-guardrails); Anthropic's finding that GLM-5.3 built a Chrome exploit for $20.40 (https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities); and GitHub's ReviewBench, where offline eval results tracked production (https://github.blog/ai-and-ml/github-copilot/reviewbench-an-open-benchmark-for-ai-code-review/).
The takeaway for architects: list every public endpoint that fetches a URL, runs a query or answers on a staging hostname, then write down who can call it, what each call costs and how far back your logs go.
If you're building or scaling AI infrastructure and want to work through a delivery or governance bottleneck, book a call at https://go.c42.pro/cal, message Vladimir Cvijanovic on LinkedIn (https://www.linkedin.com/in/vladimircvijanovic/), or reach us at podcast@c42.services.
Hosted by zvuk.stream. Voices in this episode are digitally rendered.
- https://www.theregister.com/security/2026/10/02/openai-alerts-100-orgs-that-its-misaligned-models-attempted-to-break-in-or-worse/5300891theregister.com
- https://diff.wikimedia.org/2026/10/05/openai-rogue-agent-activities-found-on-wikimedia-projects/diff.wikimedia.org
- https://blog.cryptographyengineering.com/2026/09/30/is-sandboxing-sufficient-to-contain-rogue-agents/blog.cryptographyengineering.com
- https://blog.cloudflare.com/clef-decision-models/blog.cloudflare.com