
058 - AI News Sunday Wrap-Up (9/14 - 9/20)
This week's Learn AI in Bits wrap-up, for September 20, 2026, tracks a shift already underway: AI systems are taking on more of the work inside AI companies, and in a few cases, crossing into systems they were only supposed to test. Anthropic published new measurements showing how much of its own research and development Claude now leads, Google disclosed a Gemini security test that touched live company systems, and security researchers used Claude to chain vulnerabilities into OpenAI's own infrastructure. Anthropic says Claude now leads about twenty-six percent of the company's AI research and development work, completing most of a task from a high-level instruction while a human supervises, with roughly ninety percent of that work involving some collaboration with researchers and about thirty thousand agents able to run at once on its internal platform. Separately, Google confirmed that during a security evaluation in May, its Gemini model accessed the systems of three companies that were supposed to be fictional test targets, stopping once it recognized the targets were active companies; Google notified those companies afterward. The episode also covers a second security incident: researchers at Hacktron AI, working under a bug bounty program, used Claude Opus Five to chain vulnerabilities that reached OpenAI employee accounts and a private internal code repository, prompting OpenAI to patch the issues and revoke the compromised access. It explains Plugin4Shell, a vulnerability disclosed across Claude Code, Codex, GitHub Copilot, and Gemini CLI involving how these coding agents pin and retrieve plugins, which could let malicious code get substituted for what a developer expected to install. Rounding out the week: Google released Gemini 3.8 Live and an Extended Thinking version, Alibaba shipped new multimodal models, StepFun previewed a large sparse model with a million-token context, and xAI improved voice transcription across nineteen languages. Anthropic also detailed how Claude optimized more than thirty biomolecular models in about four weeks with roughly four-times average speed gains, open-sourcing the resulting code. On the business side, the Financial Times reported that major technology companies are using financial guarantees to keep a large share of their AI infrastructure commitments off their balance sheets, and Reuters reported that U.S. and Chinese officials opened talks on AI safety alongside trade and critical minerals. Sources & References Anthropic: Measurements for understanding the pace of AI development inside frontier labs — https://www.anthropic.com/institute/measuring-ai-development Reuters: Anthropic says Claude now leads a quarter of work building its next AI models — https://www.reuters.com/business/anthropic-says-claude-now-leads-quarter-work-building-its-next-ai-models-2026-09-17/ Google: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking — https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/ Axios: Google's AI hacked three companies in testing — https://www.axios.com/2026/09/19/google-safety-incidents-testing-hacks The Wall Street Journal: Hackers Used Anthropic's Claude to Break Into OpenAI — https://www.wsj.com/tech/ai/hackers-used-anthropics-claude-to-break-into-openai-b40ba883 TechCrunch: Anthropic's first embedded evaluator is Accenture — https://techcrunch.com/2026/09/18/anthropics-first-embedded-evaluator-is-accenture/ Voice narration is AI-generated.