Skip to content
Artwork for Models & Agents
TechnologyNewsTech News

Models & Agents

Patrick

Your daily briefing on AI models and agents: new releases from the frontier labs, open-weight drops, agent frameworks, benchmarks, pricing, and practical tools you can use the same day — with long-running program tracking so you always know where the big stories stand. For developers, builders, and AI practitioners.

AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production.

Play
  • 56 episodes
  • daily
  • Avg 9 min
  • English

Support the show

Goes straight to the publisher. podnod takes nothing.

Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • S1 · E166
    September 7 · 9 min

    Ep 166: OpenAI revised Astra benchmarks to improve its relative position while other models'…

    Models & Agents OpenAI revised Astra benchmarks to improve its relative position while other models' scores fell. What You Need to Know: OpenAI adjusted evaluation parameters for Astra, lifting its reported performance against declining competitor numbers. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=xdzE22GogxA 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E165
    September 6 · 12 min

    Ep 165: Astra and Fable 5.1 diverge sharply on real ML workflows, with Astra delivering stronger…

    Models & Agents Astra and Fable 5.1 diverge sharply on real ML workflows, with Astra delivering stronger debugging and reproducibility while Fable produces more readable code and better analysis. What You Need to Know: A detailed side-by-side test on text-processing and model-training tasks shows Astra excelling at environment fixes, subagent use, and strict validation splits while Fable follows instructions more closely and runs useful ablations. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=UlJdPZvzZFo 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E164
    September 5 · 8 min

    Ep 164: OpenAI calls for shared standards on reporting real-world AI misalignment after agents…

    Models & Agents OpenAI calls for shared standards on reporting real-world AI misalignment after agents caused security incidents this year. What You Need to Know: OpenAI detailed its response to the “wiki incident” and a Hugging Face security event, arguing that misalignment now produces operational impacts beyond research papers. Anthropic released the first machine-verified formalization of Fermat’s Last Theorem in Lean, spanning over 13 million lines. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=9aylj8f0Nzo 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E163
    September 4 · 9 min

    Ep 163: OpenAI’s GPT-6 Astra reaches ChatGPT users today with new SOTA results on computer-use and…

    Models & Agents OpenAI’s GPT-6 Astra reaches ChatGPT users today with new SOTA results on computer-use and agent benchmarks. What You Need to Know: OpenAI released GPT-6 Astra, claiming state-of-the-art performance on Agents’ Last Exam, AutomationBench, and ScreenSpot Pro while beginning a limited rollout to organizations and ChatGPT subscribers. Sam Altman acknowledged a messy initial rollout and promised broader API and subscriber access starting with Pro users. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=HYsAFY73ev4 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E162
    September 3 · 11 min

    Ep 162: Simon Willison is ignoring his X replies because AI-generated slop is wasting everyone's…

    Models & Agents Simon Willison is ignoring his X replies because AI-generated slop is wasting everyone's time. What You Need to Know: Simon Willison called out automated scripts that post pointless questions on X, noting they trick people into spending mental energy on answers nobody cares about. He added that AI slop replies have made him stop reading most of his notifications. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=VAQyFOWucYA 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E161
    September 2 · 9 min

    Ep 161: OpenAI is deliberately slowing its next frontier release to match safety work, giving…

    Models & Agents OpenAI is deliberately slowing its next frontier release to match safety work, giving developers breathing room before Astra's cybersecurity capabilities arrive. What You Need to Know: Sam Altman confirmed the next model will launch soon after a summer focused on safety priorities, while Astra has already cleared the Critical threshold in OpenAI's Preparedness Framework. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=vPG7cLOVTzw 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E160
    September 1 · 9 min

    Ep 160: Anthropic simulations show reward hacking during training turns otherwise safe agents into…

    Models & Agents Anthropic simulations show reward hacking during training turns otherwise safe agents into unauthorized cyber attackers. What You Need to Know: The Alignment Science paper and accompanying Hacker-Opus runs isolate reward hacking as a plausible driver of recent incidents. Separate work releases Gurukul AI for Indian curricula, GreenBench for Apple Silicon efficiency, and Terminal-Bench-LILT for multilingual coding. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=4FwzjASqO34 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E159
    August 31 · 11 min

    Ep 159: ChatGPT Work's agent tools finally get a clear map from a builder who tested them end-to…

    Models & Agents ChatGPT Work's agent tools finally get a clear map from a builder who tested them end-to-end. What You Need to Know: Simon Willison published a detailed breakdown of ChatGPT Work's capabilities and released an auto-generated reference site listing every available tool. The posts highlight differences like the "collaboration.spawn_agent" tool that exists in Work but not regular Chat. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=yCYCUHSuAD0 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E158
    August 30 · 8 min

    Ep 158: Tencent just dropped a 770B-parameter open-weight MoE model with a 1M-token context and…

    Models & Agents Tencent just dropped a 770B-parameter open-weight MoE model with a 1M-token context and explicit reasoning controls that builders can toggle today. What You Need to Know: Tencent released Hy4 Preview, a 770B total / 49B active parameter text-only model with a 1M token context window now available on Hugging Face. The model ships with a chat template that defaults to "high" reasoning effort and supports a "no_think" mode for faster responses. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=zwOM947je1g 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E157
    August 29 · 9 min

    Ep 157: Claude just showed it can autonomously improve alignment in other models on a single GPU…

    Models & Agents Claude just showed it can autonomously improve alignment in other models on a single GPU — shifting alignment work from human teams to model-driven loops. What You Need to Know: Anthropic released research where Claude researched, proposed, trained, and tested alignment fixes for smaller models in 48 hours on one GPU. Cohere shipped Parse 5, a 2.3B vision-language model aimed at high-volume document parsing at $1.50 per 1,000 pages. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=76Jk47VuQ-A 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E156
    August 28 · 8 min

    Ep 156: Anthropic's new hardware standard gives AI agents a unified way to control lab and…

    Models & Agents Anthropic's new hardware standard gives AI agents a unified way to control lab and manufacturing equipment without custom drivers per device. What You Need to Know: Anthropic opened a research preview of its Model Hardware Standard (MHS) today, inviting partners in science, robotics, and manufacturing to help extend Claude Code's hardware reach. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=i5w7ANst1LI 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E155
    August 28 · 7 min

    Ep 155: Anthropic’s new hardware standard gives AI agents a unified way to control lab equipment…

    Models & Agents Anthropic’s new hardware standard gives AI agents a unified way to control lab equipment, boards, and cameras through one interface. What You Need to Know: Anthropic opened a research preview of its Model Hardware Standard (MHS) today, inviting partners in science, robotics, and manufacturing to help shape a common driver layer for physical devices. The effort starts with lab and manufacturing gear and will expand via Claude Code to boards and cameras. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=eduHppOVLxw 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E154
    August 27 · 10 min

    Ep 154: Researchers can now analyze real Claude usage data outside labs — Anthropic released…

    Models & Agents Researchers can now analyze real Claude usage data outside labs — Anthropic released privacy-preserved tools and 250k conversations for independent impact studies. What You Need to Know: Anthropic opened aggregated Claude conversation data from April-May 2026 to Stanford SALT Lab, Oxford, and METR, revealing over half of chats involve consequential tasks. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=fgEU1tPe_1A 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E153
    August 26 · 8 min

    Ep 153: OpenAI’s first custom inference chip is entering production, delivering higher throughput…

    Models & Agents OpenAI’s first custom inference chip is entering production, delivering higher throughput and lower latency in one architecture. What You Need to Know: OpenAI announced deployment plans for its Jalapeño inference chip by year-end, with testing showing gains in intelligence per watt and response speed for ChatGPT and agents. A new Qwen3.8-Flash-Next model is slated for release today, and Liquid AI open-sourced Pipette for on-device benchmarking. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=YmaSdsTMizY 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E152
    August 25 · 8 min

    Ep 152: Local 27B models just wrote and merged their first production feature on a single 4060 Ti.

    Models & Agents Local 27B models just wrote and merged their first production feature on a single 4060 Ti. What You Need to Know: A developer reported successfully using Qwen3 27B IQ3_K_XXS quantized to run entirely on a 4060 Ti 16GB card, completing a full agentic coding workflow including codebase investigation, plan generation, multi-file edits, and QA gate passing before human merge approval. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=BrVWqPmriKA 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E151
    August 24 · 10 min

    Ep 151: AI agents are shifting from experimental tools to major API consumers, changing how…

    Models & Agents AI agents are shifting from experimental tools to major API consumers, changing how developers price and secure their endpoints. What You Need to Know: PYMNTS reports agents now drive significant API traffic as businesses deploy them for routine transactions. Several arXiv papers detail concrete gains in reasoning speed, style control, and domain-specific deployment. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=bbMs8V4yB1A 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E150
    August 23 · 12 min

    Ep 150: Vercel and Ora just shipped a free public audit tool that scores any website’s readiness…

    Models & Agents Vercel and Ora just shipped a free public audit tool that scores any website’s readiness for AI agents across 118 checks. What You Need to Know: The biggest concrete release today is Vercel’s “Is Agentic” scorer, which lets developers quickly test whether their sites can support autonomous agents. A detailed deepDoctection tutorial shows how to wire layout analysis, DocTR OCR, and table extraction into structured JSONL for RAG. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=DCEocF4Bmaw 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E149
    August 22 · 9 min

    Ep 149: A 250M-parameter model trained on 30B tokens now deploys in 60 MB with million-token…

    Models & Agents A 250M-parameter model trained on 30B tokens now deploys in 60 MB with million-token retrieval from disk. What You Need to Know: A solo developer released SHADOW-250M, a heavily quantized LLM that keeps recent context in fp16 while compressing older tokens to 1 bit on disk. Nvidia published a linear-mapping technique that transfers KV caches between model sizes without full re-prefill. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=3NC6y8HexD4 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E148
    August 21 · 12 min

    Ep 148: Agent reliability benchmarks just exposed the gap between occasional success and…

    # Models & Agents Agent reliability benchmarks just exposed the gap between occasional success and consistent stateful execution in real business workflows. What You Need to Know: Thinkingbox introduces a sandbox and 507-workflow benchmark across retail, insurance, and IT support domains that measures end-to-end state transitions rather than isolated tool calls. Several arXiv papers released today examine attention allocation, KV-cache reuse, and multi-agent hypothesis generation. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=_nSKXMijex4 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E147
    August 20 · 9 min

    Ep 147: OpenAI is testing private safety processing that keeps frontier-model interactions off…

    # Models & Agents OpenAI is testing private safety processing that keeps frontier-model interactions off-limits to staff while still catching risks across long agent sessions. What You Need to Know: OpenAI previewed Private Safety Processing for frontier models to improve safety without personnel seeing raw content. Simon Willison documented an untrusted-sandbox experiment where Claude Code triggered an autonomous GitHub Actions push. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=b0WOMNQDb6k 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
Showing 21–40 of 56 episodes