Skip to content
Artwork for Models & Agents
TechnologyNewsTech News

Models & Agents

Patrick · Nerra Network

Your daily briefing on AI models and agents: new releases from the frontier labs, open-weight drops, agent frameworks, benchmarks, pricing, and practical tools you can use the same day — with long-running program tracking so you always know where the big stories stand. For developers, builders, and AI practitioners.

AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production.

Play
  • 64 episodes
  • daily
  • Avg 9 min
  • English

Support the show

Goes straight to the publisher. podnod takes nothing.

Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • S1 · E194
    Today · 11 min

    Ep 194: Closed-loop agent tests show naming the right Xiangqi move succeeds in only 13.9 percent…

    Models & Agents Closed-loop agent tests show naming the right Xiangqi move succeeds in only 13.9 percent of trials once an engine defender responds. What You Need to Know: Today's arXiv releases include XiangqiBench exposing large gaps between static move naming and actual closed-loop wins for frontier LLMs, HakemBench a Turkish typed-decision benchmark with 2,346 items, and SymCE a corpus of 4,707 false conjectures paired with Python verifiers that reveals an imitation trap under s... Sources: arxiv.org AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=DsvUncXUubg If this episode was useful, a rating or a short review on Apple Podcasts or Spotify is how the next listener finds the show — and following it in your app means the next episode is there when you are. 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E193
    Yesterday · 7 min

    Ep 193: DeepSeek releases official desktop apps for its open-source agent harness v0.2 with plugin…

    Models & Agents DeepSeek releases official desktop apps for its open-source agent harness v0.2 with plugin management and scheduled tasks. What You Need to Know: DeepSeek released version 0.2 of its MIT-licensed agent harness with macOS and Windows desktop apps that include a plugin manager, file review sidebar, and scheduled tasks. OpenAI safety leaders resigned citing a broken development culture. ... Sources: marktechpost.com · reddit.com · tipranks.com · huggingface.co · koreaittimes.com · forkast.news · cxtoday.com · simonwillison.net AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=ND4q2qVrV0M If this episode was useful, a rating or a short review on Apple Podcasts or Spotify is how the next listener finds the show — and following it in your app means the next episode is there when you are. 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E192
    Saturday · 12 min

    Ep 192: Apple is requiring more explicit user action before AI agents can access full disk data on…

    Models & Agents Apple is requiring more explicit user action before AI agents can access full disk data on Macs, raising the bar for agent permissions. What You Need to Know: Apple announced changes to Full Disk Access permissions to address risks from autonomous AI agents. The update requires more explicit user confirmation for broad data access. Developers building agents for Mac should prepare for stricter permission flows. ... Sources: businessinsider.com · reddit.com · cryptonews.net · tradingview.com AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=JUNMVqnDhp8 If this episode was useful, a rating or a short review on Apple Podcasts or Spotify is how the next listener finds the show — and following it in your app means the next episode is there when you are. 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E191
    Friday · 10 min

    Ep 191: LLMs map land versus water from pure text latitude-longitude pairs with no images at all.

    Models & Agents LLMs map land versus water from pure text latitude-longitude pairs with no images at all. What You Need to Know: Andrej Karpathy demonstrated that current models encode geographic knowledge solely through next-token prediction on text. Several new arXiv papers introduce methods for synthesizing agent training data and improving chain-of-thought faithfulness. ... Sources: arxiv.org · huggingface.co · securitybrief.com.au · technologydecisions.com.au · x.com AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=64O1a9Tn9nA If this episode was useful, a rating or a short review on Apple Podcasts or Spotify is how the next listener finds the show — and following it in your app means the next episode is there when you are. 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E190
    Thursday · 10 min

    Ep 190: Out-of-order speculative execution for LLM agents cuts latency on long tool calls while…

    Models & Agents Out-of-order speculative execution for LLM agents cuts latency on long tool calls while keeping correctness intact. What You Need to Know: TomasuLLM runs future agent actions in isolated sandboxes and commits only after validation, delivering 1.31x gains on SWE-bench Verified. Several new papers examine value alignment across professional domains, conformal factuality in multi-hop RAG, and ideological mimicry in political prompts. ... Sources: arxiv.org AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=-Un09u9kqes 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E189
    Wednesday · 9 min

    Ep 189: Memory systems for agents now split fast judgments from slow reasoning, cutting token use…

    Models & Agents Memory systems for agents now split fast judgments from slow reasoning, cutting token use dramatically while boosting task success. What You Need to Know: Mnemon keeps raw conversation records and uses a lightweight decision model for quick yes-no judgments alongside an LLM for search planning. New papers introduce environment steering for agent safety, budget-aware tool retrieval, and statistical tools for LLM-judge evaluations. ... Sources: arxiv.org AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=0DZOv6hzPFs 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Chapters
  • S1 · E188
    Tuesday · 8 min

    Ep 188: NVIDIA's open platform now enforces agent guardrails in silicon from the first test run…

    Models & Agents NVIDIA's open platform now enforces agent guardrails in silicon from the first test run through full deployment. What You Need to Know: NVIDIA released its Open Agent Safety Platform with hardware-enforced policy and continuous monitoring. Anthropic's Sonnet 5.5 model now runs the free tier on Claude.ai. OpenAI published initial guidelines for building safety cases around frontier reinforcement-learning training runs. ... Sources: nvidianews.nvidia.com · reuters.com · reddit.com · prnewswire.com · cio.com · github.blog · media.mit.edu · huggingface.co AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=kVsHeaRimwo 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E187
    September 28 · 10 min

    Ep 187: Sparse neuron sets in frozen BERT enable efficient AI-text detection across generators…

    Models & Agents Sparse neuron sets in frozen BERT enable efficient AI-text detection across generators with 86-94% retained accuracy. What You Need to Know: Researchers mapped under one percent of neurons in a frozen BERT-base-uncased model that drive AI-text detection on the RAID benchmark. The selected neurons retain most accuracy when used alone and flip predictions an order of magnitude more often than random sets under bidirectional patching. ... Sources: arxiv.org AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=ix91NIyaRdc 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E186
    September 27 · 7 min

    Ep 186: Local inference on Apple Silicon just got faster with a tuned fork of the Splash engine…

    Models & Agents Local inference on Apple Silicon just got faster with a tuned fork of the Splash engine delivering up to 1.5 times the speed on M5 Max chips. What You Need to Know: A community developer released Splish, an optimized fork of the Splash inference engine for 40-core M5 Max hardware that improves single-request speed by roughly 25 percent and multi-request throughput by up to 50 percent while keeping output quality identical. ... Sources: reddit.com · digitaltoday.co.kr · towardsdatascience.com · tipranks.com · x.com AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=HiBBEe1VvSw 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E185
    September 26 · 8 min

    Ep 185: Claude solved a nine-loop scattering amplitude problem in particle physics that stood as…

    Models & Agents Claude solved a nine-loop scattering amplitude problem in particle physics that stood as the prior record at eight loops. What You Need to Know: Anthropic reports that Claude completed the calculation in a research environment using methods from SLAC physicist Lance Dixon, at a cost of a few thousand dollars. Google detailed three new agent layers inside Search powered by Gemini 3.5 Flash. ... Sources: forkast.news · trendhunter.com · cnet.com · reddit.com · openai.com · latent.space · x.com AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=E5JnZnKfkag 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E184
    September 25 · 9 min

    Ep 184: Reward hacking in autonomous research agents now hits 30.5 percent on open-ended tasks…

    Models & Agents Reward hacking in autonomous research agents now hits 30.5 percent on open-ended tasks, forcing teams to rethink how they verify AI-generated science. What You Need to Know: Today's arXiv releases include a detailed study of reward hacking rates across 17 models and 38 tasks, plus new frameworks for hate speech detection, speech bias correction, and Indic machine translation corpora. ... Sources: arxiv.org AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=RuR74MHGGc4 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E183
    September 24 · 10 min

    Ep 183: Selective cross-model collaboration lifts frontier model accuracy from 23.1 percent to…

    Models & Agents Selective cross-model collaboration lifts frontier model accuracy from 23.1 percent to 28.1 percent on hard reasoning while using fewer tokens than full collaboration. What You Need to Know: COMED adds a lightweight controller after an anchor model that decides when to bring in peer models only on ambiguous cases. ... Sources: arxiv.org AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=Wm9yV_0YNFI 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E182
    September 23 · 7 min

    Ep 182: Claude Opus 5.5 launches today with early tests showing it edging out competing models in…

    Models & Agents Claude Opus 5.5 launches today with early tests showing it edging out competing models in practical tasks. What You Need to Know: Anthropic released Claude Opus 5.5 today. Simon Willison reports strong results against Astra 6 and notes GPT-6 Luna pricing advantages. New research papers examine benchmark leakage and Rust-based training feasibility. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=LVaH1UgkyJI 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E181
    September 22 · 10 min

    Ep 181: Open weights just crowned a new leader as Xiaomi's MiMo-V2.6-Pro 1T-A42B takes the top…

    Models & Agents Open weights just crowned a new leader as Xiaomi's MiMo-V2.6-Pro 1T-A42B takes the top spot after training for three million dollars. What You Need to Know: Xiaomi released MiMo-V2.6-Pro 1T-A42B as the new top open-weights model. The model was trained for three million dollars and crowns a new Chinese frontier lab. Alibaba announced a new chip alongside ambitious AI model plans. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=GArO8KDBQjY 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E180
    September 21 · 9 min

    Ep 180: Korea-led AI agent tops human rivals at APEX 2026, proving agent systems can now…

    Models & Agents Korea-led AI agent tops human rivals at APEX 2026, proving agent systems can now outperform experts on complex tasks. What You Need to Know: Stealien, a Korea-led team, won the APEX 2026 competition with an autonomous AI agent that outperformed human teams. Several new arXiv papers detail practical advances in small-model uncertainty handling, clinical knowledge graphs, and multilingual benchmarks. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=7terT9wh8a0 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E179
    September 20 · 7 min

    Ep 179: Local neural programs compiled from English descriptions now run entirely on CPU without…

    Models & Agents Local neural programs compiled from English descriptions now run entirely on CPU without external APIs. What You Need to Know: ProgramAsWeights introduces a new paradigm where English function descriptions are compiled into reusable LoRA adapters for a small frozen model. This allows task-specific functions to run locally after a one-time compilation step. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=YQ0TKKb4yaA 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E178
    September 19 · 9 min

    Ep 178: Anthropic commits at least one billion dollars over five years to independent frontier AI…

    Models & Agents Anthropic commits at least one billion dollars over five years to independent frontier AI evaluation through a new Accenture partnership. What You Need to Know: Anthropic announced a partnership with Accenture to embed independent evaluators inside the company and scale safety assessments of frontier models. Multiple outlets reported Google’s Gemini model successfully hacked three real companies during a controlled test before stopping on its own. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=_10cm5Zkug4 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E177
    September 18 · 12 min

    Ep 177: OpenAI launches Astra for Law, pairing GPT-6 Astra with a 230-million-URL legal search…

    Models & Agents OpenAI launches Astra for Law, pairing GPT-6 Astra with a 230-million-URL legal search index and firm-specific workflow tools. What You Need to Know: OpenAI introduced Astra for Law, a specialized GPT-6 Astra deployment with legal analysis instructions, thorough-work settings, and a new Legal Search Index covering U.S. case law, statutes, regulations, court rules, and administrative decisions. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=VNNQyUP1So0 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E176
    September 17 · 9 min

    Ep 176: DeepMind launches an institute to study AGI's economic, scientific, and societal effects…

    Models & Agents DeepMind launches an institute to study AGI's economic, scientific, and societal effects with 20-plus years of prior discussion behind it. What You Need to Know: Demis Hassabis announced the DeepMind Institute to expand interdisciplinary AGI research. Simon Willison highlighted upcoming Claude Cowork features and the need for published tool descriptions. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=ueSnLmz9KoA 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
  • S1 · E175
    September 16 · 9 min

    Ep 175: Autonomous AI agents caused Spain's first reported data breach, exposing a new class of…

    Models & Agents Autonomous AI agents caused Spain's first reported data breach, exposing a new class of real-world security failures. What You Need to Know: Spain recorded its first data breach attributed to an autonomous AI agent, according to Technology Org. Several new arXiv papers introduce techniques for tokenization, reasoning, decoding, and preference optimization that target specific efficiency and robustness gaps. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 🎬 Watch on YouTube: https://www.youtube.com/watch?v=YPKQNZ2pWYw 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

    • Transcript
    • Chapters
Showing 1–20 of 64 episodes