Skip to content
Artwork for Nerd Snipe with Theo and Ben

Nerd Snipe with Theo and Ben

Theo and Ben

#48 in Technology this week

"The only dev podcast hosted by real devs" - Theo (not a real dev)

Play
  • 24 episodes
  • weekly
  • Avg 1 hr 38 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • S1 · E24
    September 30 · 2 hr 8 min

    Opus 5.5 is the best model ever made, Grok 4.7 exists, GPT 6 Sol & Luna Impress, and why Jev isn't the solution to everything

    Theo and Ben break down a packed week of AI releases, from Opus 5.5's insane release to Grok 4.7 falling short of the hype. They then dig into why post-training matters, where Jev, GPT-6 Sol, and Luna fit, and more! Thank you to General Translation & Paper for Sponsoring! General Translation: ⁠https://nerdsnipe.link/gt Paper: https://nerdsnipe.link/paper Sources available on our Substack: https://nerdsnipe.substack.com/ Listen wherever you get your podcasts: Spotify: https://nerdsnipe.link/spotify Apple: https://nerdsnipe.link/apple Elsewhere: https://nerdsnipe.link/listen Timestamps: 0:00 Intro 3:18 GPT-6 Sol & Luna 14:58 Grok 4.7 38:47 Jev explained 52:48 Why model routing fails 1:05:06 Opus 5.5 1:26:17 Pricing & usage limits 1:34:23 Building games with Opus 1:41:10 Porting TypeScript to Rust 1:48:57 Agent workflows & testing

  • S1 · E23
    September 19 · 2 hr 32 min

    Why Pacing AI Won't Stop the Apocalypse, Apple Ruins UI/UX with the IPhone Duo, and How Unlimited Tokens Changes Basically Nothing

    Theo & Ben break down Jacob Coxon's resignation after three years in pretraining at OpenAI and Anthropic—and his warning that both labs are racing toward AGI and gambling with our lives—along with Dario's latest blog post about pacing the frontier and how Apple's iPhone Duo could have ruined app development (if AI didn't exist). Thanks to this episode's sponsor, PostHog: PostHog: https://nerdsnipe.link/posthog Listen wherever you get your podcasts: Spotify: https://nerdsnipe.link/spotify Apple: https://nerdsnipe.link/apple Elsewhere: https://nerdsnipe.link/listen Sources available on our Substack: https://nerdsnipe.substack.com/ Timestamps: 0:00 Intro 2:53 New iPhones 22:24 Apple and Carriers 32:38 Astra’s Reliability 47:55 Codex Usage and Caching 1:19:44 Anthropic Resignation 1:38:38 Pacing the Frontier 1:40:02 Embedded AI Evaluators 1:43:10 The Hugging Face Hack 2:02:29 React Native Debate 2:09:18 Fast Inference Tradeoffs 2:12:40 Democratic AI Coordination 2:18:22 China and Global Pacing

  • S1 · E22
    September 10 · 2 hr 5 min

    Anthropic's Answer to Astra, Gemini 3.8 Flash Killed Benchmarks, and Muse Spark 1.3's Pretty Good

    Theo & Ben break down their latest thoughts on how Astra and Fable 5.1's releases have changed the way they work, how they interact in T3 Code, and then circle back to the latest with Codex usage limits, Gemini 3.8 Flash, Muse Spark 1.3, and the growing gap between model benchmarks and real coding workflows. Thanks to this episode's sponsor, General Translation: General Translation: https://nerdsnipe.link/gt Listen wherever you get your podcasts: Spotify: https://nerdsnipe.link/spotify Apple: https://nerdsnipe.link/apple Elsewhere: https://nerdsnipe.link/listen Sources available on our Substack: https://nerdsnipe.substack.com/ Timestamps: 00:00 Intro 04:57 Gemini 3.8 Flash 14:48 Muse Spark & small models 22:08 AI subscriptions 34:14 GPT-6 Astra launch & pricing 56:04 Rate limits & autonomous coding 1:10:14 Fable 5.1 1:23:23 AI design & 3D demos 1:40:53 Astra vs. Fable: Which to use?

  • S1 · E21
    September 3 · 1 hr 51 min

    We've Been Using GPT-6 Astra for a Few Weeks, Here's What We Think...

    Theo & Ben break down OpenAI's latest model, Astra, and why it's their new benchmark for AI coding, computer use, and multimodal work. But it's not all good: they explain why its UI, stopping behavior, and agent reliability still create friction in real software workflows. From DEFCON puzzles to coding-agent PRs, we're comparing Astra with Fable and asking what a trustworthy OpenAI model should do next. Thank you to PostHog for sponsoring today's episode! PostHog, all-in-one suite of product tools: https://nerdsnipe.link/posthog Listen wherever you get your podcasts: Spotify: https://nerdsnipe.link/spotify Apple: https://nerdsnipe.link/apple Elsewhere: https://nerdsnipe.link/listen Sources available on our Substack: https://nerdsnipe.substack.com/ Timestamps 00:00 Meet Astra 04:03 Best Model Ever, With Catches 06:33 Reasoning and 3D Benchmarks 20:00 Astra Rebuilds Ping.gg 30:04 Instruction-Following Problems 44:47 The Uncommitted Fix Debate 01:09:37 The PR Babysitting Failure 01:22:41 Why Fable Still Wins 01:30:41 Multimodal and Computer Use 01:39:29 Astra vs. Fable Fleet Data

  • S1 · E20
    August 28 · 2 hr 42 min

    Ox Alpha Revealed, OpenAI's Latest Pricing Updates, and Our Coding Model Tier List

    Theo & Ben break down OpenAI's GPT-5.6 Sol price cut, breakdown the "Ox Alpha" stealth model we now know is GLM 5.3 Flash, then take a drink every time they say "Grok" while ranking every current AI model on a tier list! What could go wrong? Thanks to this episode's sponsor, General Translation: General Translation: https://nerdsnipe.link/gt Listen wherever you get your podcasts: Spotify: https://nerdsnipe.link/spotify Apple: https://nerdsnipe.link/apple Elsewhere: https://nerdsnipe.link/listen Sources available on our Substack: https://nerdsnipe.substack.com/ Timestamps 0:00 Intro 3:30 OpenAI vs. Anthropic 14:44 Anthropic’s delayed models 20:53 Kimi K3 and open weights 30:08 The Alpha stealth model 48:26 Model tier list begins 1:09:53 GPT-5.6 Luna 1:13:09 Opus and Sonnet 1:20:25 Gemini models 1:31:15 DeepSeek and local models 1:59:58 Fable vs. Sol 2:28:46 Final rankings

  • S1 · E19
    August 20 · 2 hr 18 min

    Anthropic Doesn't Think You Can Be Trusted, China is closing the gap, SpaceXAI Leaps Ahead, and DEF CON

    Anthropic started watermarking their model's outputs, Dario posted on X, SpaceXAi is getting really good really fast, Meta is shipping again, and we won DEF CON 2026. Thank you PostHog, the all in one suite of product tools for sponsoring today's Episode! Check them out at: nerdsnipe.link/posthog Listen wherever you get your podcasts: - Spotify: nerdsnipe.link/spotify - Apple: nerdsnipe.link/apple - Elsewhere: nerdsnipe.link/listen Sources available on our Substack: nerdsnipe.substack.com Timestamps: 0:00 Intro 2:56 Claude Watermarks 15:04 Meta + Muse 48:48 GLM-5.3 55:15 Qwen + DeepSeek 1:08:48 Grok 4.6 + Bot 1:20:15 Gavin vs Dario 1:43:44 DEFCON 2:15:35 Viewer Q&A

  • S1 · E18
    July 28 · 1 hr 23 min

    Opus 5 Releases, China Catches Up, and the OpenAI Model Sandbox Escape

    Theo and Ben break down why Opus 5 feels like GPT-5.5 crossed with Fable rather than GPT-5.6 crossed with Fable, what Fable found when it audited an Opus thread line by line, where Opus still clearly wins (3D, animation, and Claude Code limits at half Fable's price with no 50% weekly cap), plus Kimi K3, GLM 5.2, the Hugging Face hack, Grok 4.5, Codex, and T3 Code. Thanks to this episode's sponsor, General Translation: General Translation: https://nerdsnipe.link/gt Sources available on our Substack: https://nerdsnipe.substack.com/

  • S1 · E17
    July 14 · 2 hr 6 min

    5 different models dropped last week & the GPT-5.6 usage limits are brutal

    This week Grok 4.5, Muse Spark 1.1, GPT-Live, and GPT 5.6 all dropped. ⁠Theo⁠⁠⁠ and ⁠⁠Ben⁠⁠ break down which models you should care about, how OpenAI fumbled the Codex to ChatGPT app transition, the abysmal usage numbers for 5.6, and the latest drama surrounding OpenAI and Sam Altman. Plus, how many satellites would it take to trap humanity on Earth, and how many data centers to heat up the ocean? Thank you to this episode's sponsors: Composio & WorkOs. https://nerdsnipe.link/composio https://nerdsnipe.link/workos Sources available on Substack: https://nerdsnipe.substack.com/

  • S1 · E16
    July 9 · 1 hr 11 min

    We Tested GPT 5.6 Sol Early

    We've spent six figures in tokens testing OpenAI's 5.6 Sol model to see whether if its better than Fable and what OpenAI have done to make it even better than 5.5. Also, we breakdown why we both moved our agents to Linux boxes, how to actually burn $65k on a single loop, and the Codex vs Claude Code subagent gap that's now bigger than the model gap itself. Thanks to this episode's sponsors Clerk and General Translation: Clerk, the auth platform with the best DX: https://nerdsnipe.link/clerk General Translation: https://nerdsnipe.link/gt Sources Available on our Substack: https://nerdsnipe.substack.com/

  • S1 · E15
    July 8 · 1 hr 16 min

    Fable Is Back...kinda?

    Fable was supposed to give us 14 days and instead we got 3 before the export ban. Now it's back at half the rate limits. Also, we break down why AI-generated Claude Code skills fall apart, what actually triggers Fable-to-Opus rerouting, the Mythos export-control timeline, and whether GLM 5.2 can really touch Sonnet 5. Thank you to PostHog for sponsoring today's episode! PostHog, all in one suite of product tools: ⁠https://nerdsnipe.link/posthog Sources available on our Substack: https://nerdsnipe.substack.com/

  • S1 · E14
    June 30 · 1 hr 11 min

    GPT-5.6 is here! And none of us can use it.

    A new "government-approved rollout" for GPT-5.6 is coming, and it's got us wondering: are we entering the next dark age of AI model access? Additionally, we break down the launch of T3 Code x Grok CLI, Apple's price hikes, the memory supply crunch driving RAM and SSD costs up, a repo-poisoning and AI PR spam wave hitting open source, the GPT-5.6 non-launch, and what frontier model access looks like in a Fable/Mythos-tier world. Thank you to Composio for sponsoring this episode! Composio, connect your agents to everything: https://nerdsnipe.link/composio Sources Available on our Substack: https://nerdsnipe.substack.com/

  • S1 · E13
    June 24 · 1 hr 33 min

    The US Government Banned Claude Fable 5...

    The US government banned Anthropic's Mythos and Fable models just after launch so we break down exactly how Project Glass Wing, a panicked AWS engineer, and Dario's failure to communicate with Washington triggered the chaos. Plus: SpaceX's $60B Cursor acquisition, GLM 5.2 beating Google, the Codex trick for running 200 parallel agents, and why your next GPU will ship with a GPS tracker Thank you to Composio and WorkOS for sponsoring this episode! Composio, connect your agents to everything: https://nerdsnipe.link/composio WorkOS, your enterprise ready solution: https://nerdsnipe.link/workos Sources Available on our Substack: https://nerdsnipe.substack.com/

  • S1 · E12
    June 15 · 1 hr 17 min

    Our impressions of Claude Fable/Mythos (we filmed this before the ban)

    RIP Fable 5. We recorded this before it got taken offline, but it's still worth talking about. The model is incredible. We really miss it. Thank you, Firecrawl, Depot, and Clerk for sponsoring! Firecrawl, the best api for searching and crawling the web: nerdsnipe.link/firecrawl Depot, better CI in every way: nerdsnipe.link/depot Clerk, the best dx in auth: nerdsnipe.link/clerk Sources https://x.com/thsottiaux/status/2043177597434306699 https://cognition.ai/blog/frontier-code https://x.com/paradite_/status/2064585901351792887 https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf (page 13 has the “prompt modification” quote) https://www.anthropic.com/institute/recursive-self-improvement Timestamps 0:00 Intro 3:29 Fable First Impressions 10:24 Benchmark Drama 26:42 Claude Code & Workflows 36:07 Pricing & June 22 Cutoff 39:54 Data Retention 45:58 Hidden Safeguards 1:07:01 The Claude Constitution

  • S1 · E11
    June 10 · 1 hr 35 min

    Now even Google's buying GPUs from SpaceX?

    Cloudflare buys Void0, Google buying up compute from xAI, and Claude seems to be getting more anxious so we're here to break down everything this week on another episode of Nerd Snipe! Thank you to Composio for sponsoring today's episode! Composio, connect your agents to everything: https://nerdsnipe.link/composio Sources: https://x.com/vite_js/status/2062525206158078047 https://x.com/EdLudlow/status/2062970770612199542 https://x.com/nrehiew_/status/2063099050719846719 https://www.reddit.com/r/EconomyCharts/comments/1lp34n4/china_vs_us_energy/ https://x.com/elonmusk/status/1963443919150330139 https://x.com/AnthropicAI/status/2062568862479208923 https://x.com/HSVSphere/status/2060396271756595666 https://x.com/theo/status/2061018426152530232 https://x.com/Teknium/status/2062522290504613944 01:58 Cloudflare vs Vercel 13:03 Convex angle 19:59 SpaceX compute 35:49 AI self-improvement 44:45 Article reactions 50:28 Claude anxiety 57:42 Guardrails 01:08:21 Cursed image gen 01:19:18 Hermes agents

  • S1 · E10
    June 3 · 1 hr 15 min

    We (mostly) like Claude Opus 4.8

    Opus 4.8 + a ton of new stuff in Claude Code dropped this week, and we actually kinda like it. There's also a new benchmark that's actually good, and we have a lot of thoughts about the future of these AI labs... Thank you PostHog and Clerk for sponsoring! PostHog, the all in one suite of product tools: nerdsnipe.link/posthog Clerk, the best dx in auth: nerdsnipe.link/clerk SOURCES https://www.anthropic.com/news/claude-opus-4-8 https://x.com/theo/status/2060120708815139241 https://x.com/Baconbrix/status/2060065875911422343 ⁠https://x.com/zoink/status/2060769829133721974 https://x.com/maria_rcks/status/2060937270824153251 https://x.com/_catwu/status/2060054180379689074 https://x.com/datacurve/status/2060834005998793199 https://deepswe.datacurve.ai/ https://deepswe.datacurve.ai/blog#results https://github.com/scaleapi/SWE-agent/blob/402a7b8fdac8193f3f255bb53859ba274234f596/config/benchmarks/anthropic_filemap_multilingual.yaml https://deepswe.datacurve.ai/ https://x.com/AnthropicAI/status/2060061347522433422 https://x.com/theo/status/2060199299632472494 https://x.com/theo/status/2060901326058561795 https://x.com/38twelveDaily/status/2060408696945975631 00:00 - Shower Thoughts 02:44 - Deep SWE Benchmark 10:45 - Opus vs GPT-5.5 19:57 - Anthropic’s Huge Raise 25:39 - Token Maxing 40:02 - AI Slot Machine 43:49 - Claude Code Friction 50:01 - Opus, Mythos, and Safety

  • S1 · E9
    May 28 · 1 hr 22 min

    Google is Not a Serious Company

    Not only did Google accidentally ban Railway's account, but their new flagship model Gemini 3.5 Flash is absurdly bad. Oh and apparently Theo's building his own cloud... Thank you Macroscope and GT for sponsoring today's episode! Macroscope: ⁠nerdsnipe.link/macroscope General Translation: nerdsnipe.link/gt SOURCES ⁠https://x.com/theo/status/2057359424378097823⁠ ⁠https://x.com/JustJake/status/2056881510939283776⁠ ⁠https://x.com/KorduGG/status/2059141337895604626⁠ ⁠https://x.com/unboringtech/status/2059144145273491610⁠ TIMESTAMPS 00:00 - Gemini fallout 05:21 - Video gen 10:10 - Google Cloud 14:04 - Windsurf/Cursor 25:10 - Manus/Meta 30:00 - China lock-in 35:00 - Cursor/xAI 45:00 - Cloud workflows 50:32 - Lakebed 65:22 - Hermes/security

  • S1 · E8
    May 20 · 1 hr 9 min

    How the OpenClaw creator uses $1.3 million of tokens

    Peter, the creator of openclaw is apparently going through $1,300,000 worth of tokens every month. Seems like we're not using nearly enough tokens. Oh and the security psychosis is getting worse. (Anthropic is being bad again too) Thank you to today's sponsors! - AgentMail, email inboxes for your agents: nerdsnipe.link/agentmail - Clerk, the best experience for auth, orgs, billing, and more: nerdsnipe.link/clerk TIMESTAMPS 00:00 - Intro 02:44 - Token Spend 06:47 - Token Future 12:12 - Secure Agents 22:50 - Anthropic Rules 31:35 - Security 39:36 - Token Tax 49:49 - AI Psychosis 59:47 - macOS 01:01:43 - AI Video

  • S1 · E7
    May 14 · 1 hr 36 min

    Anthropic solved their compute problem by buying it from Elon?

    Anthropic seems to have finally solved their compute problems (kinda) by buying it from Elon, the security problem is getting so much worse, and apparently Bun's getting re-written in rust? Thank you to PostHog and Composio for sponsoring today's episode! - PostHog, the all in one suite of product tools: nerdsnipe.link/posthog - Composio, connect your agents to everything: nerdsnipe.link/composio Sources/references: https://x.com/claudeai/status/2052060691893227611 https://openai.com/index/elon-musk-wanted-an-openai-for-profit/#december-2018-elon-told-us-to-raise-billions-per-year-immediately-or-forget-it https://x.com/jarredsumner/status/2053391824702898475 https://x.com/jarredsumner/status/2051595933704761618 https://x.com/jarredsumner/status/2053047748191232310 https://ze3tar.github.io/post-zcrx.html https://www.jefftk.com/p/ai-is-breaking-two-vulnerability-cultures https://x.com/thdxr/status/2053570581807722968 https://x.com/badlogicgames/status/2052691176373805534 https://x.com/garrytan/status/2052996691586932783 TIMESTAMPS 00:00:00 - Anthropic/xAI00:08:59 - OpenAI Lawsuit00:25:40 - Bun in Rust00:34:34 - Security00:59:28 - Pottery Coding01:11:11 - Local Models

  • S1 · E6
    May 6 · 1 hr 47 min

    Theo Almost Lost $1 Million

    This week Theo nearly lost a million dollars and Ben got AI psychosis (from gstack)... Thank you to Coderabbit and Clerk for sponsoring today's episode! - Coderabbit, the ultimate AI code reviewer: ⁠nerdsnipe.link/coderabbit - Clerk, the auth platform with the best DX: nerdsnipe.link/clerk Sources/references: - https://x.com/theo/status/2014863266888233193 - https://x.com/theo/status/2050305813894648289 - https://x.com/theo/status/2050314995561611357 - https://x.com/sama/status/2050671161915371998 - https://x.com/davis7/status/2050718508372431026 - https://x.com/thdxr/status/2050719575033983323 - https://x.com/davis7/status/2050762009592148375 - https://x.com/naval/status/2050560057675522500 - https://x.com/theo/status/1952229335416623592 00:00 - Intro / studio return 00:52 - Azure $1M / Microsoft 11:22 - Cloud platform talk 18:08 - Coding agents / SDKs 29:06 - GPT-5.5 pricing debate 45:14 - OpenClaw workflows 01:12:37 - G Stack / G Brain01:29:00 - Dynamic UI / wrap-up

  • S1 · E5
    May 1 · 1 hr 35 min

    We need to talk about OpenAI

    OpenAI and Microsoft are breaking up, Sam's drunk posting, Anthropic is being stupid again, and we still disagree about GPT-5.5Thanks to this episode's sponsors: - Clerk, the auth platform with the best DX: https://nerdsnipe.link/clerk- Coderabbit, the ultimate AI code reviewer: ⁠https://nerdsnipe.link/coderabbit- PlanetScale, the fastest and most scalable cloud databases: https://nerdsnipe.link/planetscale 00:00 Intro 04:05 Sam drunk tweets1 3:13 Anthropic billing woes 28:27 OpenAI divorce 42:34 GitHub can't stop dying 59:40 GPT-5.5 retrospective Sources/references: https://x.com/sama/status/2046808114561974567 https://x.com/sama/status/2046808217133670800 https://x.com/sama/status/2048160404376105179 https://x.com/sama/status/2047403771416940715 https://x.com/om_patel5/status/2048204411986469232 https://x.com/MSFTnews/status/2048749108127506936 https://x.com/GergelyOrosz/status/2048834949667537369 https://x.com/theo/status/2047721472521621991 https://x.com/mitchellh/status/2049213597419774026 https://x.com/kdaigle/status/2047803291988590609 https://x.com/ryanflorence/status/2048538797638599109 https://x.com/davis7/status/2048239401059434710 https://x.com/davis7/status/2048077518725366173 https://x.com/badlogicgames/status/2048444292562026713 https://x.com/0xSero/status/2048744545853030690

Showing 1–20 of 24 episodes