
Nerd Snipe with Theo and Ben
Theo and Ben
#48 in Technology this week"The only dev podcast hosted by real devs" - Theo (not a real dev)
- 24 episodes
- weekly
- Avg 1 hr 38 min
- English
S1 · E24September 30 · 2 hr 8 min- NS
S1 · E23September 19 · 2 hr 32 minWhy Pacing AI Won't Stop the Apocalypse, Apple Ruins UI/UX with the IPhone Duo, and How Unlimited Tokens Changes Basically Nothing
Theo & Ben break down Jacob Coxon's resignation after three years in pretraining at OpenAI and Anthropic—and his warning that both labs are racing toward AGI and gambling with our lives—along with Dario's latest blog post about pacing the frontier and how Apple's iPhone Duo could have ruined app development (if AI didn't exist). Thanks to this episode's sponsor, PostHog: PostHog: https://nerdsnipe.link/posthog Listen wherever you get your podcasts: Spotify: https://nerdsnipe.link/spotify Apple: https://nerdsnipe.link/apple Elsewhere: https://nerdsnipe.link/listen Sources available on our Substack: https://nerdsnipe.substack.com/ Timestamps: 0:00 Intro 2:53 New iPhones 22:24 Apple and Carriers 32:38 Astra’s Reliability 47:55 Codex Usage and Caching 1:19:44 Anthropic Resignation 1:38:38 Pacing the Frontier 1:40:02 Embedded AI Evaluators 1:43:10 The Hugging Face Hack 2:02:29 React Native Debate 2:09:18 Fast Inference Tradeoffs 2:12:40 Democratic AI Coordination 2:18:22 China and Global Pacing
- NS
S1 · E22September 10 · 2 hr 5 minAnthropic's Answer to Astra, Gemini 3.8 Flash Killed Benchmarks, and Muse Spark 1.3's Pretty Good
Theo & Ben break down their latest thoughts on how Astra and Fable 5.1's releases have changed the way they work, how they interact in T3 Code, and then circle back to the latest with Codex usage limits, Gemini 3.8 Flash, Muse Spark 1.3, and the growing gap between model benchmarks and real coding workflows. Thanks to this episode's sponsor, General Translation: General Translation: https://nerdsnipe.link/gt Listen wherever you get your podcasts: Spotify: https://nerdsnipe.link/spotify Apple: https://nerdsnipe.link/apple Elsewhere: https://nerdsnipe.link/listen Sources available on our Substack: https://nerdsnipe.substack.com/ Timestamps: 00:00 Intro 04:57 Gemini 3.8 Flash 14:48 Muse Spark & small models 22:08 AI subscriptions 34:14 GPT-6 Astra launch & pricing 56:04 Rate limits & autonomous coding 1:10:14 Fable 5.1 1:23:23 AI design & 3D demos 1:40:53 Astra vs. Fable: Which to use?
- NS
S1 · E21September 3 · 1 hr 51 minWe've Been Using GPT-6 Astra for a Few Weeks, Here's What We Think...
Theo & Ben break down OpenAI's latest model, Astra, and why it's their new benchmark for AI coding, computer use, and multimodal work. But it's not all good: they explain why its UI, stopping behavior, and agent reliability still create friction in real software workflows. From DEFCON puzzles to coding-agent PRs, we're comparing Astra with Fable and asking what a trustworthy OpenAI model should do next. Thank you to PostHog for sponsoring today's episode! PostHog, all-in-one suite of product tools: https://nerdsnipe.link/posthog Listen wherever you get your podcasts: Spotify: https://nerdsnipe.link/spotify Apple: https://nerdsnipe.link/apple Elsewhere: https://nerdsnipe.link/listen Sources available on our Substack: https://nerdsnipe.substack.com/ Timestamps 00:00 Meet Astra 04:03 Best Model Ever, With Catches 06:33 Reasoning and 3D Benchmarks 20:00 Astra Rebuilds Ping.gg 30:04 Instruction-Following Problems 44:47 The Uncommitted Fix Debate 01:09:37 The PR Babysitting Failure 01:22:41 Why Fable Still Wins 01:30:41 Multimodal and Computer Use 01:39:29 Astra vs. Fable Fleet Data
- NS
S1 · E20August 28 · 2 hr 42 minOx Alpha Revealed, OpenAI's Latest Pricing Updates, and Our Coding Model Tier List
Theo & Ben break down OpenAI's GPT-5.6 Sol price cut, breakdown the "Ox Alpha" stealth model we now know is GLM 5.3 Flash, then take a drink every time they say "Grok" while ranking every current AI model on a tier list! What could go wrong? Thanks to this episode's sponsor, General Translation: General Translation: https://nerdsnipe.link/gt Listen wherever you get your podcasts: Spotify: https://nerdsnipe.link/spotify Apple: https://nerdsnipe.link/apple Elsewhere: https://nerdsnipe.link/listen Sources available on our Substack: https://nerdsnipe.substack.com/ Timestamps 0:00 Intro 3:30 OpenAI vs. Anthropic 14:44 Anthropic’s delayed models 20:53 Kimi K3 and open weights 30:08 The Alpha stealth model 48:26 Model tier list begins 1:09:53 GPT-5.6 Luna 1:13:09 Opus and Sonnet 1:20:25 Gemini models 1:31:15 DeepSeek and local models 1:59:58 Fable vs. Sol 2:28:46 Final rankings
- NS
S1 · E19August 20 · 2 hr 18 minAnthropic Doesn't Think You Can Be Trusted, China is closing the gap, SpaceXAI Leaps Ahead, and DEF CON
Anthropic started watermarking their model's outputs, Dario posted on X, SpaceXAi is getting really good really fast, Meta is shipping again, and we won DEF CON 2026. Thank you PostHog, the all in one suite of product tools for sponsoring today's Episode! Check them out at: nerdsnipe.link/posthog Listen wherever you get your podcasts: - Spotify: nerdsnipe.link/spotify - Apple: nerdsnipe.link/apple - Elsewhere: nerdsnipe.link/listen Sources available on our Substack: nerdsnipe.substack.com Timestamps: 0:00 Intro 2:56 Claude Watermarks 15:04 Meta + Muse 48:48 GLM-5.3 55:15 Qwen + DeepSeek 1:08:48 Grok 4.6 + Bot 1:20:15 Gavin vs Dario 1:43:44 DEFCON 2:15:35 Viewer Q&A
- NS
S1 · E18July 28 · 1 hr 23 minOpus 5 Releases, China Catches Up, and the OpenAI Model Sandbox Escape
Theo and Ben break down why Opus 5 feels like GPT-5.5 crossed with Fable rather than GPT-5.6 crossed with Fable, what Fable found when it audited an Opus thread line by line, where Opus still clearly wins (3D, animation, and Claude Code limits at half Fable's price with no 50% weekly cap), plus Kimi K3, GLM 5.2, the Hugging Face hack, Grok 4.5, Codex, and T3 Code. Thanks to this episode's sponsor, General Translation: General Translation: https://nerdsnipe.link/gt Sources available on our Substack: https://nerdsnipe.substack.com/
- NS
S1 · E17July 14 · 2 hr 6 min5 different models dropped last week & the GPT-5.6 usage limits are brutal
This week Grok 4.5, Muse Spark 1.1, GPT-Live, and GPT 5.6 all dropped. Theo and Ben break down which models you should care about, how OpenAI fumbled the Codex to ChatGPT app transition, the abysmal usage numbers for 5.6, and the latest drama surrounding OpenAI and Sam Altman. Plus, how many satellites would it take to trap humanity on Earth, and how many data centers to heat up the ocean? Thank you to this episode's sponsors: Composio & WorkOs. https://nerdsnipe.link/composio https://nerdsnipe.link/workos Sources available on Substack: https://nerdsnipe.substack.com/
- NS
S1 · E16July 9 · 1 hr 11 minWe Tested GPT 5.6 Sol Early
We've spent six figures in tokens testing OpenAI's 5.6 Sol model to see whether if its better than Fable and what OpenAI have done to make it even better than 5.5. Also, we breakdown why we both moved our agents to Linux boxes, how to actually burn $65k on a single loop, and the Codex vs Claude Code subagent gap that's now bigger than the model gap itself. Thanks to this episode's sponsors Clerk and General Translation: Clerk, the auth platform with the best DX: https://nerdsnipe.link/clerk General Translation: https://nerdsnipe.link/gt Sources Available on our Substack: https://nerdsnipe.substack.com/
- NS
S1 · E15July 8 · 1 hr 16 minFable Is Back...kinda?
Fable was supposed to give us 14 days and instead we got 3 before the export ban. Now it's back at half the rate limits. Also, we break down why AI-generated Claude Code skills fall apart, what actually triggers Fable-to-Opus rerouting, the Mythos export-control timeline, and whether GLM 5.2 can really touch Sonnet 5. Thank you to PostHog for sponsoring today's episode! PostHog, all in one suite of product tools: https://nerdsnipe.link/posthog Sources available on our Substack: https://nerdsnipe.substack.com/
- NS
S1 · E14June 30 · 1 hr 11 minGPT-5.6 is here! And none of us can use it.
A new "government-approved rollout" for GPT-5.6 is coming, and it's got us wondering: are we entering the next dark age of AI model access? Additionally, we break down the launch of T3 Code x Grok CLI, Apple's price hikes, the memory supply crunch driving RAM and SSD costs up, a repo-poisoning and AI PR spam wave hitting open source, the GPT-5.6 non-launch, and what frontier model access looks like in a Fable/Mythos-tier world. Thank you to Composio for sponsoring this episode! Composio, connect your agents to everything: https://nerdsnipe.link/composio Sources Available on our Substack: https://nerdsnipe.substack.com/
- NS
S1 · E13June 24 · 1 hr 33 minThe US Government Banned Claude Fable 5...
The US government banned Anthropic's Mythos and Fable models just after launch so we break down exactly how Project Glass Wing, a panicked AWS engineer, and Dario's failure to communicate with Washington triggered the chaos. Plus: SpaceX's $60B Cursor acquisition, GLM 5.2 beating Google, the Codex trick for running 200 parallel agents, and why your next GPU will ship with a GPS tracker Thank you to Composio and WorkOS for sponsoring this episode! Composio, connect your agents to everything: https://nerdsnipe.link/composio WorkOS, your enterprise ready solution: https://nerdsnipe.link/workos Sources Available on our Substack: https://nerdsnipe.substack.com/
- NS
S1 · E12June 15 · 1 hr 17 minOur impressions of Claude Fable/Mythos (we filmed this before the ban)
RIP Fable 5. We recorded this before it got taken offline, but it's still worth talking about. The model is incredible. We really miss it. Thank you, Firecrawl, Depot, and Clerk for sponsoring! Firecrawl, the best api for searching and crawling the web: nerdsnipe.link/firecrawl Depot, better CI in every way: nerdsnipe.link/depot Clerk, the best dx in auth: nerdsnipe.link/clerk Sources https://x.com/thsottiaux/status/2043177597434306699 https://cognition.ai/blog/frontier-code https://x.com/paradite_/status/2064585901351792887 https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf (page 13 has the “prompt modification” quote) https://www.anthropic.com/institute/recursive-self-improvement Timestamps 0:00 Intro 3:29 Fable First Impressions 10:24 Benchmark Drama 26:42 Claude Code & Workflows 36:07 Pricing & June 22 Cutoff 39:54 Data Retention 45:58 Hidden Safeguards 1:07:01 The Claude Constitution
- NS
S1 · E11June 10 · 1 hr 35 minNow even Google's buying GPUs from SpaceX?
Cloudflare buys Void0, Google buying up compute from xAI, and Claude seems to be getting more anxious so we're here to break down everything this week on another episode of Nerd Snipe! Thank you to Composio for sponsoring today's episode! Composio, connect your agents to everything: https://nerdsnipe.link/composio Sources: https://x.com/vite_js/status/2062525206158078047 https://x.com/EdLudlow/status/2062970770612199542 https://x.com/nrehiew_/status/2063099050719846719 https://www.reddit.com/r/EconomyCharts/comments/1lp34n4/china_vs_us_energy/ https://x.com/elonmusk/status/1963443919150330139 https://x.com/AnthropicAI/status/2062568862479208923 https://x.com/HSVSphere/status/2060396271756595666 https://x.com/theo/status/2061018426152530232 https://x.com/Teknium/status/2062522290504613944 01:58 Cloudflare vs Vercel 13:03 Convex angle 19:59 SpaceX compute 35:49 AI self-improvement 44:45 Article reactions 50:28 Claude anxiety 57:42 Guardrails 01:08:21 Cursed image gen 01:19:18 Hermes agents
- NS
S1 · E10June 3 · 1 hr 15 minWe (mostly) like Claude Opus 4.8
Opus 4.8 + a ton of new stuff in Claude Code dropped this week, and we actually kinda like it. There's also a new benchmark that's actually good, and we have a lot of thoughts about the future of these AI labs... Thank you PostHog and Clerk for sponsoring! PostHog, the all in one suite of product tools: nerdsnipe.link/posthog Clerk, the best dx in auth: nerdsnipe.link/clerk SOURCES https://www.anthropic.com/news/claude-opus-4-8 https://x.com/theo/status/2060120708815139241 https://x.com/Baconbrix/status/2060065875911422343 https://x.com/zoink/status/2060769829133721974 https://x.com/maria_rcks/status/2060937270824153251 https://x.com/_catwu/status/2060054180379689074 https://x.com/datacurve/status/2060834005998793199 https://deepswe.datacurve.ai/ https://deepswe.datacurve.ai/blog#results https://github.com/scaleapi/SWE-agent/blob/402a7b8fdac8193f3f255bb53859ba274234f596/config/benchmarks/anthropic_filemap_multilingual.yaml https://deepswe.datacurve.ai/ https://x.com/AnthropicAI/status/2060061347522433422 https://x.com/theo/status/2060199299632472494 https://x.com/theo/status/2060901326058561795 https://x.com/38twelveDaily/status/2060408696945975631 00:00 - Shower Thoughts 02:44 - Deep SWE Benchmark 10:45 - Opus vs GPT-5.5 19:57 - Anthropic’s Huge Raise 25:39 - Token Maxing 40:02 - AI Slot Machine 43:49 - Claude Code Friction 50:01 - Opus, Mythos, and Safety
- NS
S1 · E9May 28 · 1 hr 22 minGoogle is Not a Serious Company
Not only did Google accidentally ban Railway's account, but their new flagship model Gemini 3.5 Flash is absurdly bad. Oh and apparently Theo's building his own cloud... Thank you Macroscope and GT for sponsoring today's episode! Macroscope: nerdsnipe.link/macroscope General Translation: nerdsnipe.link/gt SOURCES https://x.com/theo/status/2057359424378097823 https://x.com/JustJake/status/2056881510939283776 https://x.com/KorduGG/status/2059141337895604626 https://x.com/unboringtech/status/2059144145273491610 TIMESTAMPS 00:00 - Gemini fallout 05:21 - Video gen 10:10 - Google Cloud 14:04 - Windsurf/Cursor 25:10 - Manus/Meta 30:00 - China lock-in 35:00 - Cursor/xAI 45:00 - Cloud workflows 50:32 - Lakebed 65:22 - Hermes/security
- NS
S1 · E8May 20 · 1 hr 9 minHow the OpenClaw creator uses $1.3 million of tokens
Peter, the creator of openclaw is apparently going through $1,300,000 worth of tokens every month. Seems like we're not using nearly enough tokens. Oh and the security psychosis is getting worse. (Anthropic is being bad again too) Thank you to today's sponsors! - AgentMail, email inboxes for your agents: nerdsnipe.link/agentmail - Clerk, the best experience for auth, orgs, billing, and more: nerdsnipe.link/clerk TIMESTAMPS 00:00 - Intro 02:44 - Token Spend 06:47 - Token Future 12:12 - Secure Agents 22:50 - Anthropic Rules 31:35 - Security 39:36 - Token Tax 49:49 - AI Psychosis 59:47 - macOS 01:01:43 - AI Video
- NS
S1 · E7May 14 · 1 hr 36 minAnthropic solved their compute problem by buying it from Elon?
Anthropic seems to have finally solved their compute problems (kinda) by buying it from Elon, the security problem is getting so much worse, and apparently Bun's getting re-written in rust? Thank you to PostHog and Composio for sponsoring today's episode! - PostHog, the all in one suite of product tools: nerdsnipe.link/posthog - Composio, connect your agents to everything: nerdsnipe.link/composio Sources/references: https://x.com/claudeai/status/2052060691893227611 https://openai.com/index/elon-musk-wanted-an-openai-for-profit/#december-2018-elon-told-us-to-raise-billions-per-year-immediately-or-forget-it https://x.com/jarredsumner/status/2053391824702898475 https://x.com/jarredsumner/status/2051595933704761618 https://x.com/jarredsumner/status/2053047748191232310 https://ze3tar.github.io/post-zcrx.html https://www.jefftk.com/p/ai-is-breaking-two-vulnerability-cultures https://x.com/thdxr/status/2053570581807722968 https://x.com/badlogicgames/status/2052691176373805534 https://x.com/garrytan/status/2052996691586932783 TIMESTAMPS 00:00:00 - Anthropic/xAI00:08:59 - OpenAI Lawsuit00:25:40 - Bun in Rust00:34:34 - Security00:59:28 - Pottery Coding01:11:11 - Local Models
- NS
S1 · E6May 6 · 1 hr 47 minTheo Almost Lost $1 Million
This week Theo nearly lost a million dollars and Ben got AI psychosis (from gstack)... Thank you to Coderabbit and Clerk for sponsoring today's episode! - Coderabbit, the ultimate AI code reviewer: nerdsnipe.link/coderabbit - Clerk, the auth platform with the best DX: nerdsnipe.link/clerk Sources/references: - https://x.com/theo/status/2014863266888233193 - https://x.com/theo/status/2050305813894648289 - https://x.com/theo/status/2050314995561611357 - https://x.com/sama/status/2050671161915371998 - https://x.com/davis7/status/2050718508372431026 - https://x.com/thdxr/status/2050719575033983323 - https://x.com/davis7/status/2050762009592148375 - https://x.com/naval/status/2050560057675522500 - https://x.com/theo/status/1952229335416623592 00:00 - Intro / studio return 00:52 - Azure $1M / Microsoft 11:22 - Cloud platform talk 18:08 - Coding agents / SDKs 29:06 - GPT-5.5 pricing debate 45:14 - OpenClaw workflows 01:12:37 - G Stack / G Brain01:29:00 - Dynamic UI / wrap-up
- NS
S1 · E5May 1 · 1 hr 35 minWe need to talk about OpenAI
OpenAI and Microsoft are breaking up, Sam's drunk posting, Anthropic is being stupid again, and we still disagree about GPT-5.5Thanks to this episode's sponsors: - Clerk, the auth platform with the best DX: https://nerdsnipe.link/clerk- Coderabbit, the ultimate AI code reviewer: https://nerdsnipe.link/coderabbit- PlanetScale, the fastest and most scalable cloud databases: https://nerdsnipe.link/planetscale 00:00 Intro 04:05 Sam drunk tweets1 3:13 Anthropic billing woes 28:27 OpenAI divorce 42:34 GitHub can't stop dying 59:40 GPT-5.5 retrospective Sources/references: https://x.com/sama/status/2046808114561974567 https://x.com/sama/status/2046808217133670800 https://x.com/sama/status/2048160404376105179 https://x.com/sama/status/2047403771416940715 https://x.com/om_patel5/status/2048204411986469232 https://x.com/MSFTnews/status/2048749108127506936 https://x.com/GergelyOrosz/status/2048834949667537369 https://x.com/theo/status/2047721472521621991 https://x.com/mitchellh/status/2049213597419774026 https://x.com/kdaigle/status/2047803291988590609 https://x.com/ryanflorence/status/2048538797638599109 https://x.com/davis7/status/2048239401059434710 https://x.com/davis7/status/2048077518725366173 https://x.com/badlogicgames/status/2048444292562026713 https://x.com/0xSero/status/2048744545853030690