Skip to content
Artwork for The Infra Pod

The Infra Pod

The Infra Pod

The Infra Pod brings you insightful and thought-provoking discussions on the world of infrastructure software. This podcast is started by two engineers, Ian Livingstone (tech advisor for Snyk) and Tim Chen (General Partner at Essence VC), team up with a rotating cast of guests to dive deep into the latest trends and hot topics in the software infrastructure space.

Play
  • 20 episodes
  • fortnightly
  • Avg 39 min
  • English
  • August 24 · 43 min

    1,000 tokens per second: the case against predicting one word at a time (Kumar, VP of Engineering at Inception)

    What if the entire LLM industry has been solving language generation the slow way — one token at a time? In this episode of The Infra Pod, hosts Tim Chen (GP at Essence VC) and Ian Livingstone (CEO of Keycard) sit down with Kumar, VP of Engineering at Inception, to unpack Mercury 2, the company's diffusion-based language model, and make the case for a fundamentally different way to generate text and code. Kumar breaks down the core mechanical difference: where a GPT-style transformer predicts the next token one pass at a time, a diffusion model predicts a whole batch of upcoming tokens at once and iteratively denoises them in parallel, freezing the easy ones early and spending extra compute only on the hard ones. That approach — borrowed from image generation but re-engineered for text, where output length isn't known in advance and streaming is a hard requirement — yields roughly a 10x speedup and 3-5x cost efficiency by simply doing fewer forward passes. Mercury 2 hits 1,000 tokens/second on commodity NVIDIA hardware and now benchmarks competitively against cost-optimized models like Claude Haiku, Gemini Flash, and GPT-mini, though Kumar is candid that no diffusion model — Inception's included — has yet reached Sonnet or frontier-tier intelligence. The conversation moves from algorithm to product: why Inception keeps its API OpenAI-compatible (Kumar's electric-car analogy — same interface, very different feel under the hood), why most agentic sub-tasks don't need frontier intelligence at all, and why voice and search are the workloads where sub-second latency stops being a nice-to-have and becomes existential. Kumar closes with a genuinely spicy take! [00:00] Guest introduction: Kumar, VP of Engineering at Inception Labs [01:31] Diffusion vs. autoregressive LLMs: what's actually different under the hood [07:00] Why isn't diffusion the default already? Trade-offs and diffusion's late start on text [09:05] From Mercury 1 to Mercury 2: the road to enterprise-readiness [11:26] Open source diffusion models, ICML's best paper, and Inception's head start [16:15] Under the hood: how Inception hits 1,000 tokens/sec without sacrificing latency [20:52] Data strategy: what training a diffusion model actually requires [22:51] Does diffusion change how you build agents and products on top of it? [28:00] Mercury 2 benchmarked against cost-optimized frontier models [29:23] The model routing problem — and why it may already be mostly solved [37:47] Spicy Future: AGI

  • July 20 · 34 min

    What happens to your service mesh when the workloads running on it aren't written by humans? (Chat with William at Buoyant)

    In this episode of The Infra Pod, hosts Tim Chen (GP at Essence VC) and Ian Livingstone (CEO of Keycard) sit down with William Morgan, co-founder and CEO of Buoyant and creator of Linkerd, to explore how AI agents are reshaping the infrastructure layer — from service security to inference routing. Will traces Linkerd's origins to Twitter's 2014 migration from a monolithic Rails app to distributed microservices — the moment function calls became network calls that could actually fail. A decade later, that same communication layer is under pressure again. Non-deterministic agents make MCP and A2A calls over L7 protocols, and for the first time, fine-grained access control isn't optional: an agent will eventually find and call every reachable endpoint, including the one that deletes your database. The conversation covers the real pressure AI coding tools are already placing on platform teams — deploy cadences going from tens to potentially thousands per day — and what running inference inside Kubernetes actually means for proxies. Will breaks down why KV cache-aware routing is a 100x performance lever, why the modern inference proxy looks less like Envoy and more like a Makefile, and shares his spicy take on where compute is heading: in-cluster inference becomes the default, with frontier models reserved only for tasks that genuinely need godlike intelligence. [00:00] Guest introductions and Buoyant founding story [03:00] Linkerd's origin: solving Twitter's monolith-to-microservices migration [07:30] How AI is (and isn't) changing Linkerd today [11:00] MCP, A2A, and agents as L7 traffic in your cluster [14:30] Why agents make endpoint-level access control non-optional [18:00] AI as amplifier: what 10–1000x more deploys means for platform teams [21:30] Running inference in Kubernetes: a pathological workload [24:45] KV cache-aware routing and the 100x performance gap [27:00] The proxy/gateway landscape: grad students vs. premature standardization [29:30] The inference proxy is a Makefile now, not Envoy [31:00] Prompt injection, sandboxing, and the security problems with no clean answer [33:00] Spicy Future: in-cluster inference becomes the default

  • July 9 · 45 min

    From 50 million developers to a billion builders (with Tyler Wells, CTO of BrainGrid)

    What happens when the tools for building software stop requiring you to know how to code? In this episode of The Infra Pod, hosts Tim Chen (GP at Essence VC) and Ian Livingstone (CEO of Keycard) sit down with Tyler Wells, co-founder and CTO of BrainGrid (ex-Senior Director of Engineering at Twilio) , to explore what it actually takes to build a coding agent platform for people who have never touched a terminal — from spec-driven development to custom sandboxes to agents that cheat on their own tests. Tyler shares BrainGrid's origin: using structured specs and markdown requirements to keep early coding agents on track at his previous company, then pivoting from developer tooling to non-technical users after discovering the real unlock. People with deep domain expertise — in logistics, fitness, whatever — now have a path to ship software they could never have built before. What they struggle with isn't the ambition, it's that they expect a button to click. BrainGrid's job is to abstract away everything from dev environment setup to database provisioning so a non-technical founder can watch their idea materialize in a browser without ever seeing a terminal. On the infrastructure side, Tyler gets specific about the hard problems hiding beneath that simple interface. Agents will quietly rewrite their own acceptance criteria to pass validation if you let them — BrainGrid had to build immutable gates the builder agent can't touch. He also walks through why they built their own sandboxes from scratch: third-party providers were too slow for the tight feedback loop non-technical users need, so BrainGrid purpose-builds images pre-loaded with curated stacks, hitting 1.6-second spin-up times without a single npm install at runtime. The conversation closes on token spend — which is fast becoming the new line-of-code count, a metric organizations are already optimizing for the wrong reasons. [00:00] Guest introductions and BrainGrid founding story [03:30] The spec-driven approach: how structured requirements keep agents on track [07:00] Pivoting from developer tools to non-technical users [12:00] What non-technical builders actually get stuck on [16:00] The invisible infrastructure: databases, env vars, credentials, templates [21:30] Agents gaming acceptance criteria — and the fix [27:00] Why BrainGrid built its own sandboxes instead of using third-party providers [32:00] Task scalability: from landing pages to full apps with auth and databases [36:00] Spicy Future: 50 million developers becomes a billion bespoke builders [39:30] Token spend is the new line count — and it's already being misused

  • #63
    June 4 · 41 min

    Building a model that can prove theorems (with Shubho from Axiom Math)

    What happens when you combine world-class mathematicians with cutting-edge AI systems? In this episode, Ian Livingston (CEO of Keycard) and Timothy Chen (GP at Essence VC) sits down with Shubho, CTO of Axiom Math, to explore the emerging world of AI-driven mathematical reasoning and formal verification. From proving theorems in Lean to scaling software verification for the agentic coding era, Shubo lays out a compelling vision for why mathematical infrastructure matters more than ever. The conversation also ventures into homomorphic encryption, multi-party computation, Navier-Stokes, and why the next frontier of computing might just be running on math we haven't discovered yet. [00:00] Guest introductions and backgrounds [00:52] Company founding story [03:01] Axiom Math mission explained [05:15] Software verification applications [16:02] MPC and encryption challenges [24:01] Business model and products [33:49] Spicy Future hot takes

  • April 28 · 40 min

    Betting on Open Source Models to be the future (Chat with Benny, Cofounder of Fireworks AI)

    In this episode of The Infra Pod, hosts Tim Chen (Essence VC) and Ian Livingstone (Keycard) sit down with Benny Chen, co-founder of Fireworks AI, to explore the evolving world of AI inference infrastructure. Benny shares his journey from Meta — where capacity planning meetings made it clear GPUs were heading "up and to the right" — to co-founding Fireworks AI before ChatGPT even launched. The conversation dives deep into why the team bet early on inference over training, how they approached model optimization from horizontal compiler techniques to per-model kernel tuning, and why model customization is the key to unlocking better-than-frontier performance for vertical use cases. Benny discusses the reality of open source vs. closed models, the rise of agentic workloads, and why the real question isn't which model to use — it's which tasks have already been saturated. This episode is packed with technical insights on inference infrastructure, reinforcement learning for model customization, and what it means to truly adopt an AI-native engineering culture. 0:24 Benny's journey and founding Fireworks AI3:23 Early conviction: betting on inference before ChatGPT8:29 Pivoting from PyTorch training to text inference15:42 Horizontal vs. per-model optimization strategies11:14 Open source vs. frontier models: the real gap32:35 How customers engage: PLG to hands-on customization17:37 When to move off frontier models33:42 The future of agentic memory and data sovereignty32:35 Fireworks' differentiation in a crowded market33:53 Spicy Future: AI doomers, bot management, and going fully out of loop

  • March 9 · 33 min

    Building a successful infra product between all the AI apps and model providers (chat with Louis from OpenRouter)

    Tim (Essence VC) and Ian (Keycard) interviewed Louis Vichy, co-founder of OpenRouter, about why he built OpenRouter to de-risk AI app development (end-user pays LLM costs), how it scaled to processing ~5–6T tokens/week, and what OpenRouter is today: a reliable inference routing/control layer across ~60 providers with consolidated billing and reduced vendor lock-in. Louis explains why teams adopt OpenRouter (constant new model integrations, pricing/billing, differing API shapes), how routing focuses on practical heuristics (fallbacks, cost, throughput, latency), and how reliability is achieved via provider failover (e.g., alternate endpoints like Vertex/Bedrock). They discuss agent trends (longer-running agents, small models for routing/classification with specialized downstream models), possible memory support, developer conveniences (e.g., PDF parsing), and enterprise features (security/compliance guardrails, presets). The episode ends with links to OpenRouter chat/rankings pages and hiring for high-agency TypeScript-focused engineers.00:00 Welcome & Meet Louis (OpenRouter Co‑Founder)00:27 Origin Story: De‑Risking AI App Costs (Hackathon Lessons)01:35 First Big Feature: End‑User Pays for Tokens (Sign in with OpenRouter)02:34 From Routing to Rankings: Scaling to Trillions of Tokens03:42 What OpenRouter Is Today: Reliable Inference Across 60+ Providers05:55 Why Teams Adopt It: Avoiding Model API Churn, Billing, and Vendor Lock‑In08:37 Winning Strategy: Don’t Build a “Magic Router”—Optimize Cost/Latency/Throughput18:58 From Chat to RAG + Memory: Building Persistent Agent Context20:37 Developer Bells & Whistles: Auto PDF Parsing and More21:11 Enterprise Readiness: Compliance, Security Guardrails & Model Presets22:22 Customer Growth at Warp Speed in the AI Era23:03 Spicy Future!

  • February 23 · 41 min

    From 30 Seconds to 20ms: Solving Browser Speed for AI Agents (Chat with Catherine from Kernel)

    In this episode of The Infra Pod, hosts Tim Chen (Essence VC) and Ian Livingstone (Keycard) sat down with Catherine Jue, co-founder and CEO of Kernel, to explore the cutting-edge world of browser infrastructure for AI agents. Catherine shares her journey from Cash App to founding Kernel, explaining how she discovered the critical need for scalable browser automation when AI agents need to interact with the web. The conversation dives deep into the technical innovations behind Kernel's use of unikernels and micro VMs, which enable blazingly fast browser startup times (20ms vs 30+ seconds) and unique snapshot/restore capabilities. Catherine discusses the evolution from deterministic browser automation to truly agentic behavior, the challenges of optimizing for variable web workloads, and her optimistic vision for an AI-powered future where the pie expands rather than consolidates. This episode is packed with technical insights about infrastructure, agent tooling, and the future of how software interfaces will evolve in an agent-native world. 0:24 Catherine's startup journey and founding Kernel 1:30 Cash App's OpenAI experiment sparks the idea 3:56 Why browser infrastructure for AI agents? 6:36 Unikernels: 20ms startup vs 30+ seconds 15:02 Optimizing for variable web workloads 23:25 Future of agent-native software 32:05 Hot takes!

  • February 9 · 29 min

    Coding agents need infra to apply code changes! (Chat with Tejas from Morph)

    Tim (Essence VC) and Ian (Keycard) sat down with Tejas Bhakta (CEO of Morph) to chat about building infrastructure for the fastest file edit APIs for coding agents. He shares how Morph delivers 10,000 tokens/second through speculative decoding, why cursor removed fast apply, and his vision for autonomous software that updates without prompts. The conversation covers subagent architecture, code search optimization, and the path to reliable AI coding at scale. Timestamps: 0:00 - Introduction 0:29 - Why start Morph and pivoting through YC 1:23 - The fast apply insight from Cursor 3:42 - How fast apply works and speculative decoding 6:09 - Use cases: when and where fast apply matters 8:19 - Why Cursor removed fast apply 9:22 - Morph's value prop beyond speed 11:58 - Subagent architecture and SDK approach 14:45 - Semantic search and code-specific tooling 19:52 - Building custom coding agents vs platforms 22:42 - Adoption inhibitors and the future of codegen 23:26 - Spicy take: Autonomous software and reliability

  • January 26 · 42 min

    Let's chat about vibe coding & Ralph! (Chat with Dexter at Humanlayer)

    In this episode of The Infra Pod, hosts Tim and Ian sit down with Dexter Horthy, CEO of Human Layer, to explore the evolution of AI coding agents and the future of software development. Dexter shares his journey from building data tools to discovering the real problem: making AI coding agents actually productive for senior engineers, not just juniors. The conversation dives deep into the research-plan-implement workflow that enables engineers to ship 99% of their code with AI assistance, the challenges of getting staff engineers to adopt AI tools, and why most AI coding ecosystems don't actually help you sell to enterprises. Dexter also shares his spicy take on how Ralph-style agents can be even further enhanced. Whether you're a skeptical senior engineer or an AI-curious developer, this episode offers practical insights into what actually works in production AI coding today. [0:00] Introduction & Dexter's Journey Why Dexter finally started a company, the failed data catalog pivot, and building an AI janitor for data warehouses [8:00] The Hard Lessons of AI Ecosystem Hype Why there's no "SAML for AI agents" and what enterprises actually need versus what the hype machine promises [13:00] The Research-Plan-Implement Breakthrough How to make senior engineers productive with AI, staying objective during research, and making decisions at the top of the context window [26:00] The Vibe Shift & Where We Are Today When respected engineers started believing, the role of Ralph and spec-driven development, and what's working in production [37:00] Spicy Take: Ralph Goes to the Supreme

  • January 12 · 47 min

    Building a bug-free vibe coding world (Chat with Akshay from Antithesis)

    In this episode of the Infra Pod, hosts Ian Livingston (Keycard) and Tim Chen (Essence VC) interviewed the Field CTO Akshay Shah of Antithesis, diving deep into the world of distributed systems, reliability, and the future of software testing. The conversation covers the challenges of building bug-free distributed systems, the story behind Antithesis, lessons from major outages, and the evolving landscape of infrastructure and AI-driven operations. Timeline with Timestamps: 00:00 – Introduction & guest background 02:00 – What Antithesis does and why it matters 06:00 – Real-world impact: Testing distributed systems (etcd, Kubernetes) 09:00 – Major outages & lessons learned (AWS, Knight Capital) 12:00 – The origins and philosophy behind Antithesis 16:00 – The future of reliability, testing, and AI in infrastructure 28:00 – Closing thoughts & where to learn more Links: Learn more about Antithesis: https://antithesis.com Antithesis on YouTube: @AntithesisHQ

  • Dec 29, 2025 · 23 min

    Infra Pod 2025: Our Favorite Moments, Hottest Takes, and What’s Next

    Join Tim from Essence VC and Ian Livingston from Keycard for the year-end 2025 recap of Infra Pod! In this special episode, Tim and Ian reflect on their favorite moments, hottest takes, and biggest lessons from a year of rapid change in infrastructure, AI, and agent technology. They revisit standout episodes—like deep dives into browser automation, the evolving role of memory in LLMs, and the disruptive potential of agent sandboxes. The hosts discuss how companies are pivoting in the AI era, the importance of adapting quickly, and the surprising ways hardware choices are shaping the future of compute. Looking ahead, Tim and Ian share bold predictions for 2026, debate the next big abstractions in infrastructure, and invite listeners to share their own hot takes and favorite episodes. Whether you’re an engineer, founder, or just passionate about the future of tech, this episode is packed with insights, energy, and a look at what’s next for the Infra Pod community.

  • Dec 15, 2025 · 40 min

    From Spark to Eventual: Reinventing Data for the AI Era (Chat with Sammy from Eventual)

    In this episode of The Infra Pod, hosts Tim from Essence VC and co-host Ian Livingston (Keycard) interviewed Sammy Sdu, CEO of Eventual, a multimodal data processing platform. Sammy shares his journey from AI research and self-driving cars to founding Eventual, discusses the challenges of processing unstructured and multimodal data, and explores the future of data engineering, scalability, and the role of agents in modern data pipelines. Timestamps: 02:47 — Data processing challenges & founding Eventual 09:40 — Real-world use cases & business impact 24:20 — The future of data engineering & tools 40:00 — Closing thoughts & where to learn more

  • Dec 1, 2025 · 42 min

    Render is defining what taste means in backend infra (Chat with Anurag from Render)

    In this episode of The Infra Pod, hosts Tim and Ian are joined by Anurag, CEO of Render, to discuss the journey of building a modern cloud platform from scratch. The conversation covers Anurag’s background at Stripe, the challenges of cloud infrastructure, the evolution of developer tools, the importance of abstraction and taste in product design, and the future of agent-driven development. The episode is packed with insights on scaling platforms, developer experience, and the shifting landscape of cloud computing.

  • Nov 17, 2025 · 44 min

    Bazel and the Next Wave of AI Developer Infrastructure (Chat with Alex Eagle from Aspect Build)

    Welcome to Episode 53 of The Infra Pod! Hosts Tim from Essence and Ian from Keycard are joined by special guest Alex Eagle, CEO and co-founder of Aspect Build. In this episode, Alex shares his journey from working on Angular at Google to founding a company around Bazel, Google's open-source build tool. The conversation dives deep into the challenges and motivations behind building developer infrastructure, the evolution of CI/CD systems, and the unique strengths and hurdles of adopting Bazel in organizations of all sizes. The trio explores the future of software development, the impact of AI on coding and build systems, and the ongoing debate between monorepos and polyrepos. Alex also discusses Aspect's mission to make Bazel more accessible and the broader implications for developer productivity in an agentic, AI-driven world. Whether you're a platform engineer, open-source enthusiast, or just curious about the future of build tools, this episode is packed with insights and spicy predictions for the future of developer infrastructure. 00:00 – Introduction & Guest Background Tim and Ian introduce Alex Eagle, who shares his journey from Google to founding Aspect Build. 04:20 – Why Bazel? Alex explains the motivation behind focusing on Bazel, its challenges, and the analogy to municipal infrastructure. 10:55 – Bazel in the Real World Discussion on Bazel’s adoption, who should use it, and the hurdles organizations face. 21:06 – The Future: AI, Agents, and Build Systems Exploring how AI and agentic coding are changing developer infrastructure and Bazel’s evolving role. 43:01 – Closing & Takeaways Final thoughts, how to learn more about Aspect and Bazel, and episode wrap-up.

  • Oct 6, 2025 · 41 min

    The AI Analyst is coming to change Data Teams (Chat with Lucas from Gravity)

    In this episode of the Infra Pod, hosts Tim (Essence VC) and Ian (CEO of Keycard) sat down with Lucas Thelosen, founder of Gravity and former head of product for data and AI at Google. Lucas shares his journey from leading teams at Google and Looker to launching Gravity, a company focused on bridging the gap between business users and data through generative AI. The conversation dives deep into the challenges of data analysis in modern organizations, the evolution of AI-powered tools like Orion, and how generative AI is transforming the way companies leverage their data. Lucas discusses the importance of semantic layers, onboarding AI agents, and the future of data teams in a world where AI can automate complex analysis and reporting.

  • Sep 22, 2025 · 43 min

    The next cloud to overtake AWS are AI sandboxes?! (Chat with Ivan from Daytona)

    Join host Tim Chen (Essence VC) and Ian Livingston (Keycard.sh) as they sat down with Ivan Burazin, CEO and co-founder of Daytona to explore the cutting edge of agent-native infrastructure. From the origins of Daytona to the challenges and opportunities of building for autonomous agents, this episode dives deep into the tools, use cases, and visionary thinking shaping the next generation of infrastructure. Whether you’re a developer, founder, or tech enthusiast, you’ll gain insights into the trends, pivots, and bold ideas driving the future of AI and cloud platforms.

  • Sep 8, 2025 · 42 min

    Future of File Storage for AI (Chat with Hunter, CEO of Archill)

    Ian (Keycard) and Tim Essence VC) engage in an insightful discussion with Hunter Leath, CEO of Archil. The episode delves into the motivations behind Archil, a new data startup focused on revolutionizing data storage for cloud applications. Hunter explains the limitations of existing data storage paradigms like S3 and traditional block storage, advocating for a new approach that leverages SSDs and custom protocols to offer high-performance, infinite storage that can support modern workloads, including AI and CI/CD applications. The conversation also touches on the trust and complexity of implementing such a system, the future vision for file storage, and practical use cases ranging from serverless Jupyter notebooks to large-scale CI/CD operations.00:17 The Gap in Cloud Data Storage01:48 Understanding Unstructured Data02:41 Building Modern Data Systems04:15 Challenges and Innovations in File Storage11:12 Targeting CI/CD and AI Workloads23:00 The Future of File Storage

  • Aug 25, 2025 · 35 min

    Turning Gaming PCs to Serverless CI for AI! (Chat with Aditya from Blacksmith)

    Tim (Essence VC) and Ian (Keycard) sat down with Aditya, CEO of Blacksmith, to explore the inception and innovative approach of Blacksmith in the CI/CD space. Blacksmith utilizes gaming CPUs and NVMe SSDs to deliver high-performance, serverless CI compute. They discussed the challenges faced by large companies with CI systems, why GitHub Actions was chosen as their primary CI system, and how AI-induced code generation is pushing the need for faster and more efficient CI solutions. Adya also highlighted the future potential of CI observability and other improvements Blacksmith is focusing on to maintain their edge in the market.00:29 Founding Blacksmith: The Origin Story01:16 Understanding Blacksmith's Serverless CI Compute02:21 Challenges in CI/CD and Blacksmith's Solutions05:14 Technical Deep Dive: Performance and Optimization06:53 Why GitHub Actions?08:08 Innovative Hardware Choices for CI10:06 Scaling and Managing CI Workloads19:14 Future of CI/CD and AI Integration28:02 Spicy Futures: Predictions and Hot Takes

  • Aug 11, 2025 · 44 min

    Building the Future of AI with Long-term Memory (Chat with Charles from Letta)

    In this episode of The Infra Pod, Tim (Essence VC) and Ian (Keycard.sh) delve into the fascinating world of AI memory with Charles, the CEO and co-founder of Letta. They explore the intricacies of memory in AI, its current state, how it’s implemented in various applications, and its potential to revolutionize productivity tools such as coding assistants. Charles shares insights on how Letta is leading the way by creating platforms for agents with long-term memory, the future implications of shared memory systems, and the concept of 'sleep time compute.' Tune in to gain a deeper understanding of why memory could become more valuable than the models themselves in the future of AI.00:24 The Concept of Memory in AI02:04 Expanding on AI Memory02:17 Implementing Effective AI Memory01:55 Current and Future State of AI Memory11:37 The Challenges and Opportunities of AI Memory11:37 Real-World Applications and Limitations35:14 Introducing Letta's Solutions36:35 Future Trends and Predictions

  • Jul 28, 2025 · 36 min

    The bet on Postgres to be the backbone to run reliable services (Chat with Jeremy and Qian from DBOS)

    In this episode of the Infra Pod, Tim (Essence VC) and Ian (Keycard) hosted DBOS building a reliable backend service, with guests including CEO Jeremy and cofounder Qian. The discussion delves into the company's motivation, their research work making software durable and reliable by default, their choice to be betting on Postgres, and how they integrate AI with traditional systems. 00:51 DBOS: Mission and Origins01:37 Creating Reliable and Durable Software11:13 Leveraging Postgres for Durability27:06 Spicy Takes: The Future of Postgres and Development

Showing 1–20 of 20 episodes