
1,000 tokens per second: the case against predicting one word at a time (Kumar, VP of Engineering at Inception)
What if the entire LLM industry has been solving language generation the slow way — one token at a time? In this episode of The Infra Pod, hosts Tim Chen (GP at Essence VC) and Ian Livingstone (CEO of Keycard) sit down with Kumar, VP of Engineering at Inception, to unpack Mercury 2, the company's diffusion-based language model, and make the case for a fundamentally different way to generate text and code. Kumar breaks down the core mechanical difference: where a GPT-style transformer predicts the next token one pass at a time, a diffusion model predicts a whole batch of upcoming tokens at once and iteratively denoises them in parallel, freezing the easy ones early and spending extra compute only on the hard ones. That approach — borrowed from image generation but re-engineered for text, where output length isn't known in advance and streaming is a hard requirement — yields roughly a 10x speedup and 3-5x cost efficiency by simply doing fewer forward passes. Mercury 2 hits 1,000 tokens/second on commodity NVIDIA hardware and now benchmarks competitively against cost-optimized models like Claude Haiku, Gemini Flash, and GPT-mini, though Kumar is candid that no diffusion model — Inception's included — has yet reached Sonnet or frontier-tier intelligence. The conversation moves from algorithm to product: why Inception keeps its API OpenAI-compatible (Kumar's electric-car analogy — same interface, very different feel under the hood), why most agentic sub-tasks don't need frontier intelligence at all, and why voice and search are the workloads where sub-second latency stops being a nice-to-have and becomes existential. Kumar closes with a genuinely spicy take! [00:00] Guest introduction: Kumar, VP of Engineering at Inception Labs [01:31] Diffusion vs. autoregressive LLMs: what's actually different under the hood [07:00] Why isn't diffusion the default already? Trade-offs and diffusion's late start on text [09:05] From Mercury 1 to Mercury 2: the road to enterprise-readiness [11:26] Open source diffusion models, ICML's best paper, and Inception's head start [16:15] Under the hood: how Inception hits 1,000 tokens/sec without sacrificing latency [20:52] Data strategy: what training a diffusion model actually requires [22:51] Does diffusion change how you build agents and products on top of it? [28:00] Mercury 2 benchmarked against cost-optimized frontier models [29:23] The model routing problem — and why it may already be mostly solved [37:47] Spicy Future: AGI