Qualcomm's Durga Malladi: HBC vs HBM, Dragonfly AI 250, and Winning on TCO
Qualcomm re-entered the data center inference market late — competitive GPUs, XPUs, and custom accelerators are already out there. Durga Malladi (EVP/GM, Technology Planning, Edge & Data Center) explains why they didn't build another me-too accelerator and instead placed a differentiating technical bet: High Bandwidth Compute (HBC). The difference between HBC and HBM: HBM is stacked DRAM that shuttles data across wide buses to a separate accelerator — very high bandwidth, but expensive and power-hungry. HBC bonds stacked DRAM directly on top of a custom logic die using wafer-on-wafer techniques, so a lot of the compute runs literally next to memory. The result is lower latency and lower power per bit — and 18x effective bandwidth on the AI 250 versus the AI 200 at the SAME 768 GB LPDDR capacity and the SAME 160 kW per rack. Why the late-mover bet works: Qualcomm's differentiation isn't just architectural. Their foundry and memory-vendor relationships let them credibly promise the multi-megawatt-to-gigawatt supply hyperscalers need. Memory vendors are developing their own HBC-equivalent variants as partners rather than pure competitors, and Qualcomm expects the two technologies to co-evolve. Silicon is back in the lab; proof points land in Q1/Q2 2027. Chapters: 0:00 Introduction 2:49 The Memory Wall Problem 8:05 HBC Physical Architecture 9:33 HBC vs. Custom HBM 11:50 Silicon Proof Points Coming 12:08 Managing Thermals 13:34 Dragonfly AI 250 14:48 Trillion-Parameter Model on One Card 16:56 Tokenomics and TCO 18:26 Why TCO is Key 19:26 Manufacturing at Scale Follow Chipstrat: Newsletter: https://www.chipstrat.com X: https://x.com/chipstrat Follow Vik: Newsletter: https://www.viksnewsletter.com/ X: https://x.com/vikramskr Follow Semi Doped: Get more of Austin and Vik daily, free: https://daily.semidoped.com/
- Transcript







