Skip to content
Artwork for My Weird Prompts
My Weird Prompts · September 1 · 27 min

CPU vs GPU Inference: When to Skip the GPU

The received wisdom says serious AI inference needs a GPU — but that's not the full story. This episode unpacks the real bottlenecks: memory bandwidth vs. parallel compute, prefill vs. decode, and the surprising workloads where CPU actually wins. From classical ML models like XGBoost to Whisper transcription and Mixtral on a laptop, we explore when the GPU advantage compresses from 100x to 6x, and why batch size one changes everything. If you're running internal chatbots, batch processing audio, or deploying small transformers, the economics may surprise you. Episode #043993 — open it directly at myweirdprompts.com/043993

0:00-27:57

transcript

No transcript — this publisher did not publish one.

show notes

The received wisdom says serious AI inference needs a GPU — but that's not the full story. This episode unpacks the real bottlenecks: memory bandwidth vs. parallel compute, prefill vs. decode, and the surprising workloads where CPU actually wins. From classical ML models like XGBoost to Whisper transcription and Mixtral on a laptop, we explore when the GPU advantage compresses from 100x to 6x, and why batch size one changes everything. If you're running internal chatbots, batch processing audio, or deploying small transformers, the economics may surprise you.

Episode #043993 — open it directly at myweirdprompts.com/043993