Skip to content
Artwork for the localhost

the localhost

the localhost

Five tech industry veterans argue about AI that runs on your own hardware. Open weight models, NPUs, home labs, and the real question underneath it all: who controls intelligence, you or a data center?

Every week, one host takes the hot seat to defend a provocative claim while the other four try to take it apart. Then everyone votes. We run the models, test the vendor claims ourselves, and share what actually happened, including the failures.

No hype, no scripts, and nobody here is pretending to have all the answers. We're learning this stuff alongside you, one episode at a time.

New episodes weekly. There's no place like 127.0.0.1.

Play
  • 5 episodes
  • weekly
  • Avg 41 min
  • English
  • #5
    Yesterday · 40 min

    the localhost:0005 | what actually works

    Five guys, five machines, and one question: what actually works? No debate this week. No vote. We just opened up our laptops and showed each other what we're actually running, what broke, and what turned out to be worth the trouble. Jacob figured out the whole local-plus-cloud workflow on a train from Munich to Paris with a spotty hotspot. Robert and Jacob are both splitting AI work across multiple machines with NVIDIA's new Personal AI Router, which means that graphics card in your desk drawer just became useful again. And Frank built an app that tells Uber drivers where to be, running entirely on a server in his office, controlled from his phone. Plus: Mozilla says the gap between the best cloud models and the ones you can download has narrowed to about four months. Neil takes over the Surface AI Factory. Robert finds out about it live on the show. And if you've ever thought "I could never run AI myself," the whole middle of this episode is the answer. There's no syntax anymore. It's just a conversation. CHAPTERS 0:00 Cold open: what actually works 1:00 Around the horn, and Neil's news 2:18 Robert's next chapter 3:16 The Rundown: the four month gap 11:35 How do you actually start? 12:10 Chauncey: okay but what do I DO with it? 13:00 Jacob's train from Munich to Paris 14:50 Get over the fear factor 23:00 NVIDIA PAIR: your old GPU just got useful 25:49 A brief and important pop culture quiz 26:40 Perplexity Computer, and the bicycle company 36:27 The Uber driver app running on The Beast 39:00 Wrap: there's no place like 127.0.0.1 the localhost is five tech industry veterans arguing about AI that runs on your own hardware. Open weight models, NPUs, home labs, and the question underneath it all: who controls intelligence, you or a data center? New episodes weekly. https://thelocalhost.show

  • #4
    September 14 · 41 min

    the localhost:0004 | the escape hatch

    An escape hatch is the thing that lets you leave. You build on something that also works on somebody else's hardware, so your vendor can't lock you in. Open weight models were supposed to be that hatch for AI. Download them, run them yourself, walk away whenever you want. Then NVIDIA bought Hugging Face, the place everyone downloads them from. Frank worked at Silicon Graphics in the '90s and says he's seen this movie before. Five of us argue about it for forty minutes: is the hatch still open, or did the company you'd escape from just buy the exit? Plus what happened when we ran a 180 billion parameter model at home, and why one of us cancelled a $200 a month subscription. (0:00) Cold open: I've seen this movie before (2:18) The Rundown: NVIDIA closes the Hugging Face deal (12:31) More from the week's news (22:43) The Tension: is the escape hatch still open? (27:46) Neil's counter: the ovens, the recipes, and the CUDA moat (29:15) Robert flips it: a door instead of a hatch? (30:10) The vote (33:52) On My Device (39:38) Wrap the localhost is five tech industry veterans arguing about AI that runs on your own hardware. Open weight models, NPUs, home labs, and the question underneath it all: who controls intelligence, you or a data center? New episodes weekly. thelocalhost.show

  • #3
    September 6 · 44 min

    the localhost:0003 | the ai hardware boom is beginning

    Last week we argued the $25 token was dying. This week: that's exactly when the hardware boom starts. Frank takes the hot seat with a paradox. If intelligence keeps getting cheaper, we don't use less of it — we invent more things to do with it. More agents, more inference, more workloads. Which leaves every enterprise with one question they haven't had to ask before: which intelligence should we rent, and which should we own? Citi and Vercel are reportedly running open-weight models more than half the time. In June it was 29%. By late August, 53%. A new family of models out of the UAE ships six versions built for six classes of hardware — phone to server — with no quantization involved. NVIDIA's Hugging Face acquisition is official. And OpenAI shipped Astra. We also had four people instead of five, a guest from London, and for the first time in three episodes, a completely unanimous vote. ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ CHAPTERS 00:00 If intelligence gets cheap, what do you buy? 00:47 Welcome back — who's here, who's at Disneyland 01:28 Laura Osborn joins from London 02:14 Neil, and the Bears/Seahawks problem 03:07 Chauncey, and the disclosure 03:43 The Rundown: NVIDIA and Hugging Face is official 04:25 Neil: what GitHub can tell us about this 05:48 Citi and Vercel pass 50% open weights 07:50 The three-tier model: endpoint, edge, cloud 08:56 Six models for six classes of hardware 10:22 Purpose-built beats shrinking things down? 11:21 What UK and European customers actually ask for 14:41 From cloud architect to local AI 15:18 OpenAI ships Astra 16:26 Chauncey: be careful what you point it at 17:51 Neil: code generation at the edge 19:33 Baseline: what should people actually care about? 20:20 Open weights are no longer philosophical 21:27 What's coming out of IFA 22:52 The medical school of the future 25:00 Why doesn't anyone know this is possible yet? 26:53 "We have more agents than human employees" 28:23 Tension of the Week: the hardware boom begins 30:27 Chauncey: fixed cost beats a variable one 32:39 Laura: I spent five years telling everyone to go cloud 34:18 Neil: when pennies turn into a million dollars 35:00 The brewery test 36:06 The data center next door 36:34 The vote 38:42 On My Device: the 36B gets the job 40:07 Chauncey: Copilot CLI, running locally 41:44 The Dream Team and the one-person virtual business 41:56 Laura: an on-device speaker coach 42:58 Neil: Mac vs the Beast, round two ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ WHAT CAME UP • NVIDIA acquiring Hugging Face is confirmed. NVIDIA says it stays open and that NVIDIA hardware won't be required. Neil's read: this looks like Microsoft buying GitHub — the same fears, and possibly the same outcome. His angle is that Hugging Face has been effectively off-limits to a lot of enterprises, and being inside a vendor with a security and compliance apparatus may be what finally gets those models sanctioned. • Citi and Vercel are reported to be running open-weight models for more than half their AI workloads — 29% in June, 53% by late August. Ten weeks. • The MBZUAI Institute of Foundational Models released six models spanning roughly 1B to 375B parameters, each trained for a specific class of hardware rather than quantized down from something larger. Frank ran the 36B on his dual-3090 workstation against Qwen3.8-27B: about 68 tokens/sec against 40, on a comparable qualifying profile. • Neil's three tiers: the endpoint, the cloud, and an emerging middle — departmental edge servers. Local doesn't grant you HIPAA or FERPA compliance by itself, but no round trip to the cloud is a smaller attack surface. • The economics nobody models: an insurance agent filing 50 claims a day, times 10,000 agents. Even at pennies per call, that's over a million dollars a year in workloads that could run at the edge. • Laura, five years a cloud architect, on doing a full 180: the constraint in Europe isn't just sovereignty rules, it's that cloud capacity is at its limits. • And Neil's field test of public sentiment, conducted on a bartender who asked him point blank whether he was anti-AI. The vote was 4–0. First unanimous verdict of the series. ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ THIS WEEK'S CREW Frank Buchholz — host, independent Laura Osborn — guest, joining from London Neil Misak Chauncey Larsen Jacob Rhoades and Robert Henry are off this week. Jacob is celebrating his first anniversary. Robert was on a plane, sending Teams messages anyway. DISCLOSURE: Laura, Neil and Chauncey are employed by Microsoft. Frank is independent. All views expressed are their own, nothing discussed is unannounced or non-public, and none of this is financial advice. ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ localhost is a weekly show about running AI on hardware you own. Independent and unsponsored. New episodes weekly: https://thelocalhost.show There's no place like 127.0.0.1. #LocalAI #OpenWeights #AI #EdgeAI #AIHardware #Sovereignty #OnDeviceAI #AIPodcast

    • Transcript
  • #2
    August 30 · 37 min

    the localhost:0002 | the $25 token is dead

    Last week the conversation was cost vs. control. This week the cost floor fell out. Frontier models are getting dramatically cheaper, the models you can run on your own hardware are getting dramatically better — and Robert makes the case that within 18 months the only thing frontier labs will really be able to sell is speed. The crew gets into NVIDIA's reported interest in Hugging Face, a 170B+ parameter model coaxed onto a single consumer card, what Apple's latest silicon means for local AI, and the question underneath all of it: when intelligence becomes a commodity, where does the money go? Plus: an aerospace CEO in the back of an Uber who can't send his engineering IP to the cloud, and what happened when Robert turned on autopilot and walked away from his desk. the localhost is Frank Buchholz (independent), Jacob Rhoades, Robert Henry, Neil Misak, and Chauncey Larsen (Microsoft, opinions their own). Independent and unsponsored. New episodes weekly: https://thelocalhost.show There's no place like 127.0.0.1.

    • Transcript
    • Chapters
  • #1
    August 22 · 41 min

    the localhost:0001 | open is a position

    Five tech industry veterans, one hot seat, and a vote. In our first episode: Meta gave away a new AI model and promised the big one soon, NVIDIA's CEO used his first ever post on X to defend open weights, and 270 companies signed on. One did not. So Frank takes the hot seat to argue that openness is positional, not principled, and the crew tries to take him apart. Along the way: what open weight actually means (up to the secret sauce), the Flock camera controversy as edge AI's cautionary tale, and a job interview for five AI models on a used $700 graphics card, with a winner nobody was talking about. Plus the stat of the night: one host's router setup runs 86 percent of his AI locally and has already saved him $1,600. We run the models, test the claims, and share what actually happened, including our own mistakes. New episodes weekly. There's no place like 127.0.0.1. Creators & Guests Chauncey Larsen - Host Frank Buchholz - Host Jacob Rhoades - Host Neil Misak - Host Robert Henry - Host (00:00) - cold open: from cost to control (01:05) - meet the hosts (and who writes the paychecks) (04:50) - the rundown: qwen 3.8 lands (07:55) - open weight vs open source, explained (14:45) - flock cameras: edge capture, central control (17:40) - this week's tension: open is a position (21:05) - the debate: manifestos, chess games, and IPOs (29:15) - the vote (31:40) - on my device: five models walk into a job interview (34:45) - what's running on the crew's devices (37:45) - the 86 percent stat (40:20) - wrap: there's no place like 127.0.0.1 Mark Zuckerberg's manifesto, "The Future Is for Everyone": https://www.meta.com/thefutureisforeveryone/ The Open Weights letter coverage and Spark 1.2 status: https://www.orcarouter.ai/blog/meta-muse-spark-1-2-explained Glimmer launch coverage: https://www.cnbc.com/2026/08/10/meta-muse-glimmer-open-weight-ai.html The May "not suitable for open sourcing" quote: https://www.implicator.ai/meta-releases-30b-open-weight-muse-glimmer-and-promises-spark-1-2-weights/ Flock's announced changes: https://www.nbcnews.com/tech/security/flock-safety-police-abuse-oversight-data-retention-rcna592217 Qwen 3.8-27B: https://huggingface.co/Qwen/Qwen3.8-27B The 2.4T open release: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B Frank's full benchmark, raw traces and all: https://github.com/frankcx1/bakeoff The results post: https://www.linkedin.com/posts/frankjbuchholz_on-monday-meta-gave-away-a-brand-new-ai-share-7493881833236123648-1kd1/ the localhost is five friends talking about AI on your own hardware: Frank Buchholz (independent), Jacob Rhoades, Robert Henry, Neil Misak, and Chauncey Larsen (Microsoft, opinions their own). The show is independent and unsponsored. Find us at https://thelocalhost.show

    • Transcript
    • Chapters
Showing 1–5 of 5 episodes