
Building Real Confidence in Chiplet Stacks with Evelyn Landman of proteanTecs
Chiplets aren't just a trend anymore, they're becoming the default architecture for the AI era. But as the industry stacks more dies into a single package to keep up with AI and HPC demands, it's inheriting a reliability problem nobody has fully solved: when four, eight, or more dies from different processes (and sometimes different vendors) share one package, a single weak link can disqualify the entire stack — and the cost of that failure is enormous. In this episode, OCP's Rob Coyle sits down with Evelyn Landman, co-founder and CTO of proteanTecs, to unpack why "known-good-die" testing is no longer enough and what it actually takes to reach "known-good-stack" confidence. They get into the hidden risks of multi-die packaging — thermal coupling, voltage droop, mechanical stress from through-silicon vias, and interface faults that traditional pass/fail testing simply can't catch — and how embedded telemetry is giving engineers real visibility inside the package, from bring-up and debug all the way through in-field operation. In this conversation: Why AI, HBM, and co-packaged optics are forcing the move to multi-die packages How "shift-left" testing catches faults before you commit to an expensive package Mix-and-match and binning: building a more consistent, higher-performing stack What embedded telemetry reveals that lab-condition pass/fail testing misses Real-time actuation and fleet-level predictive maintenance in the field How power and thermal insight extends useful life across data centers, automotive, and physical AI The role of OCP's Open Chiplet Economy in moving from siloed integration to open, interoperable chiplet operations Interested in the chiplet ecosystem? Reach out to Evelyn and the proteanTecs team, and get involved with the Open Chiplet Economy work stream at the Open Compute Project. Chapters 00:00 Intro: chiplets as the default AI-era architecture 01:11 What's pushing the industry to modular dies 03:22 The hidden cost: the problem shifts to the package 04:56 Shift-left: catching faults before you package 05:28 Single-die pass/fail vs. multi-die reality 07:25 Where reliability breaks down in the field today 10:06 What "known-good-stack" confidence looks like 12:20 Mix-and-match and binning before assembly 15:37 Embedded telemetry vs. traditional pass/fail 17:30 Thermals, hotspots, and heat-sensitive optics & memory 19:54 Real-time actuation to prevent imminent failure 20:59 Board- and fleet-level predictive maintenance 24:43 Scaling to millions of devices, tighter margins 27:16 OCP's Open Chiplet Economy & the road ahead About the Open Compute Project The Open Compute Project (OCP) is a collaborative community committed to redesigning hardware technology to efficiently support the growing demands on compute infrastructure. Its members span hyperscalers, ODMs, component manufacturers, and ecosystem partners — sharing open designs and working together on shared technical challenges that no single company can solve alone.