
Why autonomous driving is still so hard, with Joana Fonseca | Builders
Joana Fonseca is an AI engineer at TRATON, working on AI for autonomous driving. In this episode of Builders, Joana explains how vision language models (VLMs) and vision language action models (VLAs) could change autonomous driving by helping vehicles understand entire scenarios rather than simply detecting individual objects. She also breaks down what makes deploying AI in the physical world so difficult, from latency and model size to safety, edge cases, testing, and the huge amounts of diverse real-world data required. Joana holds a PhD in machine learning and robotics from KTH Royal Institute of Technology. Her work has taken her from developing algorithms for autonomous submarines to building AI systems for autonomous trucks at TRATON.In this episode: How VLMs differ from traditional computer vision How AI can understand an entire driving scenario The role of VLAs in autonomous vehicles Why 99.99% reliability may still not be enough Why latency becomes critical when AI operates in the physical world The data problem behind autonomous driving Why moving from prototype to production remains so difficult What production-ready AI looks like for autonomous systems Subscribe to Builders for more conversations with the people building the future of technology, AI, and engineering. Chapters (00:00) Introduction (01:41) From robotics to autonomous driving (03:49) Building AI for autonomous trucks (04:47) What are vision language models? (05:48) VLMs vs traditional computer vision (07:34) From VLMs to vision language action models (09:51) The black box problem (12:38) Why 99.99% safety isn’t enough (13:57) When AI latency becomes dangerous (15:01) Testing AI in the physical world (16:12) The autonomous driving data problem (17:28) Where current AI models fall short (18:49) Why prototypes struggle to reach production (19:52) Why robotaxis haven’t scaled everywhere (21:36) What makes AI production-ready














