EP29 Binary Similarity in the LLM Era with Jonas Wagner and Endre Bangerter of ThreatRay
"Throwing more LLM prompts at a packed binary won’t make your analysis faster. But combining code similarity with AI will." In Episode 29 of Behind the Binary, Jonas Wagner and Endre Bangerter of ThreatRay join the podcast to lay out the real-world state of binary similarity in the LLM era. They demystify the current AI hype cycle, explaining why large language models are structurally unsuited for raw, unguided binary analysis, and how traditional binary similarity techniques are finding a massive resurgence as the ultimate intelligence pre-filter. Learn how they automate the classification of complex compiled samples by separating known-good libraries from unique malware code, and how these combined workflows are shrinking threat attribution timelines down from days to minutes. THE SESSION: Beyond BinDiff: Tracing the transition from expert-crafted syntactic tools to deep semantic vector matching. The Token-Saver Pipeline: How to stop wasting LLM context windows on standard library runtimes (like Go and Rust) and target only custom, unique malware functions. Dynamic Unpacking: Leveraging sandbox execution and memory dumps to capture clean binary representations before static analysis begins. Anti-AI Obfuscation: The emerging landscape of adversarial techniques designed to throw off both similarity algorithms and LLMs. Scale and Attribution: Building indexed code repositories to search back in time and attribute code reuse across decades of malware families. Join the Community Research Hub: Threat research, training events and news: https://cloud.google.com/security/flare The FLARE Insider: Get community updates and announcements. To subscribe, email flare-external@google.com FOLLOW THE SHOW: Subscribe: Apple Podcasts | Spotify | YouTube