
Evals - Beyond the Vibe Check | LangChain, Langfuse, Mercor, CoreWeave & Galileo
Everyone says they do evals. Almost nobody does. LangChain actually put a number on it: around 89% of teams have observability set up, and only about a third ever do anything with what it collects. So I got five people who do this for a living in a room, from LangChain, Mercor, Galileo, Langfuse, and CoreWeave. Side note, three of those companies got acquired in the past year. Galileo to Cisco, Langfuse to ClickHouse, Weights & Biases to CoreWeave. The layer is getting bought up faster than most teams can figure out how to use it. A few things I didn't expect going in. Grading an agent while it runs can cost you more than the agent does. The model you're using as a judge is probably wrong in ways you'll never notice. And the guy whose company sells eval tooling told the room to stop writing evals before shipping, just put it out and learn from what breaks. We also got into the parts nobody blogs about. How many traces a person still has to read by hand. Why the same agent is a weekend project internally and a year of work at a bank. What happens when your users start doing things you never thought to test. If you've shipped an agent and had no real way to know whether it was working, this one's worth your time! Chapters 00:00 Welcome 00:34 Meet the Panel 02:11 Why Nobody Runs Evals 03:35 Traces Explained 04:14 Offline Evals Basics 05:13 Online Evals in Production 08:26 Surprises From Production 11:22 Trajectory Checks and Bucketing 12:51 Verifiers and Trusting Evals 14:21 Hard Lessons at Scale 19:10 Human Review and Judge Drift 21:29 Which Agents Are Hardest 23:41 Voice and Multimodal Evals 27:10 Cheap Binary Guardrails 30:09 Enterprise Deployment Lifecycle 35:24 Tooling and Expert Knowledge 37:40 Fixing Long Horizon Agents 40:12 Error Analysis Workflow 41:11 Trajectory Evals and Runbooks 42:27 Condensing Long Traces 43:09 Real Long Running Agents 44:11 What's Hardest to Evaluate 45:06 Verifying Tool Side Effects 46:20 Evaluating Auto Research 48:20 Picking a Sampling Rate 49:59 Metric Drift and Security 53:49 Common Evals Mistakes 57:34 Regulated Industry Rollouts 01:01:37 Sourcing Domain Experts 01:03:50 Hallucination Evaluators 01:05:37 Predictions for Evals 01:11:55 Audience Q&A Begins 01:16:30 Are Complex Playbooks Ready 01:18:22 Where Models Fall Short 01:20:04 Extracting Expert Knowledge 01:22:10 Evaluating Writing Style 01:23:24 Closing Remarks GUESTS Liam Bush - Deployed Engineer @ LangChain Braden Holstege - VP Enterprise AI @ Mercor Soumya Mohan - Head of Product @ Galileo Lotte Verheyden - Head of Developer Relations @ Langfuse Emmanuel Turlay - Director of Engineering @ CoreWeave
- Transcript














