
Why Coding Agents Keep Making Your Codebase Worse
transcript
show notes
HumanLayer CEO Dexter Horthy, who coined "context engineering", on why AI agents degrade codebases over time. He breaks down a benchmark that measures AI code quality across features, and how his team ships with agents without drowning in slop.
In this video, we cover:
- SlopCodeBench, SWE-bench and which coding benchmarks to trust
- Token maxing, dark factories and the "Instagram phase" of AI engineering
- Why 2-4x productivity beats chasing 100x
- Pull requests, design docs and how engineers build taste now
- Dex's agentic engineering workflow: one-shots, overnight janitor bots and research-design-plan
For engineers and engineering leaders adopting AI coding agents who want to know what's real and what works in teams at scale.
Timestamps:
00:00:00 - The Man Who Coined the Term: Context Engineering
00:00:54 - SlopCodeBench: Proof AI Makes Codebases Worse Over Time
00:05:49 - Background Agents and the Enterprise Friction Problem
00:08:13 - Why You'll Never One-Shot a Big Feature
00:11:18 - Scaling Agentic Engineering to 1,000 Engineers
00:16:19 - Dark Factories, Fake Hype and Token Maxing
00:19:55 - Why 2-4x Faster Beats Chasing 100x
00:21:31 - Should Teams Stop Doing Pull Requests?
00:24:20 - Can Deterministic Tools Replace Code Review?
00:27:37 - How to Build Engineering Taste When AI Writes Code
00:33:05 - Dex's Agentic Workflow: One-Shots and Janitor Bots
00:35:18 - Research, Design, Plan: Shipping Big Features With Agents
00:38:35 - Why Overnight Agents and Heavy Parallelism Backfire
00:40:10 - Subscription Maxing vs Paying Per Token
00:42:40 - Parallelize Like Poker, Not a Slot Machine
Dex Horthy: https://x.com/dexhorthy
Dex's talks and deep dives: https://aithatworks.bold.video
#AgenticEngineering #AICoding #SoftwareEngineering





