Skip to content
Artwork for Beyond Coding
Beyond Coding · Yesterday · 45 min

Why Coding Agents Keep Making Your Codebase Worse

HumanLayer CEO Dexter Horthy, who coined "context engineering", on why AI agents degrade codebases over time. He breaks down a benchmark that measures AI code quality across features, and how his team ships with agents without drowning in slop. In this video, we cover: SlopCodeBench, SWE-bench and which coding benchmarks to trust Token maxing, dark factories and the "Instagram phase" of AI engineering Why 2-4x productivity beats chasing 100x Pull requests, design docs and how engineers build taste now Dex's agentic engineering workflow: one-shots, overnight janitor bots and research-design-plan For engineers and engineering leaders adopting AI coding agents who want to know what's real and what works in teams at scale. Timestamps: 00:00:00 - The Man Who Coined the Term: Context Engineering 00:00:54 - SlopCodeBench: Proof AI Makes Codebases Worse Over Time 00:05:49 - Background Agents and the Enterprise Friction Problem 00:08:13 - Why You'll Never One-Shot a Big Feature 00:11:18 - Scaling Agentic Engineering to 1,000 Engineers 00:16:19 - Dark Factories, Fake Hype and Token Maxing 00:19:55 - Why 2-4x Faster Beats Chasing 100x 00:21:31 - Should Teams Stop Doing Pull Requests? 00:24:20 - Can Deterministic Tools Replace Code Review? 00:27:37 - How to Build Engineering Taste When AI Writes Code 00:33:05 - Dex's Agentic Workflow: One-Shots and Janitor Bots 00:35:18 - Research, Design, Plan: Shipping Big Features With Agents 00:38:35 - Why Overnight Agents and Heavy Parallelism Backfire 00:40:10 - Subscription Maxing vs Paying Per Token 00:42:40 - Parallelize Like Poker, Not a Slot Machine Dex Horthy: https://x.com/dexhorthy Dex's talks and deep dives: https://aithatworks.bold.video #AgenticEngineering #AICoding #SoftwareEngineering

0:00-45:15

transcript

No transcript — this publisher did not publish one.

show notes

HumanLayer CEO Dexter Horthy, who coined "context engineering", on why AI agents degrade codebases over time. He breaks down a benchmark that measures AI code quality across features, and how his team ships with agents without drowning in slop.

In this video, we cover:

  • SlopCodeBench, SWE-bench and which coding benchmarks to trust
  • Token maxing, dark factories and the "Instagram phase" of AI engineering
  • Why 2-4x productivity beats chasing 100x
  • Pull requests, design docs and how engineers build taste now
  • Dex's agentic engineering workflow: one-shots, overnight janitor bots and research-design-plan

For engineers and engineering leaders adopting AI coding agents who want to know what's real and what works in teams at scale.

Timestamps:
00:00:00 - The Man Who Coined the Term: Context Engineering
00:00:54 - SlopCodeBench: Proof AI Makes Codebases Worse Over Time
00:05:49 - Background Agents and the Enterprise Friction Problem
00:08:13 - Why You'll Never One-Shot a Big Feature
00:11:18 - Scaling Agentic Engineering to 1,000 Engineers
00:16:19 - Dark Factories, Fake Hype and Token Maxing
00:19:55 - Why 2-4x Faster Beats Chasing 100x
00:21:31 - Should Teams Stop Doing Pull Requests?
00:24:20 - Can Deterministic Tools Replace Code Review?
00:27:37 - How to Build Engineering Taste When AI Writes Code
00:33:05 - Dex's Agentic Workflow: One-Shots and Janitor Bots
00:35:18 - Research, Design, Plan: Shipping Big Features With Agents
00:38:35 - Why Overnight Agents and Heavy Parallelism Backfire
00:40:10 - Subscription Maxing vs Paying Per Token
00:42:40 - Parallelize Like Poker, Not a Slot Machine

Dex Horthy: https://x.com/dexhorthy
Dex's talks and deep dives: https://aithatworks.bold.video

#AgenticEngineering #AICoding #SoftwareEngineering

links2