

Tokens Per Watt: Why Your Context Window Is a Power Decision
This story was originally published on HackerNoon at: https://hackernoon.com/tokens-per-watt-why-your-context-window-is-a-power-decision. On an H100, tokens per watt drops 12x between 4K and 64K context. Agents live at the fat end of that curve. The fix comes from semiconductor architecture. Check more stories related to machine-learning at: https://hackernoon.com/c/machine-learning. You can also check exclusive content about #artificial-intelligence, #agentic-ai, #semiconductors, #llm-inference, #ai-infrastructure, #tokens-per-watt, #software-engineering, #gpu, and more. This story was written by: @ajjayg. Learn more about this writer by checking @ajjayg's about page, and for more stories, please visit hackernoon.com. A March 2026 paper derives what its authors call the 1/W law: tokens per watt halves every time the serving context window doubles. On an H100 running Llama-3.1-70B, that's 17.6 tok/W at 4K context and 1.50 tok/W at 64K. Same silicon, roughly 12x worse efficiency, purely from context length (arXiv:2603.17280). Agents are the single worst workload for that law, because a tool-calling loop re-sends its entire accumulated history on every step. Chip designers hit a structurally similar wall in 2004 and answered with power domains, DVFS, and clock gating rather than a better transistor. The translation to agent architecture is real. But it breaks in one specific place that's worth knowing about before you bet your GPU budget on it.


















