Your AI Agent Is Hemorrhaging Tokens on Thoughts It Already Acted On

Your AI Agent Is Hemorrhaging Tokens on Thoughts It Already Acted On

Every reasoning step your agent takes gets stapled to its context forever — even after the decision is done, the tool called, the file written. That’s not intelligence, that’s hoarding. A new paper says you can delete most of it without breaking anything. The question is whether you should believe them.

What happened

Researchers from a multi-institution team studied a specific, expensive problem in agentic workflows: as long-horizon agents execute tasks, their context window bloats with reasoning history that has already been acted upon. Their method, ICLR (Interaction Aware Compression for Long Horizon Reasoning), is training-free and works online — it scores reasoning blocks using a frozen proxy entropy signal and drops the low-value ones, while keeping actions, tool calls, and observations intact. Tested on 260 WorkBuddyBench tasks, ICLR pushed average reward from 0.699 to 0.718 while cutting input tokens by 25.5%, output tokens by 14.4%, and cache read tokens by 33.3%. The underlying theory is that once task-relevant state has been “externalized” — into code, files, tool outputs, or environment feedback — the reasoning that produced it becomes redundant. They call this making reasoning blocks “replaceable,” and they characterize agent memory vs. context window not as a log to preserve but as dynamic working state to actively manage. They also flag a danger they call “trajectory amplification”: delete the wrong reasoning block locally and you can trigger nonlinear, compounding changes in total computation downstream. Chain-of-thought reasoning, in other words, is load-bearing in ways that aren’t obvious until something collapses.

Cold read

The benchmark here is WorkBuddyBench, a single dataset of 260 tasks — narrow enough that generalization claims should be treated as hypotheses, not findings. A 2.7% reward improvement (0.699 → 0.718) is real but modest, and “reward” in agent benchmarks is often a coarse proxy for what actually matters in production. The proxy entropy scoring mechanism is frozen and heuristic — the paper does not show it reliably predicts which reasoning is safe to drop across different agent architectures, tool sets, or task domains. The “trajectory amplification” finding is the most honest and most alarming part of the paper: the authors are telling you that local deletions can cause nonlinear downstream failures, which means the risk surface of this method is hard to characterize without extensive testing on your specific workload. Training-free sounds like “free lunch” — it isn’t; you’re trading fine-tuned safety guarantees for operational convenience, and in production agentic systems, that’s a trade with teeth.

What it means for you

  • Signal maturity: 2/5 — single benchmark, modest gains, the dangerous edge case is named but not bounded
  • Who gets hurt: Startups running multi-step agents at scale (coding assistants, autonomous research tools, enterprise automation) who adopt this naively and hit trajectory amplification in production
  • What breaks if this is true: The business case for selling “full reasoning transparency” as a compliance or audit feature — if reasoning is ephemeral working state, not a permanent log, your audit trail story changes
  • Why it might not land: Proxy entropy scoring is not task-aware; in domains where reasoning steps carry long-range dependencies (legal, financial, medical workflows), “externalized state” may be an illusion and the pruning will silently corrupt task execution
  • Watch for: Major inference providers (Anthropic, OpenAI, Google) shipping native context compression APIs that implement some form of this — if they do, that’s the market validation signal; if they don’t by mid-2027, this stays academic

Forecast as of 2026-09-25

By Q3 2027, at least one major inference or agent orchestration platform will ship a production feature implementing selective reasoning compression for long-horizon agents, citing token cost reduction in the 20–35% range — but fewer than half of enterprise deployments will enable it by default due to audit and reproducibility concerns.


Source: When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression — Mingxuan Wang, Fei Luo, Bo Wang, Guorun Yao, Yinglong Guo, Chao Ning, Hongyue Chen, Yanbiao Ma, Jungong Han. https://arxiv.org/abs/2609.29875v1

Similar Posts