Agent Learning Speed Is Doubling Every Three Months—Moore’s Law Is Back
Agent Learning Speed Is Doubling Every Three Months—Moore’s Law Is Back
A 50-author paper just dropped what may be the first empirical scaling law for AI agents after deployment—not in the lab, not on synthetic tasks, but grinding through real work for up to 12 hours at a stretch. If the curve holds, the agent you deploy today is half as capable as the one you’ll run in Q4. That’s either the best news you’ve heard this year, or a reason to stop building any moat that depends on today’s performance ceiling.
What happened

Researchers built EdgeBench, a suite of 134 real-world tasks spanning software engineering, scientific discovery, formal mathematics, combinatorial optimization, professional knowledge work, and interactive games. Every task demands at least 12 hours of continuous agent operation with rich, multi-level feedback—this is not a multiple-choice eval. Across roughly 38,000 hours of agent-environment interaction, they found that overall performance follows a log-sigmoid scaling law with an R² of 0.998—a nearly perfect statistical fit. Critically, this isn’t pretraining data scaling; it’s learning from the environment during deployment, the phase that agentic workflows actually live in. They also measured cross-generation improvement: agent learning speed roughly doubles every three months. Only 51 of the 134 tasks are publicly released alongside their evaluation framework, with the rest held back.
Cold read
An R² of 0.998 on 134 tasks sounds bulletproof, but a log-sigmoid curve is flexible enough to fit almost any monotonic S-curve—high goodness-of-fit is not the same as a law with predictive power outside the observed range. The “doubling every three months” claim is a cross-model-generation average, which means it could be masking huge variance between task types; software engineering and formal mathematics almost certainly scale differently than knowledge-work tasks. The benchmark contamination problem is not addressed in the abstract—tasks built with “substantial expert effort” at one lab are still tasks that future models may have been trained on or adjacent to. The 83 withheld tasks are also a flag: you cannot independently verify the law holds on the full 134. And “real-world” is doing heavy lifting here—these are curated, timed, feedback-rich tasks, not the messy, underspecified, low-feedback environments most production agentic AI actually encounters.
What it means for you
- Signal maturity: 2/5 — Compelling first evidence, but one paper from one lab with a partially withheld benchmark is not a verified law
- Who gets hurt: Vertical SaaS companies whose differentiation is “AI that learns your workflows over time”—if environment-learning scales this predictably, that moat compresses fast as foundation models catch up
- What breaks if this is true: Any pricing or retention model built around a long ramp-up period for agent capability; customers won’t need 6 months to see value, they’ll need 6 weeks
- Why it might not land: Most production tasks lack the “rich, multi-level feedback” this benchmark requires—without that signal, the scaling law may simply not apply
- Watch for: Independent replication on the 51 public tasks by a lab with no co-authors on this paper; if the R² holds at 0.99+ externally, start believing it
Forecast as of 2026-07-07
By Q2 2027, at least one major foundation-model lab (OpenAI, Anthropic, Google DeepMind) will publish its own environment-learning scaling analysis explicitly citing EdgeBench—either confirming the log-sigmoid law on their own task suite or publishing a rebuttal showing the curve breaks down in low-feedback real-world conditions. If neither happens, the “three-month doubling” claim will remain vendor folklore rather than science.
Source: EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments — Deyao Zhu, Xin Zhou, Shengling Qin, Xuekai Zhu, Hangliang Ding, Shu Zhong, Zixin Wen, Zhonglin Xie, Chenhui Gou, Linxuan Ren, Yueyang Wang, Junfeng Zhong, Rui Liu, Tian Gao, Yangguang Lin, Jingyuan Zhang, Maojia Song, Xuan Qi, Jinhong Wu, Chenyang Zhang, Yinzhu Piao, Ziru Niu, Hongbin Lin, Lingxiang Meng, Peng Tang, Chengyao Tang, Shanyu Wu, Huanyu Zheng, Yu Liu, Liya Zhu, He Wang, Ming Ding, Ziyu Wan, Hao Liu, Sibo Wang, Haotian Zhu, Xintian Zhang, Nan Chai, Yipeng Liu, Panhao Lai, Sihang Yuan, Zixin Su, Ge Zhang, Wangchunshu Zhou, Yantao Du, Wenhao Huang, Guang Shi. https://arxiv.org/abs/2607.05155v1
