AI Agents Can Code Your Research — But Can’t Actually Do It
AI Agents Can Code Your Research — But Can’t Actually Do It
The entire “recursive self-improvement” story — AI gets smarter by doing AI research, loop repeats, we all retire — rests on one assumption nobody had properly tested. A team of 24 researchers just tested it. The agents faceplanted.
What happened
Researchers at Princeton, DeepMind, and collaborating institutions devised what they call “shadow evaluations”: take a real, unpublished NeurIPS 2026 submission, hand its central research question to a frontier AI agent, and have the paper’s original authors grade the output — people who know exactly what good looks like and can’t be fooled by confident-sounding prose. Two papers were tested. Agents were given six full days and thousands of dollars of compute — not a toy budget. The agentic workflow completed all of the engineering autonomously — writing code, running experiments, no human hand-holding. And then it stopped mattering: both outputs were unambiguously rejected by the authors. Five specific failure modes were catalogued: poor judgment about the bar for publishable research, uncreative responses to design shortcomings, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced the same failures — this wasn’t one bad run.
Cold read
Two papers is not a dataset; it is two data points. The authors chose NeurIPS-caliber submissions, which is the hardest tier of AI research — this tells you nothing about whether agents can handle easier, more incremental research tasks that make up the bulk of actual lab work. “Unambiguously rejected” is a crisp result, but grading is still done by the papers’ own authors, who have obvious priors about what the answer should look like and may unconsciously penalize approaches that deviate from their own framing. The engineering success is also undersold in the doom framing: an agent that autonomously executes a full experimental pipeline for six days is genuinely not nothing — the line between “does the engineering” and “does the research” is exactly where billion-dollar agentic AI bets are being placed right now. Finally, frontier models in July 2026 will not be frontier models in twelve months; failure modes catalogued today are an engineering backlog, not a permanent ceiling.
What it means for you
- Signal maturity: 2/5 — two case studies with author-graded output; directionally useful, statistically thin
- Who gets hurt: Startups selling “autonomous AI researcher” or “AI co-scientist” products to pharma, biotech, and deep-tech R&D budgets on the premise that the agent closes the loop without a PhD in the room
- What breaks if this is true: The “AI does AI research → faster AI → AI does more AI research” flywheel that justifies current hyperscaler capex projections doesn’t spin; the doubling-time forecasts are fiction
- Why it might not land: Engineering is 70–80% of research cycle time in many applied ML shops; an agent that nails the engineering and flags dead ends faster than a junior researcher still has real ROI, even if it can’t originate the hypothesis
- Watch for: Any lab publishing shadow-evaluation results at scale (10+ papers, multiple domains) — that’s when this becomes evidence rather than an anecdote
Forecast as of 2026-07-30
By Q3 2027, no frontier agent will have a verified, independently-graded record of producing an accepted paper at a top-4 ML venue (NeurIPS/ICML/ICLR/CVPR) from an open-ended prompt without human co-authorship directing the research question — and the field will still be debating whether the eval methodology is fair.
Source: Can AI agents conduct open-ended AI research? Early evidence from two case studies — Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, Konstantinos Voudouris, Viet Nguyen, Toby Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, Helen Toner, Gillian Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani, Arvind Narayanan. https://arxiv.org/abs/2607.27191v1
