Your AI Agent Aced the Benchmark and Then Failed Every Real Task

Your AI Agent Aced the Benchmark and Then Failed Every Real Task

Vendors are selling you fine-tuned agents with glowing eval scores. Those scores, it turns out, measure something almost completely different from whether the agent actually does the job. The gap isn’t small—it’s the difference between looking good on paper and delivering 10% completion in production.

What happened

Researchers tested four models—Qwen3 at 4B and 14B parameters, Gemma 3 at 4B and 12B parameters—before and after supervised fine-tuning on multi-turn customer-support workflows. They measured performance two ways: the standard “next-turn” protocol (predict the next action given the gold interaction history, score it against a reference) and end-to-end agentic workflow execution (the model runs the whole task autonomously, no gold history as a crutch). SFT improved next-turn scores for every single model—clean, consistent, looks great in a slide deck. But when those same fine-tuned models were cut loose on real autonomous workflows, none of the four succeeded under holistic evaluation. Strict trajectory completion topped out at 10.4% workflow success—the best result across all four models. Tool-specific gains were also inconsistent across metrics and models, meaning even the fine-grained intermediate signals were unreliable predictors.

Cold read

This study is scoped to customer-support workflows only, so generalizing to coding agents, research agents, or other verticals requires a leap the data doesn’t support. The models tested are mid-tier open-weights (4B–14B); whether the same disconnect holds at frontier scale—GPT-4o class, Claude-tier—is simply unknown from this work. The 10.4% ceiling is damning, but the paper doesn’t decompose why the gap exists at a mechanistic level—error accumulation, hallucination, tool misuse, or something else. Most critically, the finding that “next-turn metrics don’t predict workflow success” is not the same as proving that no lightweight metric can predict it—the authors motivate separate reporting, not the abolition of turn-level evals. Treat this as a calibration warning, not a framework for scrapping your entire eval stack.

What it means for you

  • Signal maturity: 4/5 — Empirically solid within its scope; the core finding is clean and directionally trustworthy
  • Who gets hurt: Companies that bought or built customer-support automation on next-turn benchmark performance and are now wondering why CSAT isn’t moving
  • What breaks if this is true: Every vendor demo that shows golden dataset eval scores as proof-of-production-readiness is selling you a metric that doesn’t transfer; your procurement criteria are measuring the wrong thing
  • Why it might not land: Vendors and internal teams have strong incentives to keep using next-turn metrics—they’re cheap, fast, and make models look good; end-to-end eval is expensive and slow
  • Watch for: Any foundation model lab or agent platform that starts publishing both turn-level and end-to-end workflow completion rates in the same model card—that’s the tell that the industry is actually absorbing this lesson

Forecast as of 2026-09-21

By Q3 2027, at least two major enterprise AI agent platforms (likely in the customer-support or sales verticals) will publicly disclose end-to-end workflow completion benchmarks alongside turn-level metrics following customer pressure or regulatory guidance—but the majority of vendor marketing materials will still lead with next-turn scores.


Source: When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success — Md Tahmid Rahman Laskar, Xue-Yong Fu, Gundeep Singh, Karol Chang, Kevin Sanders, Shi Zong, Tania Habib, Julien Bouvier Tremblay, Shayna Gardiner, Harsh Saini, Matthias Lee, Elena Khasanova, Quinten McNamara, Shashi Bhushan TN. https://arxiv.org/abs/2609.21187v1

Similar Posts