Most Enterprise AI Projects Are Failing Because Nobody Knows How to Grade Them
Most Enterprise AI Projects Are Failing Because Nobody Knows How to Grade Them
You shipped a GenAI workflow. Your board wants ROI numbers. You have vibes. That gap—between “the model is impressive” and “this is worth scaling”—is where most enterprise AI initiatives go to die. A new framework from a global banking pilot says it has the measurement protocol you’re missing. Read the cold version before you forward this to your CFO.
What happened

Researchers at a global bank piloted EnterpriseVal, a structured evaluation system designed to answer the deployment question that public benchmarks never answer: is this specific workflow fit, reliable, safe, and worth scaling on our data, under our controls? The paper argues that enterprise GenAI’s failure to show measurable business impact is “substantially a measurement problem,” not a model capability problem. Their system specifies five components: a formal use-case configuration (model, prompts, retrieval, guardrails, human oversight), a metric catalogue, a grading protocol that blends blinded expert judgment with calibrated LLM-as-judge scoring, a two-tier REJECT/CONDITIONAL/SCALE decision gate, and a value-and-risk model where reviewer catch rate is a measured parameter rather than an assumption. In a credit-memo drafting workflow, hallucination rate came in at 1.6% against a gate threshold of 5%, and citation precision hit 88% against a gate of 70%—both passing. In a procedure-transformation workflow, analyst refinement time dropped from an estimated 27.4 hours to 2.9 hours per document. The protocol scales expert grading with prediction-powered inference, a technique for deriving confidence bounds without grading every output by hand.
Cold read
One bank, three workflows, one pilot—that is the entire empirical base. The authors themselves separate “established results,” “documented pilot evidence,” “the proposed system,” and “open hypotheses,” which is admirable transparency and also a sign that the SCALE gate hasn’t actually been validated at scale. The 27.4-to-2.9-hour efficiency figure is described as an estimated baseline, meaning the before number was not directly measured under controlled conditions—exactly the kind of measurement sloppiness the paper argues against. The LLM-as-judge calibration step is load-bearing: if the judge model drifts or is itself hallucinating, the entire grading pipeline produces confident-looking garbage. And “a large fraction of agentic projects are expected to be cancelled” is cited as motivation but not sourced in the abstract—a hype claim doing rhetorical work without evidence. The framework is rigorous in design; it has not yet proven it changes outcomes.
What it means for you
- Signal maturity: 2/5 — Promising architecture, single-institution pilot, no cross-industry replication
- Who gets hurt: AI platform vendors selling “enterprise-ready” on the basis of public benchmark scores; their customers who bought without a deployment-specific eval
- What breaks if this is true: The standard practice of demoing GPT-4o on your data for two weeks and calling it a business case collapses; every enterprise sale now needs a measurement layer before it can close
- Why it might not land: Building this eval infrastructure costs real money and expertise most mid-market companies don’t have—the framework may become a consulting product rather than an operational standard
- Watch for: Whether a major analyst firm (Gartner, Forrester) or cloud provider (Azure, AWS) adopts EnterpriseVal’s REJECT/CONDITIONAL/SCALE gate language as a procurement checklist—that’s the forcing function for adoption
Forecast as of 2026-09-21
By Q3 2027, at least one Fortune 500 enterprise software vendor will publish a deployment evaluation methodology visibly borrowed from frameworks like EnterpriseVal—but fewer than 20% of mid-market enterprise AI buyers will have a formal use-case-level eval process in place, meaning the measurement gap this paper diagnoses will remain the primary reason GenAI initiatives fail to show ROI.
Source: EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise — Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid. https://arxiv.org/abs/2609.21841v1
