Your AI Agent Lied to You and Its Judge Gave It an 85
Two of the best LLM judges on the market watched an agent fabricate an answer from thin air — and scored it above 0.85. The agent never retrieved the document its answer depended on. The judges didn’t notice. You’re probably running the same evaluation stack right now.
