Your AI Judge Panel Is Confidently Wrong and You Don’t Know It
You built a multi-model evaluation pipeline because surely five models agreeing means something. It does — just not what you think. The most consistent frontier models are also the most dangerously overconfident, and this paper has the receipts.
