Publication
Accounting for Bias Enables Sustainable LLM Evaluation
Harshita Katoch; David Antony Selby; Gerrit Großmann; Sebastian Vollmer
In: Vitor Fortes Rey; René Schuster; Niklas Baumgarten; Sungho Suh; Tobias Christian Nauen (Hrsg.). International Joint Conference on Artificial Intelligence. IJCAI Workshop on Sustainability and Resource-Efficiency of Artificial Intelligence (SuRE-2026), 35th International Joint Conference on Artificial Intelligence, located at IJCAI-ECAI 2026, August 15-21, Bremen, Germany, CEUR-WS, ISBN 1613-0073, CEUR-WS, 8/2026.
Abstract
LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards
compensate for systematic measurement bias by running ever more comparisons, an approach that is both
statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treating
LLM judges as neutral, interchangeable instruments ignores documented biases like position bias, verbosity bias,
judge severity, and self-enhancement, that no volume of additional data can eliminate. We propose a unified
latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these
confounders, recovering reliable rankings from substantially fewer comparisons. Because fitting this model costs
negligible compute relative to a single round of LLM inference, bias correction is not only more statistically
rigorous but also a more sustainable approach to trustworthy evaluation.
