Skip to main content Skip to main navigation

Publication

Accounting for Bias Enables Sustainable LLM Evaluation

Harshita Katoch; David Antony Selby; Gerrit Großmann; Sebastian Vollmer
In: Vitor Fortes Rey; René Schuster; Niklas Baumgarten; Sungho Suh; Tobias Christian Nauen (Hrsg.). International Joint Conference on Artificial Intelligence. IJCAI Workshop on Sustainability and Resource-Efficiency of Artificial Intelligence (SuRE-2026), 35th International Joint Conference on Artificial Intelligence, located at IJCAI-ECAI 2026, August 15-21, Bremen, Germany, CEUR-WS, ISBN 1613-0073, CEUR-WS, 8/2026.

Abstract

LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treating LLM judges as neutral, interchangeable instruments ignores documented biases like position bias, verbosity bias, judge severity, and self-enhancement, that no volume of additional data can eliminate. We propose a unified latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders, recovering reliable rankings from substantially fewer comparisons. Because fitting this model costs negligible compute relative to a single round of LLM inference, bias correction is not only more statistically rigorous but also a more sustainable approach to trustworthy evaluation.

Projects

More links