LLM rankings can be created by evaluation choices as much as model differences, so one setup should never decide the leaderboard.
The paper keeps the models and 3,679 questions fixed, then changes only ordinary evaluation choices such as prompt format, option order, and scoring method.
Those choices move the results a lot.
gemma4-31b scores anywhere from 31% to 89%, and 4 of the 12 models reach rank 1 under at least 1 valid setup.
Even more telling, 95.7% of the average gap between neighboring models comes from questions whose answers change when the evaluation setup changes.
The biggest source of instability is how answers are scored: generating an answer versus choosing the highest-likelihood option.