Brilliant new paper from Meta.
LLM judges get validated on accuracy against golden data. That says nothing about whether the verdict survives when questioned.
The Wiggle Framework stress-tests 9 frontier models across 14 judging tasks along three axes, stability under re-prompting, stability under a single challenge, and stability under sustained pressure.
They find that every model wiggles. Verdicts flip 25 to 71% of the time under static pushback, and 62 to 91% against an adversarial persuader.
Pressure that changes a judge's verdict is almost always net-corrupting against ground truth.
Baseline jury majority strength turns out to be the best single-shot predictor of which items will move.
Paper: https://arxiv.org/abs/2608.12645
Track more trending AI papers in our academy: https://academy.dair.ai/