Interesting new research from IBM.
If you pick models from benchmark deltas, some of that delta belongs to the phrasing rather than the model.
BenchDrift generates meaning-preserving variations of benchmark problems along linguistic, referential, pragmatic, and structural axes, holding the answer fixed, then measures how often correctness flips.
Phrasing sensitivity does not fade as models improve. It changes sign. Weak models gain more from rephrasing than they lose, while strong models lose far more than they gain, so the top models on a benchmark are the ones whose scores depend most on the wording they happened to receive.
Fragility also belongs to the rephrasing. Across eight models on GSM8K, MMLU, and MATH-Hard, they largely agree on which rephrasings cost the most correct answers even while differing in how much they drift overall.
Rephrasing breaks answers models were confident about, whether the problem gets shorter or longer.
Paper: https://arxiv.org/abs/2608.11694
Track more trending AI papers in our academy: https://academy.dair.ai/