Finally a good paper testing whether self-reflection loops are worth it.
Setup: Seven methods, open models at 1.5B, 3B and 7B, two math benchmarks, 150 questions each.
Every generated token counted, including the ones spent on critiques, reflections, debate turns and checking. Each method compared against repeated sampling at its own measured cost, paired by question, with bootstrap intervals and multiplicity correction.
All 36 comparisons came back with no reliable win for any method. Ten are reliably worse, and every one of those is a method where the model inspects its own output. All 18 self-inspection comparisons are negative.
Self-Refine and a forced Reflexion sit 3.6 to 10.1 points below the baseline at 7B. Reflexion as published never triggered its own retry on the smallest model. It judged itself correct every time and quietly became a single chain of thought.
Worth knowing before you add another critique step to your agent loop.
Paper: https://arxiv.org/abs/2607.28576
Track more trending AI papers in our academy: https://academy.dair.ai/