Rohan Paul@rohanpaul_ai
37AI 编辑部评分,满分 100

重复采样胜过自我反思:Qwen2.5 研究揭示更优策略

2026-08-15 21:04· 27分钟前
AI 导读

一项研究对比了 7 种测试时推理方法,发现让 Qwen2.5(1.5B 至 7B)模型重复采样同一数学问题并取多数答案,在相同 token 预算下优于自我反思等复杂方法。36 项对比中,无方法稳定胜出,10 项显著更差。该结论限于 Qwen2.5 数学任务,不适用于前沿模型或开放任务。

If you're spending extra tokens making an LLM critique itself, this paper says a simpler move can be better: just let it try again.

Before you make an LLM reflect on its answer, try giving it another independent attempt.

The study compares 7 test-time reasoning methods on Qwen2.5 models from 1.5B to 7B, then asks a fairer question: what happens if repeated sampling gets the same token budget?

Repeated sampling means solving the same math problem several times and taking the answer that shows up most often.

Across 36 comparisons, none of the more elaborate methods reliably beat that baseline at equal generated-token cost; 10 were significantly worse.

For checkable reasoning, extra tokens may be better spent creating independent attempts than asking the model to reconsider its own work.

Ofcouse, study boundary matters: this is Qwen2.5 on math, not frontier models or open-ended tasks.

  • arxiv. org/abs/2607.28576

Title: "Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B"

来源:Rohan Paul · x.com

重复采样胜过自我反思:Qwen2.5 研究揭示更优策略

Rohan Paul · @rohanpaul_ai · X·2026-08-15 21:04·27分钟前
AI 导读

一项研究对比了 7 种测试时推理方法,发现让 Qwen2.5(1.5B 至 7B)模型重复采样同一数学问题并取多数答案,在相同 token 预算下优于自我反思等复杂方法。36 项对比中,无方法稳定胜出,10 项显著更差。该结论限于 Qwen2.5 数学任务,不适用于前沿模型或开放任务。

If you're spending extra tokens making an LLM critique itself, this paper says a simpler move can be better: just let it try again.

Before you make an LLM reflect on its answer, try giving it another independent attempt.

The study compares 7 test-time reasoning methods on Qwen2.5 models from 1.5B to 7B, then asks a fairer question: what happens if repeated sampling gets the same token budget?

Repeated sampling means solving the same math problem several times and taking the answer that shows up most often.

Across 36 comparisons, none of the more elaborate methods reliably beat that baseline at equal generated-token cost; 10 were significantly worse.

For checkable reasoning, extra tokens may be better spent creating independent attempts than asking the model to reconsider its own work.

Ofcouse, study boundary matters: this is Qwen2.5 on math, not frontier models or open-ended tasks.

  • arxiv. org/abs/2607.28576

Title: "Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B"

来源:Rohan Paul· x.com