# 虚假奖励与噪声数据对RLVR的影响

- 来源：Dongxi 东锡 NLP (@dongxi_nlp)
- 发布时间：2026-08-05 06:27
- AIHOT 分数：29
- AIHOT 链接：https://aihot.virxact.com/items/cmsf9cdnw1k6pro2emjhrr3bx
- 原文链接：https://x.com/dongxi_nlp/status/2084768262534173033

## AI 摘要

虚假奖励（2025年6月）

随机奖励可通过放大已有先验来提高分数，但并未产生真正的能力提升。

噪声数据具有破坏性（2026年8月）

经核实的错误标签会强化错误、减少探索，且表现不如干净监督。

论文：

虚假奖励：重新思考RLVR中的训练信号
噪声数据对基于可验证奖励的强化学习具有破坏性

## 正文

Spurious Rewards （June， 2025）

Random rewards can raise scores by amplifying existing priors， without creating genuine capability gains.

Noisy Data Is Destructive （Aug， 2026）

Verified wrong labels reinforce errors， reduce exploration， and underperform clean supervision.

Papers：

Spurious Rewards： Rethinking Training Signals in RLVR
Noisy Data is Destructive to Reinforcement Learning with Verifiable Rewards
