GRPO 超越英语:多语言与非英语环境下的大规模研究

Apple Machine Learning Research(RSS)·2026-08-18 08:00·1天前
AI 导读

一项大规模实证研究考察了 GRPO 在多语言和非英语环境下的表现,覆盖多种基础模型、训练语言及推理语言奖励设置。研究发现,以母语进行推理训练与英语推理训练之间的性能差距很小,表明 RLVR 在非英语场景下同样有效。该研究为多语言推理模型的强化学习训练提供了重要参考。

Apple Machine Learning Research(RSS)
50AI 编辑部评分,满分 100

GRPO 超越英语:多语言与非英语环境下的大规模研究

2026-08-18 08:00· 1天前
AI 导读

一项大规模实证研究考察了 GRPO 在多语言和非英语环境下的表现,覆盖多种基础模型、训练语言及推理语言奖励设置。研究发现,以母语进行推理训练与英语推理训练之间的性能差距很小,表明 RLVR 在非英语场景下同样有效。该研究为多语言推理模型的强化学习训练提供了重要参考。

Reinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for improving the reasoning capabilities of pretrained language models but current studies remain heavily English-centric. We conduct a large-scale empirical study of multilingual and non-English GRPO across a wide range of base models, training languages, and different reasoning language rewards. We find that training to reason in the native language often leaves only a small gap to training for English reasoning. We further observe strong…

来源:Apple Machine Learning Research(RSS)· machinelearning.apple.com