Thinking Machines · @thinkymachines · X·2026-08-28 07:40·2小时前
AI 导读

Thinking Machines与UIUC、Bridgewater研究人员合作,将专家判断融入RLVR的每个环节,训练出首个在文本转SQL任务上超越人类基准的模型。该方法需前期投入数据清洗与奖励函数对齐,但最终在复杂任务上达到SOTA水平。

Thinking Machines@thinkymachines
35AI 编辑部评分,满分 100
2026-08-28 07:40· 2小时前
AI 导读

Thinking Machines与UIUC、Bridgewater研究人员合作,将专家判断融入RLVR的每个环节,训练出首个在文本转SQL任务上超越人类基准的模型。该方法需前期投入数据清洗与奖励函数对齐,但最终在复杂任务上达到SOTA水平。

Cleaning data and aligning the reward function for RLVR takes expertise and effort upfront, but the result is a model that's state-of-the-art on a complex task.

Guest post by researchers at UIUC and Bridgewater, in collaboration with our team. https://thinkingmachines.ai/news/putting-task-expertise-into-rl

TinkerLLMs with scaffolds have lagged on text-to-SQL, a task that relies on human judgment. By folding expert judgment into every part of RLVR on Tinker, @maxYuxuanZh...

来源:Thinking Machines· x.com