# 马东锡NLP周报：RL数值失配与递归合成

- 来源：Dongxi 东锡 NLP (@dongxi_nlp)
- 发布时间：2026-08-09 16:36
- AIHOT 分数：21
- AIHOT 链接：https://aihot.virxact.com/items/cmslkle9j01ggroqyees9frke
- 原文链接：https://x.com/dongxi_nlp/status/2086370934512836707

## AI 摘要

最佳阅读 第32周

博客：

防御 RL 中训练-推理数值失配（尤其是线性注意力）——以及它是否有助于异步 RL

论文：

长时程终端任务的递归合成

## 正文

Best Read Week 32

Blog:

Defending Against the Training-Inference Numeric Mismatch in RL (Especially Linear Attention) - and Whether It Helps Async RL

Paper:

Recursive Synthesis for Long-Horizon Terminal Tasks
