Nathan Lambert@natolambert
27AI 编辑部评分,满分 100
2026-08-07 21:54· 36分钟前
AI 导读

研究团队在TorchTitan RL + vLLM上为Gated DeltaNet(Qwen3.5-9B/35B-A3B)实现训练/生成逐位一致,logprob差为0,为开源首个线性注意力模型达成此目标。异步RL下,offpolicy=32时标准栈偏差超0.065,新方法保持约0.035。代价是训练吞吐量降低2-3倍,团队建议将其作为调试工具而非生产默认。

nice rl experiment on train-inference mismatch

Yichuan WangZero Train-Inference Mismatch - now for linear attention, and under async RL 🎯 We got bitwise-exact trainer/generator parity for Gated DeltaNet (Qwen3.5-9B / 3...

来源:Nathan Lambert · x.com

Nathan Lambert · @natolambert · X·2026-08-07 21:54·36分钟前
AI 导读

研究团队在TorchTitan RL + vLLM上为Gated DeltaNet(Qwen3.5-9B/35B-A3B)实现训练/生成逐位一致,logprob差为0,为开源首个线性注意力模型达成此目标。异步RL下,offpolicy=32时标准栈偏差超0.065,新方法保持约0.035。代价是训练吞吐量降低2-3倍,团队建议将其作为调试工具而非生产默认。

nice rl experiment on train-inference mismatch

Yichuan WangZero Train-Inference Mismatch - now for linear attention, and under async RL 🎯 We got bitwise-exact trainer/generator parity for Gated DeltaNet (Qwen3.5-9B / 3...

来源:Nathan Lambert· x.com