swyx · @swyx · X·2026-08-18 00:45·19天前
AI 导读

swyx 称赞 Trajectory 在持续学习赛道上的品味与执行,称其为该领域早期领先者之一。Trajectory 的 rronak_ 在 aiDotEngineer World Fair 上阐述了其解决持续学习主要数据问题的方法,包括为何 GRPO 不够用、必须转向 on-policy 并修复随之而来的问题,并分享了扩展 SDPO 等算法的经验。

swyx@swyx
26AI 编辑部评分,满分 100
2026-08-18 00:45· 19天前
AI 导读

swyx 称赞 Trajectory 在持续学习赛道上的品味与执行,称其为该领域早期领先者之一。Trajectory 的 rronak_ 在 aiDotEngineer World Fair 上阐述了其解决持续学习主要数据问题的方法,包括为何 GRPO 不够用、必须转向 on-policy 并修复随之而来的问题,并分享了扩展 SDPO 等算法的经验。

Trajectory have generally impressed me with their tasteful execution on ambitious goals. on our Continual Learning track, @rronak_ gave a very thoughtful overview on how they're tackling the main data problems left in CL, including why GRPO isn't enough and they had to go on-policy.... and then subsequently fix all the issues that come up with it

nice overview from one of the early leaders in this field!

(see the rest of the track for more, this one was quite stacked)

TrajectoryIt's time to rethink RL. Translating real world use into model improvements requires redesigning post-training algorithms for non-verifiable, per token rewards....