内容
精选全部 AI 动态AI 日报主题收藏
接入
Agent 接入
更多
关于更新日志反馈
原文
Nathan Lambert@natolambert
66
2026-07-21 22:11· 3小时前
跳到正文
AI 摘要

Nathan Lambert 宣布其著作《Reinforcement Learning from Human Feedback》已完成。该书旨在传授微调、对齐及后训练模型的基础知识,并附带超过 10 小时的完整课程、幻灯片、训练章节的功能代码及示例模型补全库。实体书将于 1-2 周内从 Manning 发货,亚马逊稍晚上架。

My book, Reinforcement Learning from Human Feedback is done!

This is the book I wish I had when learning to fine-tune, align, & now post-train models since ChatGPT. The resource has been built by me finding time to study and document the fundamentals on nights and weekends since 2024.

Transferring as much of the intuitions of building Olmo as I possibly can in the book format.

The book is launching with an over 10 hour, full course with slidedecks, functional code for the training chapters, an example model completions library, and of course the free online web version.

Physical orders from Manning will ship in 1-2 weeks, and Amazon a week or so after. Thanks for your support!

安全/对齐教程/实践
在 X 查看原推
Nathan Lambert@natolambert · X
66导出 Markdown
2026-07-21 22:11·3小时前
在 X 看原推· x.com
AI 摘要

Nathan Lambert 宣布其著作《Reinforcement Learning from Human Feedback》已完成。该书旨在传授微调、对齐及后训练模型的基础知识,并附带超过 10 小时的完整课程、幻灯片、训练章节的功能代码及示例模型补全库。实体书将于 1-2 周内从 Manning 发货,亚马逊稍晚上架。

My book, Reinforcement Learning from Human Feedback is done!

This is the book I wish I had when learning to fine-tune, align, & now post-train models since ChatGPT. The resource has been built by me finding time to study and document the fundamentals on nights and weekends since 2024.

Transferring as much of the intuitions of building Olmo as I possibly can in the book format.

The book is launching with an over 10 hour, full course with slidedecks, functional code for the training chapters, an example model completions library, and of course the free online web version.

Physical orders from Manning will ship in 1-2 weeks, and Amazon a week or so after. Thanks for your support!

导出 Markdown
安全/对齐教程/实践
在 X 查看原推x.com