# Olmo 3 后训练与 DPO 实战案例解析

- 来源：Nathan Lambert (@natolambert)
- 发布时间：2026-07-29 23:18
- AIHOT 分数：56
- AIHOT 链接：https://aihot.virxact.com/items/cms68mhi61d63robk1as972hu
- 原文链接：https://x.com/natolambert/status/2082485850257207307

## AI 摘要

Nathan Lambert 与 Scott Geng 联合发布播客/讲座，以 Olmo 3 为案例，深入探讨 DPO 在近前沿模型后训练中的“混乱细节”。内容涵盖偏好数据缩放、Delta 学习假设、多阶段训练中的组织挑战，以及学术研究的未来方向。

## 正文

New podcast/lecture combo -- a case study in the messy details of Olmo 3 post training & DPO with @scottgeng00. It's rare to make time for these discussions， but we cover：

What it takes for a research idea to make it into a （near） frontier model.
The messy side of DPO （usually data）.
Organizational challenges in multi-stage post-training recipes.
Reflections on where research is heading today， and how people should think of DPO.
DPO fundamentals （once more）
Other topics.

Scott is one of the great student's I've had the pleasure of working with for a few years！

Chapters：

00：00 Introduction
04：32 Preference Tuning Basics
08：21 The DPO-Algorithm Zoo
17：06 Scaling Preference Data
21：55 Delta Learning Hypothesis
32：30： Applying it to Olmo 3
36：08 Challenge 1： Data Details
47：17 Challenge 2： Organization Management
51：45 Challenge 3： Scaling Issues
59：15 The Future of Academic Research

Excited to keep making this content under the motivation of my book. Really plugging away now ：）
