Nathan Lambert@natolambert
50AI 编辑部评分,满分 100
2026-08-06 04:43· 1天前
AI 导读

Nathan Lambert 发布评测讲座,梳理其亲历的评测时代变迁:从 GPT-3 提示词式“高级自动补全”,到如今复杂的智能体沙盒环境。讲座重点剖析智能体评测(借鉴 @xeophon 观点),并探讨评测如何被操纵及其真实用途,涵盖后训练评测时代、智能体评测入门及“数字能否可信”等章节。

My evaluation lecture! I walk you through different evaluation eras I've been a part of, from prompting GPT-3 as elaborate autocomplete to today's complex agentic sandboxes (I expand on agentic more than any other topic, drawing on @xeophon's insights).

This lecture is a birds eye view of how evaluation has changed, how it can be gamed, and what it's actually used for.

00:00 Intro: frontier evaluation is harder than ever 03:39 Part 1: The eras of post-training evaluation 17:36 Part 2: An intro to agentic evals 21:19 Part 3: Can you trust the number? 30:41 Takeaways & conclusion

Thanks for watching! Just one more lecture after this :)

来源:Nathan Lambert · x.com

Nathan Lambert · @natolambert · X·2026-08-06 04:43·1天前
AI 导读

Nathan Lambert 发布评测讲座,梳理其亲历的评测时代变迁:从 GPT-3 提示词式“高级自动补全”,到如今复杂的智能体沙盒环境。讲座重点剖析智能体评测(借鉴 @xeophon 观点),并探讨评测如何被操纵及其真实用途,涵盖后训练评测时代、智能体评测入门及“数字能否可信”等章节。

My evaluation lecture! I walk you through different evaluation eras I've been a part of, from prompting GPT-3 as elaborate autocomplete to today's complex agentic sandboxes (I expand on agentic more than any other topic, drawing on @xeophon's insights).

This lecture is a birds eye view of how evaluation has changed, how it can be gamed, and what it's actually used for.

00:00 Intro: frontier evaluation is harder than ever 03:39 Part 1: The eras of post-training evaluation 17:36 Part 2: An intro to agentic evals 21:19 Part 3: Can you trust the number? 30:41 Takeaways & conclusion

Thanks for watching! Just one more lecture after this :)

来源:Nathan Lambert· x.com