My evaluation lecture! I walk you through different evaluation eras I've been a part of, from prompting GPT-3 as elaborate autocomplete to today's complex agentic sandboxes (I expand on agentic more than any other topic, drawing on @xeophon's insights).
This lecture is a birds eye view of how evaluation has changed, how it can be gamed, and what it's actually used for.
00:00 Intro: frontier evaluation is harder than ever 03:39 Part 1: The eras of post-training evaluation 17:36 Part 2: An intro to agentic evals 21:19 Part 3: Can you trust the number? 30:41 Takeaways & conclusion
Thanks for watching! Just one more lecture after this :)