The final lecture of my course is an intro to character training! This is a topic that I've been quietly very invested in for ~18 months, as it:
• Has potential for high real world impact • Clearly used extensively at frontier labs • Almost no empirical literature exists • More accessible on academic compute
This lecture covers what character training is, reviews model specs, constitutions, the differences, the motivations in real world events, some example research papers I like, and open questions in how it relates to post-training/model use generally.
Hopefully this brings more people into the field (and reach out if you have questions). It is one of the more research-y chapters in my book, but one that I felt needed the reference. There is still so little, educational content on the topic online.
0:00 Intro 6:22 Part 1: Fundamentals - character, constitutions, and model specs 19:21 Part 2: Character training in practice 23:23 Part 3: Character elicitation without gradient steps 28:03 Part 4: Open questions (and the end of the course) 32:27 The course, complete
Thanks for watching. No need to like and subscribe now that the course is done, you definitely wouldn't!
h/t to @_maiush for leading the technical work I got to do in the space, and @zafstojano for investing a lot of attention at this book chapter.