Highly recommended.
I've often claimed there's huge alpha in building agent harnesses.
Turns out harnesses are compositional generalizers. The RLM harness is an instance of this. This could lead to interesting and efficient new ways to scale generalization in models.
Transformers struggle to generalize to tasks they were not explicitly trained on. Instead, we propose in 2026 that it is the job of the harness to generalize th...