This is kind of wild.
BDH-CQ scored 29.5% on ARC-AGI 1. Fair. But with just $.0007 per task.
How it's done is particularly interesting. It reasons recurrently in latent space rather than using CoT. They've also verified Transformer-like scaling up to 600B params.