🦄New Ep alert - If you've been following along, I'm obsessed with a new benchmark out of UW called SlopCodeBench - the best models in the world (Fable, Sol, Kimi K3) top out at a 33% pass rate.
What makes SCB cool is that it forces a model to evolve a codebase over time
AND it measures how that code grows in complexity as the model tries to solve progressively harder problems
I sat down with @vaibcode to go deep on how the benchmark works and my findings from the latest family of frontier models.
we do 🦄 AI that works live every tuesday at 10:15am PT - come hang for the next one!