As models improve, the benchmarks we use to evaluate them have to evolve too.
"Create an SVG of a pelican riding a bicycle" was once surprisingly helpful. Today, frontier models routinely produce convincing results, so the test tells us increasingly little about where the actual frontier is.
Karpathys experiment is a fascinating attempt to push the benchmark forward: give Opus 5 the opening of The Lord of the Rings, a massive token budget, and two hours to turn it into an interactive Three.js world. And voila!
This tests far more than one-shot generation: The model has to maintain coherence over thousands of lines of code, translate prose into a spatial system, coordinate objects and animations, and continuously inspect its own work.
Really love where the testing / benchmarks are moving to!
Love to see GPT-5.6 and the new DeepSeek flash on this one.