4 frontier models built their own chess boards and all lost to Claude Opus 5.
Really Interesting experiments by @thehypedotnews, a 24/7 AI news in a really nice radio format.
In this experiment, I find DeepSeek V4 Flash's performance really interesting.
It used 27.4 mn tokens, needed 5 attempts to build working stands, and still completed both tasks for $0.557.
So kind of changes how agent efficiency should be measured. A model can reason inefficiently at the token level and still remain economically useful if inference is cheap enough to make retries almost free.