GPT-6 Astra (Max) leads Code Arena: WebDev at 1,797, 35 points above Claude Fable 5.1 (Max).
This benchmark tests web-app building: models need to plan with tools, build a live app, and users compare paired outputs.
But more interestingly the huge gap GPT-6 Astra has with GPT-5.6 Sol (xHigh), its now at 180 points above GPT-5.6 Sol (xHigh), OpenAI's previous entry, which now ranks #13.
Real-world results are in. There is a new #1 on Code Arena - GPT-6 Astra (Max)! It also reshapes the Pareto frontier as the best-performing model at $40/Mtoken,...