Another win for Kimi K3 (Max)
Now it ranks 1st on Arena's fullstack coding test ahead of GPT-5.6 Sol (xHigh) and Claude Fable 5.
This benchmark goes beyond isolated code snippets by asking models to build working web applications across several connected steps.
Models must plan, edit files, run commands, connect databases, authentication and APIs, and produce a deployable application.
Human evaluators then compare the apps for functionality, usability and how closely they match the requested behaviour.
So to perform on this benchmark models need stronger coordination across frontend, backend and tool decisions during a complete build.