Open weights become much more consequential once the model can reliably operate tools.
And now Kimi K3 (Max) leads open-weight models in Agent Arena with a +9.75% net-improvement score.
Agent Arena tests if selecting K3 as the orchestrator improve outcomes when the model has to choose tools and sustain a real workflow
Kimi K3 ranks 3rd across all 42 models (open and close), behind Claude Fable 5 (High) and GPT 5.6 Sol (xHigh), while GLM 5.2 (Max) is the next open-weight model at +7.12%.
Arena randomly assigns sessions to models, allowing it to estimate each orchestrator's causal effect against an average model.
The combined score averages five signals: confirmed success, user reactions, correction handling, bash recovery, and nonexistent-tool calls.