Rohan Paul@rohanpaul_ai
40AI 编辑部评分,满分 100
2026-08-05 23:43· 35分钟前
AI 导读

Claude Opus 5(Max)以1,699分登顶Fullstack Code Arena,大幅领先Kimi K3 Max。该基准要求模型通过多步规划和工具调用构建可运行的Web应用,涵盖文件编辑、命令执行、数据库、认证及外部API集成,并直接评测最终应用效果,更贴近真实编码智能体的工作方式。

Claude Opus 5 at Max effort leads Fullstack Code Arena with 1,699 points.

Significantly ahead of Kimi K3 Max

Fullstack Code Arena asks models to build working web apps through multi-step planning and tools, instead of answering isolated coding questions.

Models can create and edit files, run commands, use databases, authentication and external APIs, then produce a live application for testing.

Its specialty is testing models inside a complete tool-using build process and judging the finished app, which better resembles how coding agents work.

Arena.aiBig news: Opus 5 (Max) is now #1 in the Fullstack Code Arena with 1,699 points! The Fullstack Leaderboard shows overall rankings across AI models on full-stack ...

来源:Rohan Paul · x.com

Rohan Paul · @rohanpaul_ai · X·2026-08-05 23:43·35分钟前
AI 导读

Claude Opus 5(Max)以1,699分登顶Fullstack Code Arena,大幅领先Kimi K3 Max。该基准要求模型通过多步规划和工具调用构建可运行的Web应用,涵盖文件编辑、命令执行、数据库、认证及外部API集成,并直接评测最终应用效果,更贴近真实编码智能体的工作方式。

Claude Opus 5 at Max effort leads Fullstack Code Arena with 1,699 points.

Significantly ahead of Kimi K3 Max

Fullstack Code Arena asks models to build working web apps through multi-step planning and tools, instead of answering isolated coding questions.

Models can create and edit files, run commands, use databases, authentication and external APIs, then produce a live application for testing.

Its specialty is testing models inside a complete tool-using build process and judging the finished app, which better resembles how coding agents work.

Arena.aiBig news: Opus 5 (Max) is now #1 in the Fullstack Code Arena with 1,699 points! The Fullstack Leaderboard shows overall rankings across AI models on full-stack ...

来源:Rohan Paul· x.com