# Kimi K3 登顶 SpreadsheetBench 2，超越 Claude Fable 5

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-07-18 13:33
- AIHOT 分数：40
- AIHOT 链接：https://aihot.virxact.com/items/cmrpymvkp06tcbisrb2wumndk
- 原文链接：https://x.com/rohanpaul_ai/status/2078352555944689722

## AI 摘要

Kimi K3 在 SpreadsheetBench 2 上排名第一，超越 Claude Fable 5，成为首个超越所有闭源模型的开源权重模型。该基准测试包含 321 个专家任务，涵盖财务建模、工作簿调试和原生图表创建，平均每个任务涉及 11.8 个工作表和 593.5 次单元格变更。Kimi K3 完成了 34.8% 的工作流，在金融、规划和运营场景中表现出色。

## 正文

Kimi K3 may be super useful for finance, planning and operations, given how well it plays with spreadsheets.

Took the top spot on SpreadsheetBench 2, almost tallying with Fable 5 and completing 34.8% of workflows.

SpreadsheetBench 2 tests whether tool-using AI agents can finish entire spreadsheet jobs rather than isolated formulas.

It measures autonomous workbook execution rather than general reasoning or everyday Excel assistance.

It involves 321 expert-curated tasks spanning financial modeling, workbook debugging, and native chart creation.

Each task averages 11.8 sheets and 593.5 cell changes, forcing long chains of dependent actions.

Modeling and debugging tasks require every requested edit to match while untouched cells remain unchanged.

Chart tasks pass only when a vision model confirms at least 70% of specified requirements.

### 引用推文

> AfterQuery：Kimi K3 ranks #1 on @AfterQuery's SpreadsheetBench 2, surpassing Claude Fable 5. An open weight model now outperforms all closed-sourced models. Read more in th...
