Kimi K3 may be super useful for finance, planning and operations, given how well it plays with spreadsheets.
Took the top spot on SpreadsheetBench 2, almost tallying with Fable 5 and completing 34.8% of workflows.
SpreadsheetBench 2 tests whether tool-using AI agents can finish entire spreadsheet jobs rather than isolated formulas.
It measures autonomous workbook execution rather than general reasoning or everyday Excel assistance.
It involves 321 expert-curated tasks spanning financial modeling, workbook debugging, and native chart creation.
Each task averages 11.8 sheets and 593.5 cell changes, forcing long chains of dependent actions.
Modeling and debugging tasks require every requested edit to match while untouched cells remain unchanged.
Chart tasks pass only when a vision model confirms at least 70% of specified requirements.