Chubby♨️ · @kimmonismus · X·2026-07-07 18:56·55天前
AI 导读

Artificial Analysis 发布六项新行业能力指数(金融与会计、法律、医疗、战略与运营、工程、经济),连同已有的 Agentic 和 Coding 指数共八项。Claude Fable 5(Opus 4.8 回退)在全部八项指数上领先,Claude Opus 4.8(max)在六项中居第二,GPT-5.5(xhigh)在两项中排第二。开源模型方面,GLM-5.2(max)在五项行业指数中居首,工程指数得分 53 接近 Claude Sonnet 5(max)的 55 和 GPT-5.5(xhigh)的 55;DeepSeek V4 Pro(max)在战略与运营指数上领跑。成本方面,DeepSeek V4 Flash(max)六项任务总成本低于 $0.04,而 Claude Fable 5 在战略与运营指数上每任务 $3.48,得分高 12 分但成本超 100 倍。

Chubby♨️@kimmonismus
45AI 编辑部评分,满分 100
2026-07-07 18:56· 55天前
AI 导读

Artificial Analysis 发布六项新行业能力指数(金融与会计、法律、医疗、战略与运营、工程、经济),连同已有的 Agentic 和 Coding 指数共八项。Claude Fable 5(Opus 4.8 回退)在全部八项指数上领先,Claude Opus 4.8(max)在六项中居第二,GPT-5.5(xhigh)在两项中排第二。开源模型方面,GLM-5.2(max)在五项行业指数中居首,工程指数得分 53 接近 Claude Sonnet 5(max)的 55 和 GPT-5.5(xhigh)的 55;DeepSeek V4 Pro(max)在战略与运营指数上领跑。成本方面,DeepSeek V4 Flash(max)六项任务总成本低于 $0.04,而 Claude Fable 5 在战略与运营指数上每任务 $3.48,得分高 12 分但成本超 100 倍。

tl;dr: Fable 5 basically mocks every other model and tops every benchmark.

Curious to see whether the six new benchmarks will crown a new leader after GPT-5.6.

Artificial AnalysisIntroducing six new Artificial Analysis Capability Indices for comparing model capabilities across key industry domains The new industry indices cover Finance &...

来源:Chubby♨️· x.com