Artificial Analysis 推出六个行业能力指数

Artificial Analysis · @ArtificialAnlys · X·2026-07-07 09:09·56天前
AI 导读

Artificial Analysis 发布六个新行业能力指数,涵盖金融与会计、法律、医疗、战略与运营、工程和经济学。每个指数基于 O*NET 职业常见任务独立运行基准测试。领先模型方面,Claude Fable 5(Opus 4.8 回退)在所有八个指数中领先,Claude Opus 4.8 (max) 在六个指数中排名第二,GPT-5.5 (xhigh) 在另外两个指数中第二。开放权重模型中,GLM-5.2 (max) 在五个行业指数领先,工程指数得分53,接近 Claude Sonnet 5 (max, 55) 和 GPT-5.5 (xhigh, 55);DeepSeek V4 Pro (max) 在战略与运营指数领先。成本方面,DeepSeek V4 Flash (max) 每个任务低于 $0.04,GLM-5.2 (max) 成本 $0.26–$0.58;Claude Fable 5 成本 $3.48,分数高出 DeepSeek V4 Pro 12 分但成本超百倍。时间方面,Nova 2.0 Pro Preview (medium) 最快 1.1 分钟,Claude Sonnet 5 (max) 最慢 16.7 分钟。

Artificial Analysis@ArtificialAnlys
51AI 编辑部评分,满分 100

Artificial Analysis 推出六个行业能力指数

2026-07-07 09:09· 56天前
AI 导读

Artificial Analysis 发布六个新行业能力指数,涵盖金融与会计、法律、医疗、战略与运营、工程和经济学。每个指数基于 O*NET 职业常见任务独立运行基准测试。领先模型方面,Claude Fable 5(Opus 4.8 回退)在所有八个指数中领先,Claude Opus 4.8 (max) 在六个指数中排名第二,GPT-5.5 (xhigh) 在另外两个指数中第二。开放权重模型中,GLM-5.2 (max) 在五个行业指数领先,工程指数得分53,接近 Claude Sonnet 5 (max, 55) 和 GPT-5.5 (xhigh, 55);DeepSeek V4 Pro (max) 在战略与运营指数领先。成本方面,DeepSeek V4 Flash (max) 每个任务低于 $0.04,GLM-5.2 (max) 成本 $0.26–$0.58;Claude Fable 5 成本 $3.48,分数高出 DeepSeek V4 Pro 12 分但成本超百倍。时间方面,Nova 2.0 Pro Preview (medium) 最快 1.1 分钟,Claude Sonnet 5 (max) 最慢 16.7 分钟。

Introducing six new Artificial Analysis Capability Indices for comparing model capabilities across key industry domains

The new industry indices cover Finance & Accounting, Legal, Healthcare & Medical, Strategy & Ops, Engineering, and Economics. We aim to capture the common capabilities required across knowledge work domains and evaluate how well current models meet those needs.

Each index is grounded in common tasks from O*NET occupational classifications. Tasks range from financial modeling, to legal research and contract review, to clinical decision support and patient documentation. We derive capabilities from each task, select the benchmarks that best represent the work, and weight by how often each capability appears across the domain. This means rethinking the Artificial Analysis benchmark suite for each domain and slicing evaluations to relevant domain tasks. Every component benchmark is run independently by Artificial Analysis.

The industry indices join the existing skill-based Agentic and Coding indices, which measure capabilities that cut across every domain.

Key Results ➤ Leading models: Claude Fable 5 (with Opus 4.8 fallback) leads all eight indices, with Claude Opus 4.8 (max) in second on six of eight Capability Indices and GPT-5.5 (xhigh) on two. Below the top two, rankings reshuffle substantially by domain between Gemini 3.5 Flash, Gemini 3.1 Pro Preview, GPT-5.5 (xhigh), Claude Sonnet 5 (max), and GLM-5.2 (max). ➤ Open weights leading models: Among open weights models, GLM-5.2 (max) leads on five of the six industry indices, ranking as high as fifth overall on the Artificial Analysis Engineering Index (53), within 2 points of Claude Sonnet 5 (max, 55) and GPT-5.5 (xhigh, 55). DeepSeek V4 Pro (max, 38) takes the open weights lead on Artificial Analysis Strategy & Ops Index. ➤ Cost efficiency: DeepSeek V4 Flash (max) completes tasks for <$0.04 across all six indices while scoring mid-pack, and GLM-5.2 (max) leads open weights score with a Cost per Task of $0.26 to $0.58. Frontier capability comes at a steep premium: on the Artificial Analysis Strategy & Ops Index, Claude Fable 5 (with Opus 4.8 fallback, $3.48) scores 12 points above DeepSeek V4 Pro (max, $0.03) at over 100x the Cost per Task. ➤ Time per Task: Time per Task spreads roughly 15x within each index, from 1.1 minutes for Nova 2.0 Pro Preview (medium) to 16.7 minutes for Claude Sonnet 5 (max). Speed shows a similar frontier to cost: on the Artificial Analysis Legal Index, Gemini 3.1 Pro Preview (0.8 minutes) completes tasks ~7x faster than Claude Fable 5 (with Opus 4.8 fallback, 5.4 minutes), while scoring within 11 points.

来源:Artificial Analysis· x.com