Artificial Analysis@ArtificialAnlys
62AI 编辑部评分,满分 100

Grok 4.6 登顶 AA-Briefcase,成本远低于竞品

2026-08-13 01:55· 38分钟前
AI 导读

Grok 4.6 在 Artificial Analysis 的 AA-Briefcase 智能体知识工作基准上取得大幅提升,与 Claude Fable 5 并列领先(置信区间重叠)。其单任务成本仅 $4.42,远低于 Claude Fable 5 的 $22.30、Claude Opus 5 的 $17.79 和 Kimi K3 的 $6.73。该测试集为私有,可防止数据污染。

Grok 4.6 made large gains on AA-Briefcase, our agentic knowledge work benchmark and cost substantially less than other leading models

AA-Briefcase tests models on long-horizon agentic knowledge work tasks. The test set is private to prevent contamination.

Grok 4.6 is neck and neck with Claude Fable 5, with overlapping confidence intervals. The model is also substantially cheaper than other leading models on the benchmark at a Cost per Task of $4.42 compared to Claude Fable 5's $22.30, Claude Opus 5's $17.79 and Kimi K3's $6.73.

Impressive release @SpaceXAI and @elonmusk.

来源:Artificial Analysis · x.com

Grok 4.6 登顶 AA-Briefcase,成本远低于竞品

Artificial Analysis · @ArtificialAnlys · X·2026-08-13 01:55·38分钟前
AI 导读

Grok 4.6 在 Artificial Analysis 的 AA-Briefcase 智能体知识工作基准上取得大幅提升,与 Claude Fable 5 并列领先(置信区间重叠)。其单任务成本仅 $4.42,远低于 Claude Fable 5 的 $22.30、Claude Opus 5 的 $17.79 和 Kimi K3 的 $6.73。该测试集为私有,可防止数据污染。

Grok 4.6 made large gains on AA-Briefcase, our agentic knowledge work benchmark and cost substantially less than other leading models

AA-Briefcase tests models on long-horizon agentic knowledge work tasks. The test set is private to prevent contamination.

Grok 4.6 is neck and neck with Claude Fable 5, with overlapping confidence intervals. The model is also substantially cheaper than other leading models on the benchmark at a Cost per Task of $4.42 compared to Claude Fable 5's $22.30, Claude Opus 5's $17.79 and Kimi K3's $6.73.

Impressive release @SpaceXAI and @elonmusk.

来源:Artificial Analysis· x.com