Artificial Analysis@ArtificialAnlys
60AI 编辑部评分,满分 100

Grok 4.6 登顶智能前沿,智能体性能亮眼

2026-08-12 23:39· 54分钟前
AI 导读

SpaceXAI 的 Grok 4.6 在 Artificial Analysis 智能指数上得分 61,与 GPT-5.6 Sol 并列前沿,仅次于 Anthropic 的 Claude 系列。该模型智能体性能突出,GDPval-AA v2 Elo 达 1753,且定价不变($2/$6 每百万 token),成本较 Claude Opus 5 和 GPT-5.6 Sol 低 60% 以上。

SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier in line with GPT-5.6 Sol, with standout agentic performance at lower cost

Grok 4.6 gains 5 points over Grok 4.5 on the Intelligence Index just over one month after its release, or +23 points compared to Grok 4.3. This brings SpaceXAI back to the intelligence frontier alongside OpenAI, behind only Anthropic.

Key takeaways: ➤ Grok 4.6 joins the frontier of the Artificial Analysis Intelligence Index: It scores 61, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62), and just ahead of Kimi K3

➤ Strong agentic performance: Grok 4.6 achieves a GDPval-AA v2 Elo of 1753, behind only Claude Opus 5 and with overlapping confidence intervals with Claude Fable 5 and Qwen3.8 Max. It scores 50.7% on τ3-Banking, among the top two scores alongside Qwen3.8 Max (51.3%), and 88.4% on Terminal-Bench v2.1, in line with the leading models

➤ Frontier-level intelligence at lower cost: Headline pricing is unchanged from Grok 4.5 at $2/$6 per 1M input/output tokens, 60%+ below Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). It cost $0.84 per task, the same as Kimi K3 with slightly higher intelligence, placing it on the Intelligence vs. Cost per Task Pareto frontier

➤ Grok 4.6 sits at Fable 5-tier on AA-Briefcase, our private benchmark of long-horizon agentic knowledge work tasks, with an Elo of 1577 - behind the Claude Opus 5 family. It is notably turn-efficient, completing tasks in ~53 turns and ~0.5B input tokens on average vs. ~103 turns and ~2.0B input tokens for Claude Opus 5 (max)

Other model details: ➤ Context window of 500k tokens (unchanged from Grok 4.5)

➤ Pricing of $2/$6 per 1M tokens of input/output; cache hits discounted to $0.5 per 1M tokens, an increase over Grok 4.5's $0.3 per 1M tokens for cache hits

Congratulations to @SpaceXAI and @elonmusk on the release!

来源:Artificial Analysis · x.com

同一事件 · 1

Grok 4.6 登顶智能前沿,智能体性能亮眼

Artificial Analysis · @ArtificialAnlys · X·2026-08-12 23:39·54分钟前
AI 导读

SpaceXAI 的 Grok 4.6 在 Artificial Analysis 智能指数上得分 61,与 GPT-5.6 Sol 并列前沿,仅次于 Anthropic 的 Claude 系列。该模型智能体性能突出,GDPval-AA v2 Elo 达 1753,且定价不变($2/$6 每百万 token),成本较 Claude Opus 5 和 GPT-5.6 Sol 低 60% 以上。

SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier in line with GPT-5.6 Sol, with standout agentic performance at lower cost

Grok 4.6 gains 5 points over Grok 4.5 on the Intelligence Index just over one month after its release, or +23 points compared to Grok 4.3. This brings SpaceXAI back to the intelligence frontier alongside OpenAI, behind only Anthropic.

Key takeaways: ➤ Grok 4.6 joins the frontier of the Artificial Analysis Intelligence Index: It scores 61, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62), and just ahead of Kimi K3

➤ Strong agentic performance: Grok 4.6 achieves a GDPval-AA v2 Elo of 1753, behind only Claude Opus 5 and with overlapping confidence intervals with Claude Fable 5 and Qwen3.8 Max. It scores 50.7% on τ3-Banking, among the top two scores alongside Qwen3.8 Max (51.3%), and 88.4% on Terminal-Bench v2.1, in line with the leading models

➤ Frontier-level intelligence at lower cost: Headline pricing is unchanged from Grok 4.5 at $2/$6 per 1M input/output tokens, 60%+ below Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). It cost $0.84 per task, the same as Kimi K3 with slightly higher intelligence, placing it on the Intelligence vs. Cost per Task Pareto frontier

➤ Grok 4.6 sits at Fable 5-tier on AA-Briefcase, our private benchmark of long-horizon agentic knowledge work tasks, with an Elo of 1577 - behind the Claude Opus 5 family. It is notably turn-efficient, completing tasks in ~53 turns and ~0.5B input tokens on average vs. ~103 turns and ~2.0B input tokens for Claude Opus 5 (max)

Other model details: ➤ Context window of 500k tokens (unchanged from Grok 4.5)

➤ Pricing of $2/$6 per 1M tokens of input/output; cache hits discounted to $0.5 per 1M tokens, an increase over Grok 4.5's $0.3 per 1M tokens for cache hits

Congratulations to @SpaceXAI and @elonmusk on the release!

来源:Artificial Analysis· x.com

同一事件 · 1