Rohan Paul@rohanpaul_ai
49AI 编辑部评分,满分 100

Grok 4.6 发布:智能体任务领先 GPT-5.6 Sol Max

2026-08-13 01:39· 16小时前
AI 导读

Grok 4.6 发布,主打专业智能体任务,在 GDPVal-AA v2 和 AA-Briefcase 上分别取得 1753 和 1577 分,均领先 GPT-5.6 Sol Max 与 Fable 5 Max。

Grok 4.6 just dropped.

Clearest strength is professional agent work, with leading results against GPT Sol Max and Fable 5 Max on GDPVal-AA v2 and AA-Briefcase.

• matches GPT-5.6 Sol Max at 61 on Artificial Analysis while charging $2/$6 per million input/output tokens.

• 1753 on GDPVal-AA v2 and 1577 on AA-Briefcase, both ahead of Sol Max and Fable 5 Max. Those benchmarks measure real-world agent tasks and agentic knowledge work, making them closer to research, analysis and multi-file deliverables than isolated question answering.

On coding, 69.9% on CursorBench v3.2 beats Sol's 67.2%, while DeepSWE and Terminal-Bench leave Grok behind both Sol and Fable.

SpaceXAI attributes the jump to a longer training run, regenerated SFT trajectories, model-based trace filtering, and agentic RL across coding, web development, CAD and kernel optimization.

It also reports more self-testing on long trajectories, with the model checking its work before continuing, directly targeting error accumulation across multi-step agents.

SpaceXAIIntroducing Grok 4.6. It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price.

来源:Rohan Paul · x.com

Grok 4.6 发布:智能体任务领先 GPT-5.6 Sol Max

Rohan Paul · @rohanpaul_ai · X·2026-08-13 01:39·16小时前
AI 导读

Grok 4.6 发布,主打专业智能体任务,在 GDPVal-AA v2 和 AA-Briefcase 上分别取得 1753 和 1577 分,均领先 GPT-5.6 Sol Max 与 Fable 5 Max。

Grok 4.6 just dropped.

Clearest strength is professional agent work, with leading results against GPT Sol Max and Fable 5 Max on GDPVal-AA v2 and AA-Briefcase.

• matches GPT-5.6 Sol Max at 61 on Artificial Analysis while charging $2/$6 per million input/output tokens.

• 1753 on GDPVal-AA v2 and 1577 on AA-Briefcase, both ahead of Sol Max and Fable 5 Max. Those benchmarks measure real-world agent tasks and agentic knowledge work, making them closer to research, analysis and multi-file deliverables than isolated question answering.

On coding, 69.9% on CursorBench v3.2 beats Sol's 67.2%, while DeepSWE and Terminal-Bench leave Grok behind both Sol and Fable.

SpaceXAI attributes the jump to a longer training run, regenerated SFT trajectories, model-based trace filtering, and agentic RL across coding, web development, CAD and kernel optimization.

It also reports more self-testing on long trajectories, with the model checking its work before continuing, directly targeting error accumulation across multi-step agents.

SpaceXAIIntroducing Grok 4.6. It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price.

来源:Rohan Paul· x.com