Grok 4.6 made large gains on AA-Briefcase, our agentic knowledge work benchmark and cost substantially less than other leading models
AA-Briefcase tests models on long-horizon agentic knowledge work tasks. The test set is private to prevent contamination.
Grok 4.6 is neck and neck with Claude Fable 5, with overlapping confidence intervals. The model is also substantially cheaper than other leading models on the benchmark at a Cost per Task of $4.42 compared to Claude Fable 5's $22.30, Claude Opus 5's $17.79 and Kimi K3's $6.73.
Impressive release @SpaceXAI and @elonmusk.