# Grok 4.5 在 Zapier 自动化基准测试中登顶，成本仅为竞品四分之一

- 来源：Elon Musk (@elonmusk)
- 发布时间：2026-07-09 14:42
- AIHOT 分数：59
- AIHOT 链接：https://aihot.virxact.com/items/cmrd5kzjb00r4ih4b8b3sab36
- 原文链接：https://x.com/elonmusk/status/2075108230754116018

## AI 摘要

xAI 的 Grok 4.5 在 Zapier 的 AutomationBench-AA 基准测试中取得 51% 分数，排名第一，超越 Claude Fable 5（49%）和 Claude Opus 4.8（48%），每任务成本仅 $0.34，约为竞品四分之一。Grok 4.5 完成 79.9% 的任务目标，严格通过 21.9% 的任务；输出 token 极高效，每任务约 8k，远低于 Claude Opus 4.8 的 32k。在难度最高的金融领域完成 71% 目标，领先所有模型。但每任务违规次数 0.63，高于 Gemini 3.5 Flash 的 0.46。测试涵盖 657 个任务、40 个模拟 SaaS 环境。

## 正文

Cool that Grok 4.5 is #1 in some respects, even with respect to Fable 5

### 引用推文

> Artificial Analysis：SpaceXAI's Grok 4.5 takes the #1 spot on AutomationBench-AA with a score of 51%, ahead of Claude Fable 5 (49%) and Claude Opus 4.8 (48%) at roughly a quarter of...
