# Artificial Analysis 与 Zapier 联合发布 AutomationBench-AA 独立排行榜

- 来源：Artificial Analysis (@ArtificialAnlys)
- 发布时间：2026-07-07 02:12
- AIHOT 分数：60
- AIHOT 链接：https://aihot.virxact.com/items/cmr9jnvr700e3ihe88a65y255
- 原文链接：https://x.com/ArtificialAnlys/status/2074194764510208230

## AI 摘要

Artificial Analysis 与 Zapier 合作推出 AutomationBench-AA 排行榜，测试 AI 智能体在真实 SaaS 工作流中遵循业务规则的自动化能力。基准包含 657 项任务，覆盖财务、HR、销售等 6 部门，在 40 个模拟应用（Gmail、Slack、Salesforce 等）中运行。模型通过 REST API 自主发现端点，按近 12,000 条断言评分（目标达成+不违反护栏）。结果：Claude Fable 5 (max) 以 48.6% 领先，Opus 4.8 以 48.5% 紧随其后，Gemini 3.5 Flash 为 42.6%（$0.49/任务），GPT-5.5 (xhigh) 为 42.1%（$1.32/任务）。Fable 5 在约 18% 的任务中回退到 Opus。开放权重模型 GLM-5.2 (max) 最佳（27.8%）。所有模型均违反业务规则，金融任务难度最高。

## 正文

Announcing AutomationBench-AA, our independent leaderboard for Zapier’s AutomationBench, testing whether AI agents can automate real SaaS workflows while adhering to business rules

We partnered with @zapier to run AutomationBench-AA on their private benchmark subset. This benchmark is a complex agentic workflow automation test across simulated SaaS applications. Models must complete 657 tasks spanning Finance, HR, Marketing, Operations, Sales, and Support, working across 40 simulated app environments including Gmail, Google Sheets, Slack, Salesforce, Zendesk, Jira, and HubSpot.

Unlike Zapier’s hosted leaderboard, the headline score for AutomationBench-AA shows the share of objectives a model completes without violating any guardrails. Claude Fable 5 and Opus 4.8 from @AnthropicAI lead with scores of 48.6% and 48.5%, followed by @GoogleDeepMind's Gemini 3.5 Flash at 42.6% and @OpenAI's GPT-5.5 (xhigh) at 42.1%. With Anthropic’s new classifier, Fable 5 fell back to Opus on ~18% of tasks.

Key elements of AutomationBench:

➤ Real workflow patterns, simulated environments: Tasks are drawn from real workflow patterns on Zapier and run in simulated SaaS environments, where a single task may span a range of applications like CRM, email, calendar, and messaging platforms.

➤ Autonomous API discovery: Models interact with each app through REST APIs, discovering the endpoints they need through structured tool calls and navigating environments with irrelevant and sometimes misleading records.

➤ Objectives and guardrails: Models are scored against nearly 12,000 assertions Zapier built to test that the model completed the task correctly in full. Each assertion is classified as either an objective the agent must achieve, or a guardrail that already passes initially and must not be broken.

➤ Programmatic environment grading: Tasks are graded solely on whether the correct data ended up in the right systems, with deterministic checks against the environment. Each task runs once with a 50-turn cap.

Key results for AutomationBench-AA:

➤ Claude Fable 5 (max) leads at 48.6% but falls back to Opus 4.8 in ~18% of tasks. It completes 73% of task objectives, with the fallback behavior likely explaining the limited uplift compared to Opus.

➤ Every model breaks business rules: Guardrail violations range from 0.46 per task (Gemini 3.5 Flash) to 1.26 (Qwen3.7 Plus). Gemini 3.5 Flash completes 15.0 objectives per guardrail violation, the best ratio of any model, ahead of Claude Opus 4.8 (max, 13.5).

➤ Gemini 3.5 Flash performs well for its price: At 42.6% and $0.49 per task, it effectively matches GPT-5.5 (xhigh, 42.1%, $1.32 per task) at ~37% of the cost.

➤ GLM-5.2 (max) from @Zai_org is the leading open weights model at 27.8%. This places the open weights frontier ~10 points behind Gemini 3.1 Pro Preview, and with substantially higher guardrail violations per task.

➤ Finance workflow tasks are the most difficult to automate today: across the models we evaluated at launch, agents complete roughly half the proportion of objectives on Finance tasks, compared to Support and Operations tasks.

We would like to thank Zapier and the benchmark authors for their great work developing this evaluation for important SaaS workflows, and appreciate their collaboration in launching AutomationBench on Artificial Analysis!
