# Grok 4.6 在 DiligenceBench 金融评测中位列第二，与 Claude Opus 5 持平

- 来源：Karina (@karinanguyen)
- 发布时间：2026-08-19 05:04
- AIHOT 分数：54
- AIHOT 链接：https://aihot.virxact.com/items/cmsz6xfil04borodpt9eufv5p
- 原文链接：https://x.com/karinanguyen/status/2089820713830236327

## AI 摘要

Grok 4.6 在 DiligenceBench 金融评测中排名第二，得分约 52–53%，与 Claude Opus 5 基本持平。Grok 每任务调用工具 41 次（Opus 为 22 次），SEC 文件搜索 3,293 次，搜索更广但成本更低，约 $0.84/任务，而 Opus 5 约 $1.02。

## 正文

Grok 4.6 is a great cost-effective model for deep financial research.

It's #2 on DiligenceBench with the finance harness, effectively tied with Claude Opus 5 at ~52-53%.

A few interesting differences:
• Grok searched much more: 41 tool calls per task vs. 22 for Opus
• It made 3,293 SEC filing searches vs. 486
• That broader search helped Grok find more evidence and follow task-specific instructions more closely. Opus was more efficient and sometimes more nuanced
• Sonnet 5 trails both at 46.2%.
• Grok also gets there more cheaply: about $0.84/task vs. ~$1.02 for Opus 5, despite using far more search.
• On Vals' Finance Agent Benchmark v2 (FAB v2), Grok 4.6 leads the General Qualitative category, which is consistent with the kind of broad research and evidence-gathering DiligenceBench rewards
