# Meta 新基准 GAMUT 衡量 AI 回答的事实完整性

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-07-24 19:44
- AIHOT 分数：35
- AIHOT 链接：https://aihot.virxact.com/items/cmryvt0q505mrrolg8oaaij0t
- 原文链接：https://x.com/rohanpaul_ai/status/2080620006526652699

## AI 摘要

Meta 新论文提出，事实性不仅包括避免错误，还需包含足够的关键信息。新基准 GAMUT 通过构建结构化指南，将缺失信息转化为可衡量的 pass/fail 检查。在 14 个强模型测试中，最佳得分仅为 58.7%，信息缺失导致的失败比错误陈述更多。

## 正文

New Meta paper argues that factuality means both avoiding errors and including enough of the right information.

The hardest part of factuality may be measuring what an AI answer fails to include.

GAMUT makes missing information measurable， giving AI teams a clearer test of whether long answers actually finish the job.

Most factuality tests ask whether each stated claim is correct， but they rarely check what the answer leaves out.

The benchmark first builds a structured guide marking required facts， acceptable choices， ordered steps， relationships， and their importance.

It then converts that richer guide into small pass-or-fail checks， which an LLM judge can score more consistently.

Across 14 strong models， the best score was only 58.7%， and missing information caused more failures than false claims.

---

- arxiv. org/abs/2607.19322

Title： "Two-Level Meta-Rubrics for Evaluating Open-Ended Generation： GAMUT， a Benchmark for Factual Completeness"
