New Meta paper argues that factuality means both avoiding errors and including enough of the right information.
The hardest part of factuality may be measuring what an AI answer fails to include.
GAMUT makes missing information measurable, giving AI teams a clearer test of whether long answers actually finish the job.
Most factuality tests ask whether each stated claim is correct, but they rarely check what the answer leaves out.
The benchmark first builds a structured guide marking required facts, acceptable choices, ordered steps, relationships, and their importance.
It then converts that richer guide into small pass-or-fail checks, which an LLM judge can score more consistently.
Across 14 strong models, the best score was only 58.7%, and missing information caused more failures than false claims.
---
- arxiv. org/abs/2607.19322
Title: "Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness"