Apple Machine Learning Research(RSS)
40AI 编辑部评分,满分 100

DeepAmbigQA:用于评测 LLM 答案完整性的歧义多跳问答基准

2026-08-06 08:00· 1天前
AI 导读

苹果研究团队推出 DeepAmbigQA,一个用于评测大语言模型答案完整性的歧义多跳问答基准。该基准聚焦“同名电影《盗火线》中哪位演员获得过奥斯卡奖”这类复杂问题,要求模型同时区分同名实体并跨大量实体进行多跳推理。研究团队还发布了自动数据生成流程 DEEPAMBIGQAGEN,以构建此类评测数据。

Large language models (LLMs) with integrated search tools show strong promise in open-domain question answering (QA), yet they often struggle to produce complete answer set to complex questions such as “Which actor from the film Heat won at least one Academy Award?”, which requires (1) distinguishing between multiple films sharing the same title and (2) reasoning across a large set of actors to gather and integrate evidence. Existing QA benchmarks rarely evaluate both challenges jointly. To address this, we introduce DEEPAMBIGQAGEN, an automatic data generation pipeline that constructs QA…

来源:Apple Machine Learning Research(RSS) · machinelearning.apple.com

DeepAmbigQA:用于评测 LLM 答案完整性的歧义多跳问答基准

Apple Machine Learning Research(RSS)·2026-08-06 08:00·1天前
AI 导读

苹果研究团队推出 DeepAmbigQA,一个用于评测大语言模型答案完整性的歧义多跳问答基准。该基准聚焦“同名电影《盗火线》中哪位演员获得过奥斯卡奖”这类复杂问题,要求模型同时区分同名实体并跨大量实体进行多跳推理。研究团队还发布了自动数据生成流程 DEEPAMBIGQAGEN,以构建此类评测数据。

原文 · 保持原样,未翻译

Large language models (LLMs) with integrated search tools show strong promise in open-domain question answering (QA), yet they often struggle to produce complete answer set to complex questions such as “Which actor from the film Heat won at least one Academy Award?”, which requires (1) distinguishing between multiple films sharing the same title and (2) reasoning across a large set of actors to gather and integrate evidence. Existing QA benchmarks rarely evaluate both challenges jointly. To address this, we introduce DEEPAMBIGQAGEN, an automatic data generation pipeline that constructs QA…

来源:Apple Machine Learning Research(RSS)· machinelearning.apple.com