HuggingFace Daily Papers(社区热门论文)
51AI 编辑部评分,满分 100

SPIEval:评估大语言模型作为移动助手处理分散个人信息的能力

2026-08-11 08:00· 1天前
AI 导读

SPIEval 是一个人工策划的基准,基于推理、消歧、整合、偏好推断和多意图分解五种认知能力,包含 250 个任务、覆盖 10 个应用中 4,335 条个人记录,并通过 21 个工具支持多轮交互。评测九个大语言模型后发现,最佳模型 GPT-5.5 (xhigh) 准确率仅 57.3%,最弱模型仅 16.4%;79% 的失败源于信息定位不准确,且不到 2% 的检索动作采用高级搜索方法。

Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address this gap, we introduce SPIEval, a human-curated benchmark grounded in five cognitive capabilities (i.e., reasoning, disambiguation, integration, preference inference, and multi-intent decomposition). SPIEval comprises 250 tasks spanning 4,335 personal records distributed across 10 apps and supports multi-turn interaction through 21 tools. Analysis shows that the benchmark exhibits diverse scenarios, challenging tasks, scattered information, controllable environments, and verifiable outcomes. We evaluate nine representative LLMs and find substantial room for improvement. The best-performing model, GPT-5.5 (xhigh), achieves only 57.3% accuracy, while the weakest achieves just 16.4%. Further analysis reveals that 79% of failures stem from inaccurate information localization, as LLMs often commit to plausible but incorrect information instead of continuing retrieval for verification. We also find that fewer than 2% of retrieval actions employ advanced search methods and observe substantial variation in search efficiency across models. These findings expose fundamental limitations of current LLM-based mobile assistants and motivate future research in this direction. Data and code are available at https://huggingface.co/datasets/Junjie-Ye/SPIEval.

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org

SPIEval:评估大语言模型作为移动助手处理分散个人信息的能力

HuggingFace Daily Papers(社区热门论文)·2026-08-11 08:00·1天前
AI 导读

SPIEval 是一个人工策划的基准,基于推理、消歧、整合、偏好推断和多意图分解五种认知能力,包含 250 个任务、覆盖 10 个应用中 4,335 条个人记录,并通过 21 个工具支持多轮交互。评测九个大语言模型后发现,最佳模型 GPT-5.5 (xhigh) 准确率仅 57.3%,最弱模型仅 16.4%;79% 的失败源于信息定位不准确,且不到 2% 的检索动作采用高级搜索方法。

原文 · 保持原样,未翻译

Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address this gap, we introduce SPIEval, a human-curated benchmark grounded in five cognitive capabilities (i.e., reasoning, disambiguation, integration, preference inference, and multi-intent decomposition). SPIEval comprises 250 tasks spanning 4,335 personal records distributed across 10 apps and supports multi-turn interaction through 21 tools. Analysis shows that the benchmark exhibits diverse scenarios, challenging tasks, scattered information, controllable environments, and verifiable outcomes. We evaluate nine representative LLMs and find substantial room for improvement. The best-performing model, GPT-5.5 (xhigh), achieves only 57.3% accuracy, while the weakest achieves just 16.4%. Further analysis reveals that 79% of failures stem from inaccurate information localization, as LLMs often commit to plausible but incorrect information instead of continuing retrieval for verification. We also find that fewer than 2% of retrieval actions employ advanced search methods and observe substantial variation in search efficiency across models. These findings expose fundamental limitations of current LLM-based mobile assistants and motivate future research in this direction. Data and code are available at https://huggingface.co/datasets/Junjie-Ye/SPIEval.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org