# DataSpace 基准：智能体框架选择可提升 15.36 点准确率

- 来源：elvis (@omarsar0)
- 发布时间：2026-08-06 03:15
- AIHOT 分数：66
- AIHOT 链接：https://aihot.virxact.com/items/cmsghqgks05xuro5qnvs5hdl5
- 原文链接：https://x.com/omarsar0/status/2085082167579902233

## AI 摘要

新研究发布 DataSpace 基准，用于评估数据智能体在异构工作区中生成可验证表格结果的能力，涵盖 410 个跨语言任务、7,439 个工件（共 15.01 GB，含 CSV、JSON、SQLite、Markdown、PDF 和视频）。

## 正文

Harness choice is a big deal.

So much room to advance and improve results across the board with agent harnesses.

Great paper highlighting this.

New research releases DataSpace, a benchmark where data agents produce verifiable tabular results from heterogeneous workspaces. 410 cross-language tasks over 7,439 artifacts totalling 15.01 GB across CSV, JSON, SQLite, Markdown, PDF, and video.

Across six recent frontier multimodal models and five widely used agent harnesses, the best accuracy reaches 66.34%. With the backbone held fixed, swapping the harness moves accuracy by 15.36 points.

Multimodal evidence integration and joins reduce accuracy across all six backbones. The benchmark is nowhere near saturated.

Paper: https://arxiv.org/abs/2608.03451

Track more trending AI papers in our academy: https://academy.dair.ai/
