Harness choice is a big deal.
So much room to advance and improve results across the board with agent harnesses.
Great paper highlighting this.
New research releases DataSpace, a benchmark where data agents produce verifiable tabular results from heterogeneous workspaces. 410 cross-language tasks over 7,439 artifacts totalling 15.01 GB across CSV, JSON, SQLite, Markdown, PDF, and video.
Across six recent frontier multimodal models and five widely used agent harnesses, the best accuracy reaches 66.34%. With the backbone held fixed, swapping the harness moves accuracy by 15.36 points.
Multimodal evidence integration and joins reduce accuracy across all six backbones. The benchmark is nowhere near saturated.
Paper: https://arxiv.org/abs/2608.03451
Track more trending AI papers in our academy: https://academy.dair.ai/