上下文压缩器真能省钱吗?新研究揭示隐性成本

DAIR.AI · @dair_ai · X·2026-08-19 05:30·17天前
AI 导读

DAIR.AI 分享的一项研究质疑上下文压缩器的实际收益:在 24 轮工具调用智能体任务中,GPT-5.5 在压缩后完成率稳定在 80%-85%,但检索调用从 21.0 增至 63.9,耗时增至三倍。六组模型对比中检索均上升,语义无关内容使检索增加 57%,ALFWorld 未现此现象,需自行验证。论文:arxiv.org/abs/2608.16370。

DAIR.AI@dair_ai
47AI 编辑部评分,满分 100

上下文压缩器真能省钱吗?新研究揭示隐性成本

2026-08-19 05:30· 17天前
AI 导读

DAIR.AI 分享的一项研究质疑上下文压缩器的实际收益:在 24 轮工具调用智能体任务中,GPT-5.5 在压缩后完成率稳定在 80%-85%,但检索调用从 21.0 增至 63.9,耗时增至三倍。六组模型对比中检索均上升,语义无关内容使检索增加 57%,ALFWorld 未现此现象,需自行验证。论文:arxiv.org/abs/2608.16370。

Does your context compactor actually save anything?

If you built or tuned one, this one is worth your time.

(bookmark it)

Task completion is the standard way to validate compression, and it hides the real bill. In a bounded 24-turn tool-using agent, GPT-5.5 held completion flat from 80% to 85% while retrieval calls climbed from 21.0 to 63.9. The agent kept finishing. It just spent three times the tool calls going back for state the compactor had dropped.

Retrieval rose in all six model-regime comparisons, five of them significant after correction, while completion moved in none.

Content validity mattered more than volume. Replacing retained state with semantically irrelevant content raised retrieval 57% with no completion change, and random selection performed about as well as an offline hindsight oracle.

ALFWorld showed no retrieval surge under sliding compression, so this needs measuring in your own environment.

Paper: https://arxiv.org/abs/2608.16370

Track more trending AI papers in our academy: https://academy.dair.ai/