HuggingFace Daily Papers(社区热门论文)
58AI 编辑部评分,满分 100

SpatialCLI:让视觉语言模型先借助空间工具推理、再内化能力

2026-07-30 08:00· 1天前
跳到正文
AI 摘要

SpatialCLI 提出三阶段框架,教视觉语言模型调用空间工具并逐步内化专家感知能力:先通过 Call 暴露专家视觉模型,再用 Cold-Start SFT 和智能体强化学习改进工具使用,最后将成功轨迹内化为自身能力。

Abstract:Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason about the overall task but often miss the visual details that determine success, while specialist vision models can capture those details but cannot translate them into task-level decisions. In this work, we propose SpatialCLI, a framework that teaches VLMs to reason with spatial tools and progressively internalize the specialist perceptual capabilities they provide. SpatialCLI proceeds in three stages: (1) Call exposes specialist vision models as spatial tools to augment the VLM's perception; (2) Learn uses Cold-Start SFT and agentic RL to improve tool use; and (3) Internalize verbalizes successful tool-use trajectories to internalize specialist perceptual capabilities. We further introduce SpatialCLI-Bench, a 516-example benchmark for compositional perception across localization, segmentation, depth, and pose. On MindCube, SpatialCLI raises Qwen3-VL-8B-Instruct from 29.3% to 84.6% with tools, surpassing GPT-5.6 Sol with tools (72.1%), while retaining 73.8% without tools after internalization.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2607.27703 [cs.AI]
  (or arXiv:2607.27703v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2607.27703
arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yang Zhou [

Thu, 30 Jul 2026 05:39:01 UTC (5,347 KB)

Access Paper:

Current browse context:

References & Citations

Bookmark

BibSonomy Reddit

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

SpatialCLI:让视觉语言模型先借助空间工具推理、再内化能力

HuggingFace Daily Papers(社区热门论文)·2026-07-30 08:00·1天前
阅读原文· arxiv.org
AI 摘要

SpatialCLI 提出三阶段框架,教视觉语言模型调用空间工具并逐步内化专家感知能力:先通过 Call 暴露专家视觉模型,再用 Cold-Start SFT 和智能体强化学习改进工具使用,最后将成功轨迹内化为自身能力。

原文 · 保持原样,未翻译
Abstract:Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason about the overall task but often miss the visual details that determine success, while specialist vision models can capture those details but cannot translate them into task-level decisions. In this work, we propose SpatialCLI, a framework that teaches VLMs to reason with spatial tools and progressively internalize the specialist perceptual capabilities they provide. SpatialCLI proceeds in three stages: (1) Call exposes specialist vision models as spatial tools to augment the VLM's perception; (2) Learn uses Cold-Start SFT and agentic RL to improve tool use; and (3) Internalize verbalizes successful tool-use trajectories to internalize specialist perceptual capabilities. We further introduce SpatialCLI-Bench, a 516-example benchmark for compositional perception across localization, segmentation, depth, and pose. On MindCube, SpatialCLI raises Qwen3-VL-8B-Instruct from 29.3% to 84.6% with tools, surpassing GPT-5.6 Sol with tools (72.1%), while retaining 73.8% without tools after internalization.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2607.27703 [cs.AI]
  (or arXiv:2607.27703v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2607.27703
arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yang Zhou [

Thu, 30 Jul 2026 05:39:01 UTC (5,347 KB)

Access Paper:

Current browse context:

References & Citations

Bookmark

BibSonomy Reddit

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

阅读原文arxiv.org