Flat-Pack Bench:通过家具组装任务评估大型视觉语言模型的时空理解
阅读原文· arxiv.org现有大型视觉语言模型基准测试主要关注粗粒度任务,且依赖易于语言描述的实体。为此,研究者提出了Flat-Pack Bench,这是一个专注于家具组装任务的新基准,旨在评估模型的细粒度时空理解能力。该基准采用选择题与视觉提示的形式,考察模型在组装动作排序、状态定位、部件匹配理解与追踪等方面的表现。实验表明,最先进的模型在此类细粒度推理任务上表现欠佳,暴露出其在利用视频时序信息、进行目标追踪以及理解物理空间交互方面的不足。
The emergence of Large Vision-Language Models (LVLMs) has significantly advanced video understanding capabilities. However, existing benchmarks focus predominantly on coarse-grained tasks such as action segmentation, classification, captioning, and retrieval. Furthermore, these benchmarks often rely on entities that can be easily identified verbally, like household objects, animals, human subjects, etc., limiting their applicability to complex, in-the-wild video scenarios. But, many applications such as furniture assembly, cooking, etc., require step-by-step fine-grained spatio-temporal understanding of the video, which is not sufficiently evaluated in current benchmarks. To address this gap, we introduce Flat-Pack Bench, a novel benchmark centered on furniture assembly tasks. Our benchmark evaluates LVLMs on nuanced tasks, including temporal ordering of assembly actions, temporal localization of assembly state, understanding part mating, and tracking, using multiple-choice questions paired with visual prompts highlighting relevant parts as references for fine-grained questions. Our experiments reveal that state-of-the-art LVLMs struggle significantly with fine-grained spatio-temporal reasoning, highlighting their limitations in effectively leveraging temporal information from videos, limited tracking ability, and understanding of spatial interactions like physical contact.