# VSAS-Bench：视觉流式辅助模型的实时评估基准

- 来源：Apple Machine Learning Research（RSS）
- 发布时间：2026-05-22 08:00
- AIHOT 分数：66
- AIHOT 标记：精选
- AIHOT 链接：https://aihot.virxact.com/items/cmph748qc0lwnsljwiu60grlt
- 原文链接：https://machinelearning.apple.com/research/vsas-bench-streaming-assistant

## 精选理由

苹果搞了个实时视觉助手的评估基准，把离线评测拉到了流式场景，多模态 agent 和实时 VLM 方向的研究者值得跟进一下评估方法。

## AI 摘要

现有视觉语言模型框架主要在离线场景下评估性能，但实时视觉助手所依赖的流式模型还需考量额外指标，如反映响应时效性的“主动性”和捕捉随时间推移响应稳定性的“一致性”。为此，研究团队提出了VSAS-Bench，这是一个新的评估基准，专门针对流式视觉语言模型在实时交互任务中的表现，填补了当前评估方法在动态、持续生成场景下的空白。

## 正文

research area Computer Vision, research area Data Science and Annotationconference CVPR

content type paperpublished May 2026

VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models

AuthorsPavan Kumar Anasosalu Vasu*, Cem Koc*, Fartash Faghri*, Chun-Liang Li, Bo Feng, Zhengfeng Lai, Meng Cao, Oncel Tuzel, Hadi Pouransari*

View publication

Streaming vision-language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core mechanism for real-time visual assistants. Existing VLM frameworks predominantly assess models in offline settings. In contrast, the performance of a streaming VLM depends on additional metrics beyond pure video understanding, including proactiveness, which reflects the timeliness of the model’s responses, and consistency, which captures the robustness of its responses over time. To address this limitation, we propose VSAS-Bench, a new framework and benchmark for Visual Streaming Assistants. In contrast to prior benchmarks that primarily employ single-turn question answering on video inputs, VSAS-Bench features temporally dense annotations with over 18,000 annotations across diverse input domains and task types. We introduce standardized synchronous and asynchronous evaluation protocols, along with metrics that isolate and measure distinct capabilities of streaming VLMs. Using this framework, we conduct large-scale evaluations of recent video and streaming VLMs, analyzing the accuracy–latency trade-off under key design factors such as memory buffer length, memory access policy, and input resolution, yielding several practical insights. Finally, we show empirically that conventional VLMs can be adapted to streaming settings without additional training, and demonstrate that these adapted models outperform recent streaming VLMs. For example, Qwen3-VL-4B surpasses Dispider, the best streaming VLM on our benchmark by 3% under asynchronous protocol.

* Equal contribution

Related readings and updates.

FastVLM: Efficient Vision Encoding for Vision Language Models

July 23, 2025research area Computer Vision

Vision Language Models (VLMs) enable visual understanding alongside textual inputs. They are typically built by passing visual tokens from a pretrained vision encoder to a pretrained Large Language Model (LLM) through a projection layer. By leveraging the rich visual representations of the vision encoder and the world knowledge and reasoning capabilities of the LLM, VLMs can be useful for a wide range of applications, including accessibility…

How Far Are We from Intelligent Visual Deductive Reasoning?

May 1, 2024research area Computer Vision, research area Speech and Natural Language ProcessingHow Far Are We from AGI?

This paper was accepted at the How Far Are We from AGI? workshop at ICLR 2024.

Vision-Language Models (VLMs) such as GPT-4V have recently demonstrated incredible strides on diverse vision language tasks. We dig into vision-based deductive reasoning, a more sophisticated but less explored realm, and find previously unexposed blindspots in the current SOTA VLMs. Specifically, we leverage Raven’s Progressive Matrices (RPMs), to assess VLMs’…

Discover opportunities in Machine Learning.

Our research in machine learning breaks new ground every day.

Work with us