ConceptFormer:面向视觉文档检索的查询-文档对齐自适应潜在概念学习框架

HuggingFace Daily Papers(社区热门论文)·2026-08-16 08:00·9天前
AI 导读

ConceptFormer 提出一种潜在概念表示学习框架,将查询相关证据建模为连续、查询条件化的潜在概念,无需文本中间表示或原始视觉标注即可桥接局部视觉证据与语义相关性。在多个视觉文档检索基准上,其平均 NDCG@10 较最强视觉检索基线和最强 OCR 文本检索基线分别相对提升 16.7% 和 22.1%。代码与数据已公开。

HuggingFace Daily Papers(社区热门论文)
56AI 编辑部评分,满分 100

ConceptFormer:面向视觉文档检索的查询-文档对齐自适应潜在概念学习框架

2026-08-16 08:00· 9天前
AI 导读

ConceptFormer 提出一种潜在概念表示学习框架,将查询相关证据建模为连续、查询条件化的潜在概念,无需文本中间表示或原始视觉标注即可桥接局部视觉证据与语义相关性。在多个视觉文档检索基准上,其平均 NDCG@10 较最强视觉检索基线和最强 OCR 文本检索基线分别相对提升 16.7% 和 22.1%。代码与数据已公开。

Visual document retrieval is a critical component of multimodal retrieval-augmented generation, aiming to identify query-relevant pages from document collections where evidence is distributed across text, layout, charts, and visual structures. Recent efforts toward finer-grained supervision primarily rely on textual descriptions or localized visual regions as evidence proxies. However, such supervision signals may either overlook complex visual structures or provide incomplete and inaccurate representations of the underlying evidence. To address these limitations, we propose ConceptFormer, a latent concept representation learning framework for visual document retrieval. ConceptFormer models query-relevant evidence as continuous, query-conditioned latent concepts that explicitly bridge localized visual evidence and semantic relevance, without requiring either textual intermediate representations or direct reliance on raw visual annotations. During training, ConceptFormer employs a strong vision-language model to dynamically determine the number of latent concept tokens and uses these concepts as an intermediate representation to bridge the semantic gap between queries and documents, thereby guiding the learning of the embedding space. Experiments on diverse visual document retrieval benchmarks demonstrate that ConceptFormer achieves 16.7% and 22.1% relative improvements in average NDCG@10 over the strongest visual retrieval baseline and the strongest OCR-based text retrieval baseline, respectively. Further analysis reveals that latent concepts effectively connect localized visual evidence with semantic relevance, enabling the retriever to capture both fine-grained textual cues and complex document-level visual structures while preserving strong retrieval alignment. Codes and data are available at https://github.com/Neuir/ConceptFormer.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org

arXiv检索增强多模态论文/研究