# UniME-R1：基于失败学习的检索中心思维链统一多模态检索

- 来源：HuggingFace Daily Papers（社区热门论文）
- 发布时间：2026-08-06 08:00
- AIHOT 分数：43
- AIHOT 链接：https://aihot.virxact.com/items/cmsicrk3r19c8ronk9gperscl
- 原文链接：https://arxiv.org/abs/2608.06060

## AI 摘要

UniME-R1 提出一种 embedder-adviser 框架，通过挖掘硬负样本模拟检索失败，让 adviser 基于初始检索结果生成检索中心思维链（RC-CoT），以纠正嵌入器对细微判别线索的混淆。若目标出现在初始 top-k 集合中则直接重排，否则生成 RC-CoT 调整检索方向并执行全库重检索。在 MMEB-V2 及多个通用多模态检索基准上，UniME-R1 持续优于强基线。

## 正文

Abstract:Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw multimodal inputs often misses fine-grained discriminative cues, leading to confusion among semantically similar candidates. Recent methods mitigate this limitation by generating Chain-of-Thought (CoT) rationales to enrich the query representation. However, such reasoning is typically derived from the query alone: it explains what the query describes, but not what the retriever misunderstands. We argue that effective retrieval reasoning should instead be conditioned on retrieval feedback. Based on this insight, we introduce UniME-R1, an embedder-adviser framework that learns to reason over initially retrieved candidates and generate Retrieval-Centric Chain-of-Thought (RC-CoT). The adviser analyzes candidates individually to identify the discriminative cues confused by the embedder. If the target appears in the initial top-k set, UniME-R1 directly reranks the candidates; otherwise, it generates RC-CoT to refine the retrieval direction and performs full-corpus re-retrieval with a dual-mode embedder. To train the framework, we mine hard negatives to simulate realistic retrieval failures, jointly optimize direct retrieval and RC-CoT-augmented retrieval, and align the adviser with retrieval outcomes through supervised learning and retrieval-oriented reinforcement learning. Extensive experiments on MMEB-V2 and a diverse set of general multimodal retrieval benchmarks demonstrate that UniME-R1 consistently improves retrieval performance over strong baselines.

Comments:

Subjects: Computer Vision and Pattern Recognition (cs.CV)

Cite as: arXiv:2608.06060 [cs.CV]

(or arXiv:2608.06060v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2608.06060

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Kaicheng Yang [

Thu, 6 Aug 2026 14:04:56 UTC (12,197 KB)

Access Paper:

Current browse context:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators
