Valentin Gabeur
项目负责人,同等贡献
Shangbang Long
项目负责人,同等贡献
Songyou Peng
项目负责人,同等贡献
Paul Voigtlaender
Shuyang Sun
Yanan Bao
Karen Truong
Zhicheng Wang
Wenlei Zhou
Jonathan T. Barron
Kyle Genova
Nithish Kannen
Sherry Ben
Yandong Li
Mandy Guo
Suhas Yogin
Yiming Gu
项目顾问
Huizhong Chen
项目顾问
Oliver Wang
领导层赞助人
Saining Xie
领导层赞助人
Howard Zhou
领导层赞助人
Kaiming He
领导层赞助人
Thomas Funkhouser
领导层赞助人
Jean-Baptiste Alayrac
领导层赞助人
Radu Soricut
领导层赞助人
摘要
近期研究表明,图像与视频生成器展现出零样本视觉理解能力,其方式令人联想到大语言模型(LLM)如何从生成式预训练中涌现出语言理解与推理能力。尽管长期以来人们推测,创造视觉内容的能力必然蕴含理解能力,但鲜有研究证明通用图像生成器能在众多不同视觉任务上达到最先进的理解水平。在本研究中,我们证明图像生成训练所起的作用类似于 LLM 预训练,能让模型学习到强大且通用的视觉表征,从而在各类视觉任务上实现最先进的性能。我们推出 Vision Banana,这是一个通过指令微调构建的通用模型,其基础是 Nano Banana Pro(NBP),训练数据混合了原始训练数据与少量视觉任务数据。通过将视觉任务的输出空间参数化为 RGB 图像,我们无缝地将感知问题重新定义为图像生成。我们的通用模型 Vision Banana 在涉及 2D 和 3D 理解的多种视觉任务上取得了最先进的结果,击败或媲美零样本领域专用模型,包括在分割任务上超越 Segment Anything Model 3,在度量深度估计上超越 Depth Anything 系列。我们证明,这些结果可以通过轻量级指令微调实现,且不牺牲基础模型的图像生成能力。这些优异结果表明,图像生成预训练是一种通用的视觉学习器。同时,它也表明图像生成可作为视觉任务的统一通用接口,类似于文本生成在语言理解与推理中所扮演的角色。我们或许正在见证计算机视觉领域的一次重大范式转变——生成式视觉预训练将在构建面向生成与理解的视觉基础模型中占据核心地位。
项目页面:vision-banana.github.io
1 引言
近年来,先进的图像与视频生成模型(Google, 2025a, b, Black Forest Labs, 2025, ByteDance, 2026, Luma, 2026, OpenAI, 2026)展现了前所未有的生成能力,能够合成高度复杂、高保真的视觉内容,并实现精准的语义控制。这种卓越的视觉创作能力表明,这些模型对视觉世界的底层结构、语义及其相互关系具有深刻的内在理解。然而,视觉表征学习的主流方法通常并不属于生成式建模范畴。相反,它们包括有监督判别学习(Krizhevsky et al., 2012, Dehghani et al., 2023, Dosovitskiy et al., 2020)、对比学习(Chen et al., 2020b, He et al., 2020, Chen et al., 2020c, Zhai et al., 2023, Tschannen et al., 2025, Radford et al., 2021)、自引导学习(Caron et al., 2021, Grill et al., 2020)、自编码(He et al., 2022, Bao et al., 2021, Chen et al., 2024)等方法及其组合(Oquab et al., 2023, Siméoni et al., 2025, Zhou et al., 2021, Cao et al., 2026)。早期在生成式视觉预训练方面的探索(Chen et al., 2020a, Bai et al., 2024)已展现出有前景的扩展特性,但其有效性仍落后于非生成式模型。
在本文中,我们探究视觉生成模型是否暗中扮演着通用视觉学习者的角色——即,为图像生成而训练的模型是否发展出了适用于视觉理解任务的内部表征。为此,我们使用少量计算机视觉数据(深度估计、表面法线估计、分割等)对预训练的图像生成器进行微调。随后,我们在多种视觉基准上评估该模型。如果微调后的模型在这些基准上达到或接近当前最优水平,同时保留其图像生成能力,那么就有强有力的证据表明,该图像生成器本质上就是视觉理解的基础模型——即一个通用视觉学习者。
这并非首篇研究生成模型隐藏理解能力,或利用图像与视频生成器作为视觉理解基础模型的工作。早期研究显示,生成模型在其特征中发展出了一些隐藏的理解能力 [Bhattad et al., 2023, Du et al., 2023, Li et al., 2023, Ranzato et al., 2011, Hjelm et al., 2018, Clark and Jaini, 2023, Baranchuk et al., 2021, Chen et al., 2016, Zhao et al., 2023, Mukhopadhyay et al., 2023, Tang et al., 2023, Zhang et al., 2023b, Li et al., 2024b, Hedlin et al., 2023, Yang and Wang, 2023]。更近期的研究观察到,最先进的图像和视频生成器能够生成视觉内容,这些内容看起来像是计算机视觉任务(如分割、深度估计和表面法线估计)输出的 RGB 可视化结果 [Zuo et al., 2025, Wiedemer et al., 2025]。然而,这些方法并未在现代基准测试上提供最先进的结果。部分原因在于这些模型并未严格遵循提示词,以所需的格式生成视觉输出,从而无法将其解码回视觉输出来计算定量指标。其他研究者 [He et al., 2024, 2025, Ke et al., 2024, Ye et al., 2024, Yu et al., 2024, Zhao et al., 2025, Wang et al., 2026b, Wu et al., 2025, Garcia et al., 2025, Xu et al., 2023] 则通过添加专用模块并进行全量微调来调整生成架构,从而在特定目标任务上达到 SOTA 水平。尽管这些方法成功利用了预训练特征的隐藏理解能力,但它们牺牲了模型在其他理解和生成任务上的通用性。
我们采用了一种受大语言模型(LLM)最新进展启发的方法。在自然语言处理(NLP)领域,生成式预训练 [Brown et al., 2020, Chowdhery et al., 2023] 用于生成基础模型(通常称为大语言模型),这些模型擅长生成文本;而指令微调 [Ouyang et al., 2022, Wei et al., 2021] 则引导它们遵循特定任务、以要求的格式生成文本并保持任务专注。类似地,我们将视觉生成模型定位为“基础”模型,并执行指令微调,以使模型能够根据提示以所需格式生成视觉输出,如图 1 所示。具体来说,模型被指示生成可解码为计算机视觉输出的 RGB 图像。此类指令提示和可解码的可视化方案旨在将视觉生成结果桥接并校准到可应用可测量基准指标的格式。例如,通过提示模型“将滑板类别分割为纯黄色(<255, 255, 0>)”,我们可以通过聚类像素值接近 <255, 255, 0> 的像素来轻松解析滑板的掩码。该策略具有三个主要优势。首先,它通过单一统一模型支持多种任务——指令微调后,权重在所有任务间共享,仅需更改提示。其次,它需要相对较少的新训练数据,因为指令微调仅教会模型如何将计算机视觉输出格式化为 RGB。第三,它有助于模型保留其原始图像生成能力,因为输出仅仅是新的 RGB 图像。
| 能力 | 基准与指标 | Vision Banana | 最佳对照 |
| 二维理解 | 指代分割:RefCOCOg UMD val(cIoU) | 73.8 | 73.4(SAM3 Agent) |
| 指代分割:ReasonSeg val(gIoU) | 79.3 | 77.0(SAM3 Agent) | |
| 语义分割:Cityscapes val(mIoU) | 69.9 | 65.2(SAM3) | |
| 实例分割:SA-Co/Gold( ) | 47.5 | 24.6(OWLv2) | |
| 三维理解 | 度量深度估计:4 个数据集的平均值( ) | 0.929 | 0.918(Depth Anything 3) |
| 表面法线估计:4 个数据集的平均值(平均角度误差) | 18.928 | 19.642(Lotus-2) | |
| 视觉生成 | 文生图:GenAI-Bench(对基准模型的胜率) | 53.5% | (Nano Banana Pro) |
| 图像编辑:ImgEdit(对基准模型的胜率) | 52.2%(Nano Banana Pro) |
我们提出了 Vision Banana,这是一个通用视觉模型,通过对 Nano Banana Pro 进行轻量级指令微调而训练得到,微调数据混合了其原始的图像生成数据和我们额外的视觉任务数据。在多个基准测试的评估中,我们发现 Vision Banana 在视觉理解和生成方面均表现出色,如表 1 所总结。在理解方面,Vision Banana 在 2D 和 3D 任务上都超越或达到了当前最优结果。例如,它在多种分割任务上击败了高度专业化的分割模型 SAM 3 [Carion et al., 2025],并在度量深度估计上超越了 3D 专家 Depth Anything 3 [Lin et al., 2025]。在生成方面,它在图像生成和编辑基准测试上的表现与其基础模型持平。在 GenAI-Bench [Li et al., 2024a] 上,Vision Banana 对基础模型取得了胜率。在图像编辑基准 ImgEdit [Ye et al., 2025] 上,Vision Banana 的胜率为 。由于这些结果是通过在其基础模型上进行轻量级指令微调构建的单一统一模型实现的,这有力地证明了 Nano Banana Pro 已经具备用于视觉理解的内部表征,只需通过指令微调即可解锁。
这项研究的意义体现在两个方面。首先,它表明图像生成器本质上其实是通用的视觉学习者,生成式视觉预训练所起的基础性作用与语言模型预训练类似。其次,它表明图像生成可以作为统一视觉理解的通用接口,这与文本生成在语言理解和推理中所扮演的角色相呼应。我们可能正在见证计算机视觉领域的一次重大范式转变,生成式视觉预训练将在构建用于生成和理解的基础视觉模型中占据核心地位。
2 方法
对 Nano Banana Pro 进行指令微调
最近的图像和视频生成器已展现出零样本能力,能够生成视觉理解任务的可视化结果 [Wiedemer et al., 2025, Zuo et al., 2025]。为了严格研究和评估这些能力,我们需要对模型进行对齐,使其生成的可视化结果能够被解码回视觉任务输出,以便进行定量评估。例如,在度量深度估计中,生成的深度热力图必须能够反算出物理深度值,才能进行定量评估。因此,我们通过对基础模型 Nano Banana Pro 进行指令微调,创建了 Vision Banana。微调时,我们以极低的比例将选定的、采用这种可逆格式的视觉任务数据混入 Nano Banana Pro 自身的训练数据集中。这个过程使我们能够将模型涌现出的生成式表征对齐到可测量的物理几何和语义标签上,从而让我们的单一通用模型能够与特定任务的专业模型进行评估和比较。
以低比例混合视觉数据作为一种轻量级的指令微调策略,确保我们的视觉任务对齐不会降低模型原有的生成先验。这一策略使我们的工作区别于以往的方法,那些方法对生成模型进行全参数微调,却没有保留图像生成数据的混合 [Gan et al., 2023, Ke et al., 2024, Zhao et al., 2025]。我们通过在两个任务上对 Vision Banana 与基础 Nano Banana Pro 进行基准测试,验证了图像生成能力的保留情况:文本到图像生成(GenAI-Bench [Li et al., 2024a])和图像编辑(ImgEdit [Ye et al., 2025])。在人工评估中,我们分别获得了 53.5% 和 47.8% 的胜率,表明 Vision Banana 成功保持了其基础模型的生成能力。我们在附录 D 中详细讨论了这些生成能力。附录中图 11(文本到图像生成)和图 12(图像编辑)的定性比较证实,Vision Banana 与 Nano Banana Pro 的输出结果高度相似。这些结果验证了 Vision Banana 并未遗忘其生成本质。
视觉任务与数据。
我们在两类基础视觉理解任务上评估了我们的框架:D 场景理解与 D 结构推理。D 套件包括指代表达、语义分割和实例分割,这些任务共同测试模型将自然语言与对应物体进行关联并分割的能力。对于 D 理解,我们聚焦于单目度量深度估计和表面法线估计,这需要几何推理能力以及关于物体尺度的内部知识。为了收集指令微调所需的数据,我们利用内部模型对网络爬取的 2D 图像进行标注,并利用渲染引擎生成的合成数据来处理 3D 任务。关键在于,我们的评估基准中的任何训练数据均未包含在指令微调的数据混合中,从而确保我们的结果反映了真正的通用能力。
3 Vision Banana——源自图像生成器的通用视觉模型
在本节中,我们展示了与任务专用专家模型的定性和定量对比评估。Vision Banana 基于图像生成器构建,在没有专用架构或自定义训练损失的情况下,在广泛的视觉理解任务上达到了 SOTA 级别的结果。
| 模型 | mIoU |
| 非零样本迁移 | |
| SegMan-L [Fu et al., 2025b] | 84.2 |
| 零样本迁移 | |
| APE-D [Shen et al., 2024] | 44.2 |
| OpenSeeD [Zhang et al., 2023a] | 47.8 |
| X-Decoder [Zou et al., 2023] | 52.0 |
| SAM 3 [Carion et al., 2025] | 65.2 |
| Vision Banana | 69.9 |
| 模型 | IL_MCC | ||
| 非零样本迁移 | |||
| SAM 3 [Carion et al., 2025] | 54.1 | 0.82 | 66.1 |
| SAM 3 [Carion et al., 2025] + Llama 3.2 (ft) | 61.2 | 0.86 | 70.8 |
| 零样本迁移 | |||
| gDino-T [Liu et al., 2024] | 3.3 | 0.15 | 16.2 |
| LLMDet-L [Fu et al., 2025a] | 6.5 | 0.21 | 27.3 |
| Gemini 2.5 [Gemini Team, 2025] | 13.0 | 0.29 | 46.1 |
| APE-D [Shen et al., 2024] | 16.4 | 0.40 | 36.9 |
| DINO-X [Ren et al., 2024] | 21.3 | 0.38 | 55.2 |
| OWLv2 [Minderer et al., 2023] | 24.6 | 0.57 | 42.0 |
| Vision Banana + Gemini 3.1 Flash-Lite | 47.5 | 0.84 | 56.0 |
| 模型 | cIoU |
| 非零样本迁移 | |
| HyperSeg-Phi2-2.7B [Wei et al., 2024] | 79.4 |
| X-SAM-Phi3-3.8B [Wang et al., 2026a] | 83.8 |
| 零样本迁移 | |
| HybridGL [Liu and Li, 2025] | 51.3 |
| LocalizationHeads-LLaVA-1.5-13B [Kang et al., 2025] | 67.7 |
| SAM 3 [Carion et al., 2025] + Gemini 2.5 Pro | 73.4 |
| Vision Banana | 73.8 |
| 模型 | gIoU |
| 非零样本迁移 | |
| X-SAM-Phi-3-3.8B [Wang et al., 2026a] | 56.6 |
| LISA-13B-LLaVA1.5 [Lai et al., 2024] | 65.0 |
| 零样本迁移 | |
| SegZero-Qwen2.5-VL-7B [Liu et al., 2025] | 62.6 |
| RSVP-GPT-4o [Lu et al., 2025] | 64.7 |
| SAM 3 [Carion et al., 2025] + Gemini 2.5 Pro | 77.0 |
| Vision Banana + Gemini 2.5 Pro | 79.3 |
3.1 二维语义理解
图像分割是视觉理解的基石,传统上需要复杂且针对特定任务的模型来将像素分类为语义类别或对象实例。当前领先的方法,例如 Segment Anything 系列 [Kirillov et al., 2023, Ravi et al., 2024, Carion et al., 2025],通过高度专业化的架构和大量昂贵的人工标注掩码数据来解决这一问题。Vision Banana 挑战了这一主流范式,证明了 SOTA 分割能力可以从图像生成预训练中自然涌现。我们并非在大量精心制作的分割样本上进行训练,而是利用了基础图像生成模型所学到的丰富表征。通过指导模型生成分割掩码的多色图像,我们获得了密集的分割图,从中可以解码出单个掩码,从而通过图像生成实现分割。如表 5、5、5 和 5 所示,这种优雅的生成式方法优于经过高度调优的专用模型,在所有评估的分割基准上实现了 SOTA 零样本迁移性能。我们与其他未在领域内数据(即这些基准的训练集)上训练过的方法进行了比较。在表格中,我们将其标记为“零样本迁移”。该术语的使用遵循 Segment Anything [Kirillov et al., 2023] 和 CLIP [Radford et al., 2021]。非零样本迁移方法以灰色标记。
语义分割。
语义分割涉及将每个像素分类到预定义的类别中,而不区分单个实例。例如,Cityscapes 基准 [Cordts 等人,2016] 定义了包括道路、行人和天空在内的类别。虽然实例分割和指代表达式分割也能传达语义信息,但我们在此严格使用“语义分割”这一术语,特指这种与实例无关、基于类别层面的含义。经典语义分割任务的这种性质可以通过文本提示词来指定,我们训练模型遵循此类指令。我们提示模型生成一张可视化图像,其中每个像素根据其类别进行着色,如图 2 所示。
关键在于,我们的方法是开放词汇的:目标类别不限于固定集合,可以在提示词中动态指定,同时附带相应的颜色映射。我们支持多种提示词风格,包括自然语言描述(例如,“马卡龙蛋糕用黄色表示”)和结构化的 JSON 映射,颜色可以指定为命名颜色、十六进制代码或 RGB 元组。为了进行定量评估,我们对生成的图像进行后处理,将每个像素分配给在 RGB 空间中目标颜色最接近的类别。
我们在表 5 中将 Vision Banana 与现有方法在 Cityscapes 验证集上进行了比较。在评估过程中,我们对每个示例使用相同的文本提示词,为 19 个类别提供完整的类别到颜色映射,包括图像中不存在的类别。如表 5 所示,Vision Banana 在 mIoU 上比 SAM 3 高出若干百分点,并在开放词汇模型中取得了最佳性能,缩小了与 SegMan [Fu 等人,2025b] 等封闭集、非零样本专用模型之间的差距。
实例分割。
与语义分割不同,实例分割要求模型区分属于同一类别的不同个体对象。例如,如果一张图像中包含五只狗,我们希望模型为每只动物生成独立的掩码。这对 Vision Banana 提出了独特挑战:由于实例数量事先未知,我们无法在提示词中预先指定特定颜色。为解决这一挑战,我们仅向模型提示目标类别和背景颜色,指示其为每个独立实例分配唯一且可区分的颜色。我们让模型动态地为该类别的不同实例分配不同颜色。定性示例如图 3 所示。通过附录 A 中详述的多阶段聚类算法,可以从生成的 RGB 图像中提取出单个实例掩码。
我们在开放词汇名词短语(NP)实例分割基准 SA-Co/Gold [Carion et al., 2025] 上评估了我们的模型。我们在表 5 中总结了主要结果,并在附录 B 的表 8 中提供了全面的类别细分。我们还在附录 C 中进行了定性评估。SA-Co/Gold 基准包含 168k 个图像-名词短语对,其中绝大多数为负查询(即目标名词短语在图像中不存在)。虽然 Vision Banana 理论上可以通过生成纯黑图像来处理这些负查询,但我们并未对模型进行调优以生成此类空掩码图像。我们转而将图像-名词短语对分类为正例或负例的任务交给多模态大语言模型。为此,我们向 Gemini 3.1 Flash-Lite 提示:“这张图像中是否存在 <NP> 的实例?请从以下选项中选择答案:(A):是,(B):否。”然后,我们仅对预测为正例的样本使用 Vision Banana 生成掩码图像。该方法与 Carion 等人 [2025] 中提出的“SAM 3 + Llama 3.2 (ft)”方法(标记为“SAM 3 + EV”)相关,他们微调了 Llama 3.2 以为 SAM 3 生成存在性评分。
在零样本迁移设置下,Vision Banana(搭配 Gemini 3.1 Flash-Lite)取得了最先进的性能,超越了包括 Gemini 2.5 [Gemini Team, 2025]、APE-D [Shen et al., 2024]、DINO-X [Ren et al., 2024] 和 OWLv2 [Minderer et al., 2023] 在内的现有模型。虽然 Vision Banana 在 SA-Co/Gold 上仍落后于 SAM 3 专家模型,但我们强调,与 SAM 3 不同,我们没有将 SA-Co 数据集纳入训练数据混合中。
指代表达分割。
与传统的固定类别分割不同,指代表达分割评估模型对由长篇幅、自由形式的自然语言查询所描述的对象进行分割的能力。此任务要求模型理解并推理细微的自然语言表达,同时捕捉对象之间的复杂关系。如表 5 和表 5 所示,我们的模型在零样本迁移设置下取得了最先进的性能,在 RefCOCOg UMD [Kazemzadeh et al., 2014] 上获得了 cIoU,在 ReasonSeg [Lai et al., 2024] 上获得了 gIoU。它持续优于 SAM 3 Agent [Carion et al., 2025](该模型将 SAM 3 与 Gemini 2.5 Pro 配对)以及其他近期零样本方法,包括 HybridGL [Liu and Li, 2025]、LocalizationHeads [Kang et al., 2025]、SegZero [Liu et al., 2025] 和 RSVP [Lu et al., 2025]。在 RefCOCOg 上,与在训练集上训练过的方法(如 HyperSeg [Wei et al., 2024] 和 X-SAM [Wang et al., 2026a])相比,仍存在性能差距。
对于 ReasonSeg 中的复杂推理查询,我们遵循标准做法,将推理步骤委托给多模态大语言模型。具体来说,我们利用 Gemini 2.5 Pro 将推理查询转换为描述性指代,然后将其作为 Vision Banana 的提示词。我们在单轮推理设置中评估了这一流程,其中 Gemini 和 Vision Banana 各被调用一次。在此设置下,Vision Banana 搭配 Gemini 2.5 Pro 的表现优于几种直接在 ReasonSeg 上训练的非零样本方法,包括 X-SAM [Wang et al., 2026a] 和 LISA [Lai et al., 2024]。图 4 中的定性结果展示了 Vision Banana 能够锚定多种语言线索,从物理动作(“伸懒腰的猫”)和非常规物体角色(“用作游戏手柄的烤面包机”)到标牌上的多语言文本。这凸显了我们方法的一个关键优势:从生成式预训练中继承的丰富多模态先验知识,使 Vision Banana 能够比专门的分割模型更有效地推理“要分割什么”。
有趣的是,Vision Banana 还展现出强大的跨任务迁移能力,它能够类似地理解指代表达,并结合标准的语义分割和实例分割任务,尽管它并未在这些任务上针对自由形式查询进行过显式训练。例如,在图 2(b)(右图)中,模型理解了“墙上的图案”所指代的内容。在图 3(b)(右图)中,模型成功地将新月形牛角包与其他形状的牛角包区分开来。这些发现表明,我们的生成式预训练在不同视觉锚定范式下产生了高度鲁棒且可迁移的表征。
3.2 基于单目图像的三维理解
Vision Banana 展现出从二维单目图像推断三维结构的强大能力。我们在两个经典任务上评估了该能力:单目度量深度估计和表面法线估计。如表 1 所示,Vision Banana 在这两项任务上均达到了 SOTA 性能,超越了 Depth Anything V3 [Lin 等人,2025] 和 Lotus-2 [He 等人,2025] 等专业模型。
度量深度估计。
深度估计的目标是从单目图像生成深度图,其中每个像素的值代表从相机平面到观测物体的物理度量距离 [Eigen 等人,2014]。这是一项基础的计算机视觉任务,惠及机器人、增强/虚拟现实和自动驾驶等广泛应用。然而,深度估计本质上是不适定问题,因为二维投影本身会丢弃关键的三维几何信息。此外,单目深度估计尤其具有挑战性,因为即使已知相机内参,它也缺乏多视角设置中可用的视差线索。
在深度学习时代,研究界基本上将深度估计视为一个密集的逐像素监督回归问题,采用专门的架构和特定领域的损失函数。大多数最新的 SOTA 方法在训练、推理或两者过程中都依赖相机内参 [Yang et al., 2024, Bochkovskii et al., 2024, Wang et al., 2025b, c, He et al., 2025, 2024, Hu et al., 2024, Cai et al., 2025, Lin et al., 2025, Piccinelli et al., 2025b, a]。虽然使用内参减轻了深度估计固有的歧义性,但也需要专门的模型设计。相比之下,我们的工作基于这样一个假设:生成式建模的模态寻求特性能够自然地解决训练目标的歧义性,从而消除了对此类专门技术的需求。此外,与针对性较强的模型相比,预训练过程中获得的广泛世界知识赋予了模型更强的物体尺寸和距离先验。为了使 Nano Banana Pro 能够以公制单位估计深度,我们指示模型输出精心构建的深度值伪彩色可视化图像。
为了将深度图可视化为 RGB 图像,我们在无界深度值与有界 RGB 值之间建立映射。由于精确的公制深度对近处图像内容的实用性通常高于远处内容(例如,可抓取的物体对机器人任务更重要,立体/单目深度基准通常以视差或相对/对数深度来衡量精度),我们在 RGB 编码之前对公制深度进行“曲线化”处理。具体来说,这首先通过应用 Barron [2025] 的幂变换来扭曲深度值,然后使用这些曲线化距离生成伪彩色可视化图像。我们将幂变换限制在 范围内,并将其重新缩放,以将公制距离映射到归一化距离 :
| (1) |
在所有实验中,我们将形状参数设为 ,尺度参数设为 。这些弯曲并归一化后的距离随后用于沿一条分段线性函数进行插值,该函数沿着 RGB 立方体的边缘,从黑色到白色遍历其棱边,类似于 3D 希尔伯特曲线的第一次迭代。图 5 展示了这一过程的可视化。
从归一化距离到 RGB 颜色的映射可以通过将 RGB 值投影到最近的线段上,然后沿立方体边缘反转线性插值来求逆。由于伪彩色可视化和幂变换都是严格可逆的,它们的复合构成了度量深度空间与 RGB 颜色空间之间的双射。在训练期间,我们将此映射应用于真实度量深度,以生成 RGB 训练目标。在推理时,我们应用逆映射将模型生成的 RGB 图像解码回度量深度,从而能够在标准深度基准上进行直接评估。为了增强模型对不同颜色表示的鲁棒性,我们使用其他颜色映射(如 Plasma、Inferno、Viridis 和灰度图)来扩充训练数据。
| DepthLM-7B [Cai 等人,2025] | Depth Any. v3 [Lin 等人,2025] | Depth Pro [Bochkovskii 等人,2024] | UniK3D [Piccinelli 等人,2025a] | MoGe-2 [Wang 等人,2025c] | Vision Banana | ||
| 相机内参 | 推理 | ✓ | ✓ | ||||
| 训练 | ✓ | ✓ | ✓ | ✓ | ✓ | ||
| 平均 | 部分* | 部分* | 0.715 | 0.823 | 0.802 | 0.882 | |
| 基准 | AbsRel | 0.156 | 0.144 | 0.116 | |||
| NYU | 0.915 | 0.963 | 0.961 | 0.965 | 0.961 | 0.948 | |
| [Silberman 等人,2012] | AbsRel | 0.07 | 0.074 | 0.0733 | 0.081 | ||
| iBims1 | 0.92 | 0.913 | 0.919 | 0.830 | 0.934 | ||
| [Koch 等人,2018] | AbsRel | 0.104 | 0.136 | 0.078 | |||
| ETH3D | 0.718 | 0.917 | 0.415 | 0.687 | 0.908 | 0.935 | |
| [Schops 等人,2019] | AbsRel | 0.104 | 0.327 | 0.236 | 0.104 | 0.103 | |
| DIODE-室内 | 0.838 | 0.671 | 0.713 | 0.664 | 0.917 | ||
| [Vasiljevic 等人,2019] | AbsRel | 0.123 | 0.199 | 0.161 | 0.175 | 0.108 | |
| KITTI | 0.953 | 0.843 | 0.812 | 0.629 | 0.915 | ||
| [Uhrig 等人,2017] | AbsRel | 0.086 | 0.121 | 0.174 | 0.181 | 0.107 | |
| nuScenes | 0.865 | 0.491 | 0.840 | 0.820 | 0.643 | ||
| [Caesar 等人,2020] | AbsRel | 0.287 | 0.189 | 0.195 | 0.219 | ||
* DepthLM-7B 在其评估的 4 个数据集(NYU + iBims1 + ETH3D + nuScenes)上的平均值为 ;我们在相同 4 个数据集上的平均值为 。Depth-Anything V3 在其评估的 4 个数据集(NYU + ETH3D + DIODE + KITTI)上的平均值为 ;我们在相同 4 个数据集上的平均值为 。DepthLM 在 nuScenes 上训练过,因此并非零样本。数据由 Depth-Anything V3 [Lin 等人,2025] 报告。
表 6 展示了 Vision Banana 与专业模型在六个主要学术基准上的实证结果对比。Vision Banana 的平均准确率达到 0.882,比 Unik3D [Piccinelli 等人,2025a] 高出近 6 个百分点,同时与 MoGe-2 [Wang 等人,2025c] 相比,绝对相对误差(AbsRel)降低了 20%。值得注意的是,Vision Banana 在 Depth Anything V3 [Lin 等人,2025] 所评估的四个数据集(NYU、ETH3D、DIODE、KITTI)上的平均表现优于后者( 对比 ),在近场和远距离场景中都展现了稳健的性能。我们的模型完全使用仿真引擎生成的合成深度数据进行训练——我们未使用任何真实世界的深度数据,并且排除了来自我们评估的任何深度数据集的训练数据。请注意,这一结果是在训练和推理过程中均不依赖相机参数(无论是内参还是外参)的情况下实现的。通过利用其基础模型中嵌入的庞大几何先验知识,Vision Banana 仅凭视觉线索和物体关系就能推断出绝对尺度,从而实现对任意输入图像的零样本泛化。
![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() |
| 输入图像 | 生成的深度图像 | 可视化视角 1 | 可视化视角 2 |
定性检查进一步验证了模型的能力。如图 6 所示,Vision Banana 生成了高度精确的深度图,即使在教室等杂乱环境中也能保持清晰的几何细节。当这些二维预测被反投影到三维点云时,它们在多样场景中展现出全局一致性,保持了准确的平面表面和正确的几何结构。除了常见的学术基准测试,我们还使用一张随手拍摄的智能手机照片进行了“氛围测试”,如图 7 所示。经 Google Maps 测量的深度交叉验证,Vision Banana 成功地对这张由训练期间未见过的消费级设备拍摄的照片生成了准确的深度估计。
表面法向量估计。
表面法线估计是另一项关键的视觉任务。表面法线是取值范围在 到 之间的单位向量,可作为局部几何形状和场景结构的重要代理。与度量深度所需的复杂颜色映射不同,表面法线的可视化与 RGB 色彩空间天然对齐,从而能够直接集成到我们的模型中。
我们特别采用了基于标准右手坐标系(+x 向右,+y 向上,+z 指向图像平面外)的相机空间法线公式。在这种表示中,方向向量分量直接映射到 RGB 通道,即 , , :
-
朝左:编码为粉红色。
-
朝上:编码为浅绿色。
-
朝向相机:编码为浅蓝色。
表 7 将 Vision Banana 与 SOTA 专业方法在四个公开基准上进行了比较。在三个室内数据集的平均结果中,Vision Banana 实现了最低的平均角度误差和中位角度误差。它在室外场景上也展示了具有竞争力的精度。
| 方法 | 室内 | 室内 | 室外 | |||||||
| 平均 | NYUv2 [Silberman et al., 2012] | DIODE-indoor [Vasiljevic et al., 2019] | ScanNet [Dai et al., 2017] | VKitti [Cabon et al., 2020] | ||||||
| 均值 | 中位数 | 均值 | 中位数 | 均值 | 中位数 | 均值 | 中位数 | 均值 | 中位数 | |
| Marigold [Ke et al., 2024] | 19.606 | 11.828 | 20.864 | 11.134 | 16.671 | 12.084 | 21.284 | 12.268 | – | – |
| DSINE [Bae and Davison, 2024] | 17.017 | 10.190 | 16.4 | 8.4 | 18.453 | 13.871 | 16.2 | 8.3 | 28.9 | 9.9 |
| StableNormal [Ye et al., 2024] | 17.168 | 10.028 | 19.707 | 10.527 | 13.701 | 9.46 | 18.098 | 10.097 | – | – |
| Lotus-2-Normal [He et al., 2025] | 16.558 | – | 16.9 | N/A | 18.575 | N/A | 14.2 | N/A | 28.894 | 9.677 |
| Vision Banana | 15.549 | 9.300 | 17.778 | 8.876 | 13.818 | 11.556 | 15.052 | 7.468 | 29.063 | 10.699 |
![]() | ![]() | ![]() |
![]() | ![]() | ![]() |
| 输入图像 | Lotus-2-Normal | Vision Banana |
图 8 将 Vision Banana 的输出与领先的外部方法 Lotus-2 [He et al., 2025] 进行了视觉对比。Vision Banana 始终能生成保真度显著更高、细节粒度更精细的表面法线图。图 8 底行突出展示了来自 Virtual KITTI 2 [Cabon et al., 2020] 的一个样本。尽管在该基准测试上,Vision Banana 的定量误差略高于 Lotus-2,但其视觉质量明显更优。此外需注意,Lotus-2 在 Virtual KITTI 2 上针对表面法线估计进行了训练,而 Vision Banana 则严格遵守零样本迁移协议,从未见过任何评估基准的训练集。
4 讨论
图像生成器是通用视觉学习者。
生成式预训练 [Radford et al., 2018, 2019, Brown et al., 2020] 已从根本上改变了语言理解与推理。与此同时,近期对涌现视觉能力的观察 [Wiedemer et al., 2025, Zuo et al., 2025] 引发了人们的猜测:计算机视觉正接近类似的范式转变。通过将领先的图像生成器 Nano Banana Pro 指令微调为最先进的视觉生成与理解模型,我们证实这一转变已然发生。在大规模图像生成上预训练的模型,自然习得了强大的视觉理解能力。这些生成式先验超越了传统专业视觉模型所采用的专用架构和专门训练范式。我们正目睹一场由生成式视觉预训练驱动的计算机视觉范式转变,我们相信这为真正的视觉基础模型以及基于视觉的通用人工智能铺平了道路。
图像生成作为通用接口。
作为本研究的副产品,我们证明图像生成可以作为计算机视觉的通用接口,类似于文本生成作为自然语言中许多任务(包括语言理解、生成、推理、数学、编程、智能体任务等)的统一接口。通过将视觉任务输出表示为 RGB 图像,我们可以使用自然语言提示词无缝地指导模型。虽然我们并非首个将视觉输出编码为 RGB 图像的研究者 [Ke et al., 2024, Zhao et al., 2025, Gan et al., 2023, Wang et al., 2023, Lu et al., 2022, 2024, Xie et al., 2024, Inclusion AI, 2025],但我们证明,当与强大的预训练视觉生成器结合时,这种简单的设计足以超越现代特定领域的专家模型。
除了将视觉任务输出统一为 RGB 图像之外,生成式建模天然地为视觉任务中的歧义性提供了一种解决方案——即单个输入可能对应输出分布的多个模式。为了防止输出坍缩为模糊的平均值,专家判别模型 [Carion et al., 2025, Lin et al., 2025] 通常诉诸于定制化的架构和训练损失。例如,Segment Anything 模型 [Kirillov et al., 2023, Ravi et al., 2024, Carion et al., 2025] 会返回多个分割掩码,但仅对其中一个应用损失函数。然而,生成式模型天生就能学习完整的数据分布,通过设计优雅地处理歧义性。通过消除对定制化架构设计的需求,这种公式化方法有望催生真正统一的“全能”多模态模型。
未来工作。
尽管 Vision Banana 在单目图像的二维语义理解和三维理解等基础任务上达到了 SOTA 水平,但仍有几个令人兴奋的方向值得未来探索。首先,扩大指令微调任务的多样性,可能会解锁更多涌现性的跨任务泛化能力,类似于在大语言模型中观察到的行为 [Wei et al., 2021]。其次,我们目前的评估聚焦于单目图像输入。未来,我们可以将该框架扩展以处理多视角输入 [Wang et al., 2025a] 和视频输入 [Zhang et al., 2025]。同样,探究视频生成器是否能产生更丰富、具有时间感知能力的视觉表征,也是一个极具前景的研究方向。另一个重要的下一步是探索基础视觉模型与大语言模型的协同整合,以增强跨模态推理能力。最后,使用像 Nano Banana Pro 这样的图像生成器,目前其计算开销远高于运行轻量级专用模型。开发加速和降本策略,将是部署生成式视觉框架必须克服的关键障碍。
致谢
我们感谢 Xi Chen、Fei Xia、Kaushik Shivakumar、Abhishek Sinha、Phillip Lippe、Yilin Gao、Javier Rey、Sanghyun Woo、Renshen Wang、Wentao Yuan、Keran Rong、Rundi Wu、Manoj Kumar、Manli Shu、Francesco Piccinno、Ishita Dasgupta、Benigno Uria、Miki Rubinstein、Aäron van den Oord、Jon Shlens 在讨论、建议和技术指导方面提供的帮助。
参考文献
- Bae and Davison [2024] G. Bae 和 A. J. Davison。重新思考表面法线估计的归纳偏置。载于《IEEE/CVF 计算机视觉与模式识别会议论文集》,第 9535–9545 页,2024 年。
- Bai et al. [2024] Y. Bai、X. Geng、K. Mangalam、A. Bar、A. L. Yuille、T. Darrell、J. Malik 和 A. A. Efros。序列建模实现大规模视觉模型的可扩展学习。载于《IEEE/CVF 计算机视觉与模式识别会议论文集》,第 22861–22872 页,2024 年。
- Bao 等人 [2021] H. Bao, L. Dong, S. Piao, 和 F. Wei. Beit: Bert pre-training of image transformers. arXiv 预印本 arXiv:2106.08254, 2021.
- Baranchuk 等人 [2021] D. Baranchuk, I. Rubachev, A. Voynov, V. Khrulkov, 和 A. Babenko. 基于扩散模型的标签高效语义分割. arXiv 预印本 arXiv:2112.03126, 2021.
- Barron [2025] J. T. Barron. 一种幂变换. arXiv 预印本 arXiv:2502.10647, 2025.
- Bhattad 等人 [2023] A. Bhattad, D. McKee, D. Hoiem, 和 D. Forsyth. StyleGAN 能够理解法线、深度、反照率等更多信息. 神经信息处理系统进展, 36:73082–73103, 2023.
- Black Forest Labs [2025] Black Forest Labs. FLUX.2: 前沿视觉智能. https://bfl.ai/blog/flux-2, 2025.
- Bochkovskii 等人 [2024] A. Bochkovskii, A. Delaunoy, H. Germain, M. Santos, Y. Zhou, S. R. Richter, 和 V. Koltun. Depth Pro: 不到一秒内实现锐利的单目度量深度. arXiv 预印本 arXiv:2410.02073, 2024.
- Brown 等人 [2020] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, 等. 语言模型是少样本学习者. 神经信息处理系统进展, 33:1877–1901, 2020.
- 字节跳动 [2026] 字节跳动. Seedance 2.0. https://seed.bytedance.com/en/seedance2_0/, 2026. 访问日期: 2026-03-18.
- Cabon 等人 [2020] Y. Cabon, N. Murray, 和 M. Humenberger. Virtual KITTI 2. arXiv 预印本 arXiv:2001.10773, 2020.
- Caesar 等人 [2020] H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, 和 O. Beijbom. nuScenes: 一个用于自动驾驶的多模态数据集. 见 IEEE/CVF 计算机视觉与模式识别会议论文集, 页码 11621–11631, 2020.
- Cai 等人 [2025] Z. Cai, C.-F. Yeh, H. Xu, Z. Liu, G. Meyer, X. Lei, C. Zhao, S.-W. Li, V. Chandra, 和 Y. Shi. DepthLM: 基于视觉语言模型的度量深度. arXiv 预印本 arXiv:2509.25413, 2025.
- Cao 等人 [2026] B. Cao, K. Chen, K.-K. Maninis, K. Chen, A. Karpur, Y. Xia, S. Dua, T. Dabral, G. Han, B. Han, 等. TIPsV2: 通过增强的补丁-文本对齐推进视觉-语言预训练. arXiv 预印本 arXiv:2604.12012, 2026.
- Carion 等人 [2025] N. Carion, L. Gustafson, Y.-T. Hu, S. Debnath, R. Hu, D. Suris, C. Ryali, K. V. Alwala, H. Khedr, A. Huang, 等. Sam 3: 基于概念分割一切. arXiv 预印本 arXiv:2511.16719, 2025.
- Caron 等人 [2021] M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, 和 A. Joulin. 自监督视觉 Transformer 中的涌现特性. 见 IEEE/CVF 国际计算机视觉大会论文集, 第 9650–9660 页, 2021.
- Chen 等人 [2020a] M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, 和 I. Sutskever. 基于像素的生成式预训练. 见国际机器学习大会, 第 1691–1703 页. PMLR, 2020a.
- Chen 等人 [2020b] T. Chen, S. Kornblith, M. Norouzi, 和 G. Hinton. 视觉表征对比学习的简单框架. 见国际机器学习大会, 第 1597–1607 页. PMLR, 2020b.
- Chen 等人 [2016] X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, 和 P. Abbeel. InfoGAN: 通过信息最大化生成对抗网络实现可解释表征学习. 神经信息处理系统进展, 29, 2016.
- Chen 等人 [2020c] X. Chen, H. Fan, R. Girshick, 和 K. He. 基于动量对比学习的改进基线. arXiv 预印本 arXiv:2003.04297, 2020c.
- Chen 等人 [2024] X. Chen, Z. Liu, S. Xie, 和 K. He. 解构用于自监督学习的去噪扩散模型. arXiv 预印本 arXiv:2401.14404, 2024.
- Chowdhery 等人 [2023] A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, 等. PaLM: 基于 Pathways 扩展语言建模. 机器学习研究期刊, 24(240):1–113, 2023.
- Clark 和 Jaini [2023] K. Clark 和 P. Jaini. 文本到图像扩散模型是零样本分类器. 神经信息处理系统进展, 36:58921–58937, 2023.
- Cordts 等人 [2016] M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, 和 B. Schiele. 用于语义城市场景理解的 Cityscapes 数据集. 见 IEEE 计算机视觉与模式识别大会论文集, 第 3213–3223 页, 2016.
- Dai 等人 [2017] A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser 和 M. Nießner。ScanNet:室内场景的丰富标注三维重建。载于《IEEE 计算机视觉与模式识别会议论文集》,第 5828–5839 页,2017 年。
- Dehghani 等人 [2023] M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. P. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin 等。将视觉 Transformer 架构扩展到 220 亿参数。载于《国际机器学习会议》,第 7480–7512 页。PMLR,2023 年。
- Dosovitskiy 等人 [2020] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly 等。一张图像抵得上 16x16 个词:用于大规模图像识别的 Transformer 架构。arXiv 预印本 arXiv:2010.11929,2020 年。
- Du 等人 [2023] X. Du, N. Kolkin, G. Shakhnarovich 和 A. Bhattad。生成模型:它们知道什么?它们知道事情吗?让我们一探究竟!arXiv 预印本 arXiv:2311.17137,2023 年。
- Eigen 等人 [2014] D. Eigen, C. Puhrsch 和 R. Fergus。使用多尺度深度网络从单张图像进行深度图预测。载于《神经信息处理系统进展》(NeurIPS),第 27 卷,2014 年。
- Fu 等人 [2025a] S. Fu, Q. Yang, Q. Mo, J. Yan, X. Wei, J. Meng, X. Xie 和 W.-S. Zheng。LLMDet:在大语言模型监督下学习强大的开放词汇目标检测器。载于《计算机视觉与模式识别会议论文集》,第 14987–14997 页,2025a。
- Fu 等人 [2025b] Y. Fu, M. Lou 和 Y. Yu。SegMan:利用状态空间模型和局部注意力进行全尺度上下文建模以实现语义分割。载于《计算机视觉与模式识别会议论文集》,第 19077–19087 页,2025b。
- Gan 等人 [2023] Y. Gan, S. Park, A. Schubert, A. Philippakis 和 A. M. Alaa。InstructCV:指令微调的文生图扩散模型作为视觉通才。arXiv 预印本 arXiv:2310.00390,2023 年。
- Garcia 等人 [2025] G. M. Garcia、K. A. Zeid、C. Schmidt、D. De Geus、A. Hermans 和 B. Leibe。微调图像条件扩散模型比你想象的要简单。载于《冬季计算机视觉应用会议论文集》,第 753–762 页,2025 年。
- Gemini 团队 [2025] Gemini 团队。Gemini 2.5:以高级推理、多模态、长上下文和下一代智能体能力推动前沿发展。arXiv 预印本,2025 年。
- Google [2025a] Google。推出 nano banana pro。https://blog.google/innovation-and-ai/products/nano-banana-pro/,2025a。访问日期:2026-03-15。
- Google [2025b] Google。Veo 3 公告。https://blog.google/innovation-and-ai/products/generative-media-models-io-2025/,2025b。访问日期:2026-03-15。
- Grill 等人 [2020] J.-B. Grill、F. Strub、F. Altché、C. Tallec、P. Richemond、E. Buchatskaya、C. Doersch、B. Avila Pires、Z. Guo、M. Gheshlaghi Azar 等。自举你的潜在表示——一种自监督学习的新方法。《神经信息处理系统进展》,33:21271–21284,2020 年。
- He 等人 [2024] J. He、H. Li、W. Yin、Y. Liang、L. Li、K. Zhou、H. Zhang、B. Liu 和 Y.-C. Chen。Lotus:用于高质量密集预测的基于扩散的视觉基础模型。arXiv 预印本 arXiv:2409.18124,2024 年。
- He 等人 [2025] J. He、H. Li、M. Sheng 和 Y.-C. Chen。Lotus-2:利用强大的图像生成模型推进几何密集预测。arXiv 预印本 arXiv:2512.01030,2025 年。
- He 等人 [2020] K. He、H. Fan、Y. Wu、S. Xie 和 R. Girshick。用于无监督视觉表示学习的动量对比。载于《IEEE/CVF 计算机视觉与模式识别会议论文集》,第 9729–9738 页,2020 年。
- He 等人 [2022] K. He、X. Chen、S. Xie、Y. Li、P. Dollár 和 R. Girshick。掩码自编码器是可扩展的视觉学习器。载于《IEEE/CVF 计算机视觉与模式识别会议论文集》,第 16000–16009 页,2022 年。
- Hedlin 等人 [2023] E. Hedlin、G. Sharma、S. Mahajan、H. Isack、A. Kar、A. Tagliasacchi 和 K. M. Yi。使用稳定扩散的无监督语义对应。《神经信息处理系统进展》,36:8266–8279,2023 年。
- Hjelm 等人 [2018] R. D. Hjelm、A. Fedorov、S. Lavoie-Marchildon、K. Grewal、P. Bachman、A. Trischler 和 Y. Bengio。通过互信息估计与最大化学习深度表示。arXiv 预印本 arXiv:1808.06670,2018 年。
- Hu 等人 [2024] M. Hu、W. Yin、C. Zhang、Z. Cai、X. Long、H. Chen、K. Wang、G. Yu、C. Shen 和 S. Shen。Metric3d v2:一种用于零样本度量深度和表面法线估计的多功能单目几何基础模型。IEEE 模式分析与机器智能汇刊,2024 年。
- Inclusion AI [2025] Inclusion AI。Ming-flash-omni:一种用于多模态感知与生成的稀疏统一架构。arXiv 预印本 arXiv:2510.24821,2025 年。
- Kang 等人 [2025] S. Kang、J. Kim、J. Kim 和 S. J. Hwang。你的大型视觉语言模型仅需少量注意力头即可实现视觉定位。载于计算机视觉与模式识别会议论文集,第 9339–9350 页,2025 年。
- Kazemzadeh 等人 [2014] S. Kazemzadeh、V. Ordonez、M. Matten 和 T. Berg。Referitgame:在自然场景照片中指代物体。载于 2014 年自然语言处理经验方法会议(EMNLP)论文集,第 787–798 页,2014 年。
- Ke 等人 [2024] B. Ke、A. Obukhov、S. Huang、N. Metzger、R. C. Daudt 和 K. Schindler。将基于扩散的图像生成器重新用于单目深度估计。载于 IEEE/CVF 计算机视觉与模式识别会议论文集,第 9492–9502 页,2024 年。
- Kirillov 等人 [2023] A. Kirillov、E. Mintun、N. Ravi、H. Mao、C. Rolland、L. Gustafson、T. Xiao、S. Whitehead、A. C. Berg、W.-Y. Lo、P. Dollár 和 R. Girshick。分割一切。载于 IEEE/CVF 国际计算机视觉会议(ICCV),第 4015–4026 页,2023 年。
- Koch 等人 [2018] T. Koch、L. Liebel、F. Fraundorfer 和 M. Korner。基于 CNN 的单幅图像深度估计方法评估。载于欧洲计算机视觉会议(ECCV)研讨会论文集,第 0–0 页,2018 年。
- Krizhevsky 等人 [2012] A. Krizhevsky、I. Sutskever 和 G. E. Hinton。使用深度卷积神经网络进行 ImageNet 分类。神经信息处理系统进展,第 25 卷,2012 年。
- Lai 等人 [2024] X. Lai, Z. Tian, Y. Chen, Y. Li, Y. Yuan, S. Liu, 和 J. Jia. Lisa:通过大语言模型进行推理分割. 收录于《IEEE/CVF 计算机视觉与模式识别会议论文集》,第 9579–9589 页,2024 年.
- Li 等人 [2023] A. C. Li, M. Prabhudesai, S. Duggal, E. Brown, 和 D. Pathak. 你的扩散模型其实是一个零样本分类器. 收录于《IEEE/CVF 国际计算机视觉会议论文集》,第 2206–2217 页,2023 年.
- Li 等人 [2024a] B. Li, Z. Lin, D. Pathak, J. Li, Y. Fei, K. Wu, T. Ling, X. Xia, P. Zhang, G. Neubig, 等. Genai-bench:评估和改进组合式文本到视觉生成. arXiv 预印本 arXiv:2406.13743,2024a.
- Li 等人 [2024b] X. Li, J. Lu, K. Han, 和 V. A. Prisacariu. Sd4match:学习为语义匹配提示 Stable Diffusion 模型. 收录于《IEEE/CVF 计算机视觉与模式识别会议论文集》,第 27558–27568 页,2024b.
- Lin 等人 [2025] H. Lin, S. Chen, J. Liew, D. Y. Chen, Z. Li, G. Shi, J. Feng, 和 B. Kang. Depth Anything 3:从任意视角恢复视觉空间. arXiv 预印本 arXiv:2511.10647,2025.
- Liu 等人 [2024] S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su, J. Zhu, 和 L. Zhang. Grounding DINO:将 DINO 与接地预训练相结合用于开放集目标检测. 收录于 ECCV (47),《计算机科学讲义》,第 38–55 页. Springer,2024 年.
- Liu 和 Li [2025] T. Liu 和 S. Li. 结合增强空间引导的混合全局-局部表示用于零样本指代图像分割. 收录于《IEEE/CVF 计算机视觉与模式识别会议论文集》,第 29634–29643 页,2025 年.
- Liu 等人 [2025] Y. Liu, B. Peng, Z. Zhong, Z. Yue, F. Lu, B. Yu, 和 J. Jia. Seg-Zero:通过认知强化进行推理链引导的分割. arXiv 预印本 arXiv:2503.06520,2025.
- Lu 等人 [2022] J. Lu, C. Clark, R. Zellers, R. Mottaghi, 和 A. Kembhavi. Unified-IO:一个用于视觉、语言和多模态任务的统一模型. arXiv 预印本 arXiv:2206.08916,2022.
- Lu 等人 [2024] J. Lu, C. Clark, S. Lee, Z. Zhang, S. Khosla, R. Marten, D. Hoiem, 和 A. Kembhavi. Unified-io 2: 扩展视觉、语言、音频和动作的自回归多模态模型. 收录于 IEEE/CVF 计算机视觉与模式识别会议论文集, 第 26439–26455 页, 2024.
- Lu 等人 [2025] Y. Lu, J. Cao, Y. Wu, B. Li, L. Tang, Y. Ji, C. Wu, J. Wu, 和 W. Zhu. RSVP: 通过视觉提示与多模态思维链进行推理分割. 收录于第 63 届计算语言学协会年会论文集 (第一卷: 长文), 第 14699–14716 页, 2025.
- Luma [2026] Luma. UNI-1. https://lumalabs.ai/uni-1/, 2026. 访问日期: 2026-03-19.
- Minderer 等人 [2023] M. Minderer, A. Gritsenko, 和 N. Houlsby. 扩展开放词汇目标检测. 神经信息处理系统进展, 36:72983–73007, 2023.
- Mukhopadhyay 等人 [2023] S. Mukhopadhyay, M. Gwilliam, V. Agarwal, N. Padmanabhan, A. Swaminathan, S. Hegde, T. Zhou, 和 A. Shrivastava. 扩散模型在图像分类上超越 GAN. arXiv 预印本 arXiv:2307.08702, 2023.
- OpenAI [2026] OpenAI. GPT-Image-1.5. https://openai.com/index/new-chatgpt-images-is-here/, 2026. 访问日期: 2026-03-19.
- Oquab 等人 [2023] M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, 等人. DINOv2: 无需监督学习鲁棒视觉特征. arXiv 预印本 arXiv:2304.07193, 2023.
- Ouyang 等人 [2022] L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, 等人. 训练语言模型遵循带人类反馈的指令. 神经信息处理系统进展, 35:27730–27744, 2022.
- Piccinelli 等人 [2025a] L. Piccinelli, C. Sakaridis, M. Segu, Y.-H. Yang, S. Li, W. Abbeloos, 和 L. Van Gool. UniK3D: 通用相机单目 3D 估计. 收录于计算机视觉与模式识别会议论文集, 第 1028–1039 页, 2025a.
- Piccinelli 等人 [2025b] L. Piccinelli, C. Sakaridis, Y.-H. Yang, M. Segu, S. Li, W. Abbeloos, 和 L. Van Gool。Unidepthv2:更简化的通用单目度量深度估计。IEEE 模式分析与机器智能汇刊,2025b。
- Radford 等人 [2018] A. Radford, K. Narasimhan, T. Salimans, I. Sutskever 等。通过生成式预训练提升语言理解。2018。
- Radford 等人 [2019] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever 等。语言模型是无监督的多任务学习器。OpenAI 博客,1(8):9,2019。
- Radford 等人 [2021] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark 等。从自然语言监督中学习可迁移的视觉模型。国际机器学习大会,第 8748–8763 页。PmLR,2021。
- Ranzato 等人 [2011] M. Ranzato, J. Susskind, V. Mnih 和 G. Hinton。关于具有识别应用的深度生成模型。CVPR 2011,第 2857–2864 页。IEEE,2011。
- Ravi 等人 [2024] N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. Carion, C.-Y. Wu, R. Girshick, P. Dollár 和 C. Feichtenhofer。SAM 2:在图像和视频中分割一切。arXiv 预印本 arXiv:2408.00714,2024。
- Ren 等人 [2024] T. Ren, Y. Chen, Q. Jiang, Z. Zeng, Y. Xiong, W. Liu, Z. Ma, J. Shen, Y. Gao, X. Jiang 等。DINO-X:面向开放世界目标检测与理解的统一视觉模型。arXiv 预印本 arXiv:2411.14347,2024。
- Schops 等人 [2019] T. Schops, T. Sattler 和 M. Pollefeys。BAD SLAM:捆绑调整直接 RGB-D SLAM。IEEE/CVF 计算机视觉与模式识别会议论文集,第 134–144 页,2019。
- Shen 等人 [2024] Y. Shen, C. Fu, P. Chen, M. Zhang, K. Li, X. Sun, Y. Wu, S. Lin 和 R. Ji。一次性对齐与提示所有内容以实现通用视觉感知。IEEE/CVF 计算机视觉与模式识别会议论文集,第 13193–13203 页,2024。
- Silberman 等人 [2012] N. Silberman, D. Hoiem, P. Kohli, 和 R. Fergus。基于 RGBD 图像的室内分割与支撑推理。收录于《欧洲计算机视觉会议》,第 746–760 页。Springer, 2012。
- Siméoni 等人 [2025] O. Siméoni, H. V. Vo, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V. Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoa 等。Dinov3。arXiv 预印本,编号 arXiv:2508.10104,2025。
- Tang 等人 [2023] L. Tang, M. Jia, Q. Wang, C. P. Phoo, 和 B. Hariharan。图像扩散中涌现的对应关系。《神经信息处理系统进展》,第 36 卷,第 1363–1389 页,2023。
- Tschannen 等人 [2025] M. Tschannen, A. Gritsenko, X. Wang, M. F. Naeem, I. Alabdulmohsin, N. Parthasarathy, T. Evans, L. Beyer, Y. Xia, B. Mustafa 等。Siglip 2:具有改进语义理解、定位和密集特征的多语言视觉语言编码器。arXiv 预印本,编号 arXiv:2502.14786,2025。
- Uhrig 等人 [2017] J. Uhrig, N. Schneider, L. Schneider, U. Franke, T. Brox, 和 A. Geiger。稀疏不变卷积神经网络。收录于《国际三维视觉会议 (3DV)》,2017。
- Vasiljevic 等人 [2019] I. Vasiljevic, N. Kolkin, S. Zhang, R. Luo, H. Wang, F. Z. Dai, A. F. Daniele, M. Mostajabi, S. Basart, M. R. Walter, 和 G. Shakhnarovich。DIODE:密集室内外深度数据集。CoRR, abs/1908.00463, 2019。URL http://arxiv.org/abs/1908.00463。
- Wang 等人 [2026a] H. Wang, L. Qiao, Z. Jie, Z. Huang, C. Feng, Q. Zheng, L. Ma, X. Lan, 和 X. Liang。X-sam:从分割一切到任意分割。收录于《AAAI 人工智能会议论文集》,第 40 卷,第 26187–26196 页,2026a。
- Wang 等人 [2025a] J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, 和 D. Novotny。Vggt:视觉几何基础 Transformer。收录于《计算机视觉与模式识别会议论文集》,第 5294–5306 页,2025a。
- Wang 等人 [2026b] L. Wang, A. Zanfir, E. G. Bazavan, M. Andriluka, 和 C. Sminchisescu。Thfm:面向 4D 人体感知及更广泛领域的统一视频基础模型。arXiv 预印本,编号 arXiv:2603.25892,2026b。
- Wang 等人 [2025b] R. Wang, S. Xu, C. Dai, J. Xiang, Y. Deng, X. Tong, 和 J. Yang. Moge:通过最优训练监督实现开放域图像精确单目几何估计. 载于 IEEE/CVF 计算机视觉与模式识别会议 (CVPR), 2025b.
- Wang 等人 [2025c] R. Wang, S. Xu, Y. Dong, Y. Deng, J. Xiang, Z. Lv, G. Sun, X. Tong, 和 J. Yang. Moge-2:具有度量尺度与清晰细节的精确单目几何. arXiv 预印本 arXiv:2507.02546, 2025c.
- Wang 等人 [2023] X. Wang, W. Wang, Y. Cao, C. Shen, 和 T. Huang. 图像以图像言说:面向上下文视觉学习的通用画师. 载于 IEEE/CVF 计算机视觉与模式识别会议论文集, 第 6830–6839 页, 2023.
- Wei 等人 [2024] C. Wei, Y. Zhong, H. Tan, Y. Liu, Z. Zhao, J. Hu, 和 Y. Yang. Hyperseg:利用大语言模型实现通用视觉分割. arXiv 预印本 arXiv:2411.17606, 2024.
- Wei 等人 [2021] J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, 和 Q. V. Le. 微调后的语言模型是零样本学习器. arXiv 预印本 arXiv:2109.01652, 2021.
- Wiedemer 等人 [2025] T. Wiedemer, Y. Li, P. Vicol, S. S. Gu, N. Matarese, K. Swersky, B. Kim, P. Jaini, 和 R. Geirhos. 视频模型是零样本学习器与推理器. arXiv 预印本 arXiv:2509.20328, 2025.
- Wu 等人 [2025] C. Wu, J. Li, J. Zhou, J. Lin, K. Gao, K. Yan, S.-m. Yin, S. Bai, X. Xu, Y. Chen, 等. Qwen-Image 技术报告. arXiv 预印本 arXiv:2508.02324, 2025.
- Xie 等人 [2024] J. Xie, W. Mao, Z. Bai, D. J. Zhang, W. Wang, K. Q. Lin, Y. Gu, Z. Chen, Z. Yang, 和 M. Z. Shou. Show-o:单一 Transformer 架构统一多模态理解与生成. arXiv 预印本 arXiv:2408.12528, 2024.
- Xu 等人 [2023] J. Xu, S. Liu, A. Vahdat, W. Byeon, X. Wang, 和 S. De Mello. 基于文本到图像扩散模型的开放词汇全景分割. 载于 IEEE/CVF 计算机视觉与模式识别会议论文集, 第 2955–2966 页, 2023.
- Yang 等人 [2024] L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, 和 H. Zhao. Depth Anything V2. 神经信息处理系统进展, 37:21875–21911, 2024.
- Yang 和 Wang [2023] X. Yang 与 X. Wang。扩散模型作为表征学习器。收录于《IEEE/CVF 国际计算机视觉大会论文集》,第 18938–18949 页,2023 年。
- Ye 等人 [2024] C. Ye、L. Qiu、X. Gu、Q. Zuo、Y. Wu、Z. Dong、L. Bo、Y. Xiu 和 X. Han。Stablenormal:通过降低扩散方差实现稳定且锐利的法线。《ACM 图形学汇刊》(ToG),第 43 卷第 6 期,第 1–18 页,2024 年。
- Ye 等人 [2025] Y. Ye、X. He、Z. Li、B. Lin、S. Yuan、Z. Yan、B. Hou 和 L. Yuan。ImgEdit:统一的图像编辑数据集与基准。arXiv 预印本 arXiv:2505.20275,2025 年。
- Yu 等人 [2024] Q. Yu、P.-T. Jiang、H. Zhang、J. Chen、B. Li、L. Zhang 和 H. Lu。通过探测扩散能力实现高精度二分图像分割。arXiv 预印本 arXiv:2410.10105,2024 年。
- Zhai 等人 [2023] X. Zhai、B. Mustafa、A. Kolesnikov 和 L. Beyer。用于语言-图像预训练的 Sigmoid 损失函数。收录于《IEEE/CVF 国际计算机视觉大会论文集》,第 11975–11986 页,2023 年。
- Zhang 等人 [2025] C. Zhang、G. L. Moing、S. Koppula、I. Rocco、L. Momeni、J. Xie、S. Sun、R. Sukthankar、J. K. Barral、R. Hadsell 等。一次一个 D4RT 高效重建动态场景。arXiv 预印本 arXiv:2512.08924,2025 年。
- Zhang 等人 [2023a] H. Zhang、F. Li、X. Zou、S. Liu、C. Li、J. Yang 和 L. Zhang。用于开放词汇分割与检测的简单框架。收录于《IEEE/CVF 国际计算机视觉大会论文集》,第 1020–1031 页,2023a。
- Zhang 等人 [2023b] J. Zhang、C. Herrmann、J. Hur、L. Polania Cabrera、V. Jampani、D. Sun 和 M.-H. Yang。两种特征的故事:Stable Diffusion 补充 DINO 用于零样本语义对应。《神经信息处理系统进展》,第 36 卷,第 45533–45547 页,2023b。
- Zhao 等人 [2025] C. Zhao、Y. Sun、M. Liu、H. Zheng、M. Zhu、Z. Zhao、H. Chen、T. He 和 C. Shen。Diception:用于视觉感知任务的通用扩散模型。arXiv 预印本 arXiv:2502.17157,2025 年。
- Zhao 等人 [2023] W. Zhao、Y. Rao、Z. Liu、B. Liu、J. Zhou 和 J. Lu。释放文本到图像扩散模型用于视觉感知。收录于《IEEE/CVF 国际计算机视觉大会 (ICCV)》,第 5729–5739 页,2023 年。
- Zhou 等人 [2021] J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, 和 T. Kong. iBOT:基于在线分词器的图像 BERT 预训练. arXiv 预印本 arXiv:2111.07832, 2021.
- Zou 等人 [2023] X. Zou, Z.-Y. Dou, J. Yang, Z. Gan, L. Li, C. Li, X. Dai, H. Behl, J. Wang, L. Yuan, 等. 面向像素、图像和语言的广义解码. 见 IEEE/CVF 计算机视觉与模式识别会议论文集, 第 15116–15127 页, 2023.
- Zuo 等人 [2025] J. Zuo, H. Deng, H. Zhou, J. Zhu, Y. Zhang, Y. Zhang, Y. Yan, K. Huang, W. Chen, Y. Deng, R. Jin, N. Sang, 和 C. Gao. Nano Banana Pro 是低层次视觉全能选手吗?基于 14 项任务和 40 个数据集的综合评估. arXiv 预印本, 2025.
附录
附录 A 用于解析实例掩码的多阶段聚类算法
虽然将 RGB 图像后处理为二值掩码对于指代表达分割和语义分割而言是直接的,但对于实例分割来说,该任务更为复杂,因为需要预测的掩码数量事先未知。因此,我们指导模型动态确定掩码数量,并为每个独立实例分配一种独特的颜色。然而,由于输出是生成的图像,它会受到高频生成噪声、单个物体掩码上的轻微颜色漂移以及边界处的混合伪影的影响。标准的连通分量算法或简单的颜色阈值化在这些伪影存在时均会失效。我们提出的算法作为一种鲁棒的后处理方法,确保连续的生成输出能够被清晰地映射回离散的实例掩码。
为了从生成的多色分割图像中解析出这些独立的实例掩码,我们设计了一种多阶段连通分量分组算法。对于任何具有预定义背景颜色 的生成 RGB 掩码图像,该算法按以下步骤进行:
-
背景初始化:我们首先识别并标记所有背景像素。如果一个像素的颜色接近预定义的背景颜色 ,则该像素被分类为背景(并使用标签 进行初始化):
(2) 其中 表示像素 的 RGB 向量, 为颜色容差阈值。背景像素将从后续步骤中排除。
-
颜色相似性分组:我们使用基于种子的泛洪填充算法(16 连通性)将剩余的未标记像素分组为初始组件。如果两个相邻像素到组件种子像素的 RGB 颜色距离满足以下条件,则它们被合并到同一组件中:
(3) -
噪声修剪:为了滤除微小噪声,如果任何组件的总像素面积小于图像尺寸的某个比例,则该组件将被丢弃(重新标记为背景):
(4) 其中 (占图像面积的 0.02%)。
-
边界伪影消除:生成模型通常会在不同纯色物体之间产生细小的彩色边界光晕。为了防止这些光晕形成虚假的独立组件,我们使用大小为 的方形结构元素进行二值腐蚀。如果组件在腐蚀后的保留面积比率低于 ,则该组件被修剪:
(5) -
空间约束合并:如果空间上不相连的组件(可能代表同一物理物体的不同部分,例如由于前景遮挡)的平均颜色相似(),并且它们合并后的边界框面积与各自面积之和相比没有过度扩大,则将其合并:
(6) 其中 是最大边界框扩展因子。
附录 B SA-Co/Gold 详细定量评估
为了更深入地理解我们模型的实例分割能力,我们在表 8 中报告了 SA-Co/Gold [Carion et al., 2025] 基准测试上按类别划分的性能细分。该基准测试评估模型在七个不同子集上的多样化名词短语查询:Metaclip、SA-1B、Crowded、Food&Drink、Sports Equip.、Attributes 和 Wiki-Common。
| 平均 | Metaclip | SA-1B | Crowded | Food&Drink | Sports Equip. | Attributes | Wiki-Common | |||||||||||||||||
| cgF1 | IL_MCC | pmF1 | cgF1 | IL_MCC | pmF1 | cgF1 | IL_MCC | pmF1 | cgF1 | IL_MCC | pmF1 | cgF1 | IL_MCC | pmF1 | cgF1 | IL_MCC | pmF1 | cgF1 | IL_MCC | pmF1 | cgF1 | IL_MCC | pmF1 | |
| 在 SA-Co 上训练 | ||||||||||||||||||||||||
| SAM 3 [Carion et al., 2025] | 54.1 | 0.82 | 66.1 | 47.3 | 0.81 | 58.6 | 53.7 | 0.86 | 62.6 | 61.1 | 0.9 | 67.7 | 53.4 | 0.79 | 67.3 | 65.5 | 0.89 | 73.8 | 54.9 | 0.76 | 72.0 | 42.5 | 0.70 | 60.9 |
| SAM 3 + Llama 3.2 (ft) [Carion et al., 2025] | 61.2 | 0.86 | 70.8 | 54.2 | 0.85 | 64.0 | 56.0 | 0.89 | 62.9 | 61.3 | 0.88 | 69.8 | 67.6 | 0.86 | 78.5 | 67.5 | 0.89 | 75.6 | 71.1 | 0.91 | 77.8 | 51.1 | 0.76 | 67.1 |
| 零样本 | ||||||||||||||||||||||||
| gDino-T [Liu 等人, 2024] | 3.3 | 0.15 | 16.2 | 2.9 | 0.21 | 13.9 | 3.1 | 0.20 | 15.4 | 0.28 | 0.08 | 3.4 | 0.96 | 0.10 | 9.8 | 1.1 | 0.10 | 11.2 | 13.8 | 0.29 | 47.3 | 0.70 | 0.06 | 12.1 |
| OWLv2 [Minderer 等人, 2023] | 24.6 | 0.57 | 42.0 | 17.7 | 0.52 | 34.3 | 13.3 | 0.50 | 26.8 | 15.8 | 0.51 | 30.7 | 32.0 | 0.65 | 49.4 | 36.0 | 0.64 | 56.2 | 35.6 | 0.63 | 56.2 | 21.7 | 0.54 | 40.3 |
| LLMDet-L [Fu 等人, 2025a] | 6.5 | 0.21 | 27.3 | 4.5 | 0.23 | 19.4 | 5.3 | 0.23 | 22.8 | 2.4 | 0.18 | 13.7 | 5.5 | 0.19 | 29.1 | 4.4 | 0.17 | 25.3 | 22.2 | 0.39 | 57.1 | 1.2 | 0.05 | 23.3 |
| APE-D [Shen 等人, 2024] | 16.4 | 0.40 | 36.9 | 12.6 | 0.42 | 30.1 | 2.2 | 0.22 | 10.0 | 7.2 | 0.35 | 20.3 | 22.7 | 0.51 | 45.0 | 31.8 | 0.56 | 56.5 | 26.7 | 0.47 | 57.3 | 11.6 | 0.29 | 39.5 |
| DINO-X [Ren 等人, 2024] | 21.3 | 0.38 | 55.2 | 17.2 | 0.35 | 49.2 | 19.7 | 0.48 | 40.9 | 12.9 | 0.34 | 37.5 | 30.1 | 0.49 | 61.7 | 28.4 | 0.41 | 69.4 | 31.0 | 0.42 | 74.0 | 9.7 | 0.18 | 53.5 |
| Gemini 2.5 [Gemini Team, 2025] | 13.0 | 0.29 | 46.1 | 9.9 | 0.29 | 33.8 | 13.1 | 0.41 | 32.1 | 8.2 | 0.27 | 30.3 | 19.6 | 0.33 | 59.5 | 15.1 | 0.28 | 53.5 | 18.8 | 0.30 | 63.1 | 6.5 | 0.13 | 50.3 |
| Vision Banana + Gemini 3.1 Flash-Lite | 47.5 | 0.84 | 56.0 | 38.0 | 0.81 | 46.7 | 33.5 | 0.84 | 39.8 | 34.4 | 0.83 | 41.2 | 57.2 | 0.88 | 64.8 | 63.0 | 0.92 | 68.2 | 58.8 | 0.86 | 68.7 | 47.4 | 0.76 | 62.3 |
我们采用为 SA-Co 评估建立的官方指标:
-
正向微平均 (): 在所有正向查询(目标名词短语存在于图像中)上计算的微平均分数,评估预测掩码在 中 10 个 IoU 阈值下的精确率和召回率。
-
图像级马修斯相关系数 (): 评估模型预测图像中是否存在名词短语查询的二元分类准确率的指标。
-
分类门控 (): 主要统一指标,结合了 和 :
(7)
在零样本迁移设置下,Vision Banana(搭配 Gemini 3.1 Flash-Lite 进行正/负查询过滤)在该基准测试中展现出最先进的性能,超越了 OWLv2、DINO-X 和 APE-D 等专用模型。可以说,我们的高分可归因于 Gemini 的判别能力,而先前的工作若与多模态大语言模型搭配使用也将从中受益。尽管如此,Vision Banana 在该指标上也取得了最先进的结果,该指标直接衡量开放词汇分割质量,也是本工作的重点。
附录 C 定性实例分割分析
在图 9 和图 10 中,我们展示了 SAM 3 专家模型与我们的零样本 Vision Banana 模型在 SA-Co/Gold [Carion et al., 2025] 基准测试样本上预测的实例分割掩码的定性比较。在每个预测结果下方,我们报告了其在三个独立真实标注中的最佳得分,该得分计算为 SA-Co 基准定义的十个 IoU 阈值上得分的平均值。
我们在图 9 中展示了几个成功的分割结果。例如,在示例 (9.a)、(9.b) 和 (9.c) 中,Vision Banana 分别正确分割了 17 个“修剪和护理过的指甲”实例、3 个“白炽灯泡”实例和 1 个“手绘设计”实例。尽管生成了视觉上准确的分割掩码,但我们的量化得分可能仍然较低,显著低于 SAM 3 的得分。我们发现,这可以解释为与 SAM 标注的真实掩码存在未对齐,这在高 IoU 阈值下严重惩罚了 Vision Banana 的性能。具体来说,我们的模型生成的掩码通常比 SAM 生成的掩码略小(参见示例 (9.a))且更精细(参见示例 (9.c))。
与 SAM 3 这类判别式模型相比——后者在空间模糊情况下会出现“模式平均”伪影(参见示例 (9.b)、(9.c) 和 (10.f))——Vision Banana 通过锁定目标掩码分布中单一且连贯的模式,自然地解决了此类模糊问题。此外,与 SAM 3 不同,我们的模型始终生成不重叠的掩码,并且不会在相邻物体边界之间留下不自然的空隙(参见示例 (9.f))。
在图 10 中,我们报告了模型的失败模式。这些主要表现为:(i) 错误地将实例合并在一起(例如,在 (10.a) 中合并墙面片段,在示例 (10.c) 中合并所有邮筒,在 (10.e) 中合并家具,或在 (10.g) 中合并上下柱子),以及 (ii) 将连贯的“群体实例”分割成多个掩码(例如,在示例 (10.d) 中,尽管查询要求的是“人群”,却将人群分割成单个个体)。此外,Vision Banana 偶尔会完全遗漏目标物体,例如示例 (10.a) 中的背景墙面或示例 (10.d) 中左侧的人物。
这一定性分析表明,我们的模型仍有明显的改进空间,尤其是在高度杂乱或拥挤的场景中。不过,Vision Banana 与 SAM 3 之间部分量化性能差距也归因于 SA-Co 数据集特定的标注偏差,因为我们的模型是在零样本条件下进行评估的。
原始图像 GT a GT b GT c SAM 3 预测结果 Vision Banana 预测结果






(9.a) “浅粉色和白色美甲美足” F1=0.90 (GT a),17 个掩码 F1=0.72 (GT c),17 个掩码






(9.b) “白炽灯泡” F1=0.77 (GT a),4 个掩码 F1=0.60 (GT a),3 个掩码






(9.c) “手绘设计” F1=0.40 (GT a),2 个掩码 F1=0.10 (GT b),1 个掩码






(9.d) “混凝土地面” F1=0.90 (GT c),1 个掩码 F1=0.50 (GT b),1 个掩码






(9.e) “班尼迪克蛋” F1=0.90 (GT a),1 个掩码 F1=0.00 (GT a),1 个掩码






(9.f) “沙丁鱼” F1=0.95 (GT c),8 个掩码 F1=0.91 (GT a),8 个掩码
原始图像 真实标注 a 真实标注 b 真实标注 c SAM 3 预测结果 Vision Banana 预测结果






(10.a) “一面墙“ F1=0.82 (真实标注 b),6 个掩码 F1=0.02 (真实标注 c),2 个掩码






(10.b) “服装橱窗展示“ F1=0.95 (真实标注 a),2 个掩码 F1=0.00 (真实标注 a),3 个掩码






(10.c) “邮政信箱“ F1=1.00 (真实标注 a),28 个掩码 F1=0.00 (真实标注 a),1 个掩码






(10.d) “人群“ F1=1.00 (真实标注 c),1 个掩码 F1=0.00 (真实标注 a),24 个掩码






(10.e) “家具“ F1=0.65 (真实标注 c),18 个掩码 F1=0.00 (真实标注 a),2 个掩码






(10.f) “一个罐子“ F1=0.40 (真实标注 c),1 个掩码 F1=0.00 (真实标注 a),1 个掩码






(10.g) “柱子“ F1=0.85 (真实标注 c),24 个掩码 F1=0.08 (真实标注 c),12 个掩码
附录 D 图像生成与编辑能力
最后,我们评估了 Vision Banana 在指令微调后的生成能力,以确保微调模型并未触发其核心生成特性的灾难性遗忘。图 11 和图 12 将 Vision Banana 与其基础模型 Nano Banana Pro 在文生图和图像编辑任务上进行了比较。
在图 11 中,我们展示了针对从 GenAI-Bench [Li et al., 2024a] 采样的提示词所生成的输出结果。这些结果表明,Vision Banana 能够持续生成细节丰富且上下文准确的图像,其质量与基础生成模型相当。
此外,在图12中,我们使用来自ImgEdit [Ye et al., 2025]的提示词,在两个检查点上评估了基于指令的图像编辑任务。Vision Banana展现出对详细指令的稳健遵循能力,这证实了我们的指令微调成功地为密集视觉任务学习了输出格式,同时没有削弱预训练阶段习得的图像编辑能力。
| 原始图像 | Vision Banana | Nano Banana Pro |
![]() | ![]() | ![]() |
| 将图片中的草坡改为有海浪的沙滩。 | ||
![]() | ![]() | ![]() |
| 移除架子上的植物,并将相框放大。 | ||
![]() | ![]() | ![]() |
| 将车辆的颜色改为红色。 | ||
![]() | ![]() | ![]() |
| 将西装背景从空白墙壁改为豪华办公室环境,包括一张木质办公桌和一扇能看到城市景观的大窗户。 | ||
Valentin Gabeur
project leads and equal contributions
Shangbang Long
project leads and equal contributions
Songyou Peng
project leads and equal contributions
Paul Voigtlaender
Shuyang Sun
Yanan Bao
Karen Truong
Zhicheng Wang
Wenlei Zhou
Jonathan T. Barron
Kyle Genova
Nithish Kannen
Sherry Ben
Yandong Li
Mandy Guo
Suhas Yogin
Yiming Gu
project advisors
Huizhong Chen
project advisors
Oliver Wang
leadership sponsors
Saining Xie
leadership sponsors
Howard Zhou
leadership sponsors
Kaiming He
leadership sponsors
Thomas Funkhouser
leadership sponsors
Jean-Baptiste Alayrac
leadership sponsors
Radu Soricut
leadership sponsors
Abstract
Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how Large Language Models (LLMs) develop emergent capabilities of language understanding and reasoning from generative pretraining. While it has long been conjectured that the ability to create visual content implies an ability to understand it, there has been limited work demonstrating that generalist image generators can achieve state-of-the-art understanding capabilities on many different vision tasks. In this work, we demonstrate that image generation training serves a role similar to LLM pretraining, and lets models learn powerful and general visual representations that enable state-of-the-art performance on various vision tasks. We introduce Vision Banana, a generalist model built by instruction-tuning Nano Banana Pro (NBP) on a mixture of its original training data alongside a small amount of vision task data. By parameterizing the output space of vision tasks as RGB images, we seamlessly reframe perception as image generation. Our generalist model, Vision Banana, achieves state-of-the-art results on a variety of vision tasks involving both 2D and 3D understanding, beating or rivaling zero-shot domain-specialists, including Segment Anything Model 3 on segmentation tasks, and the Depth Anything series on metric depth estimation. We show that these results can be achieved with lightweight instruction-tuning without sacrificing the base model’s image generation capabilities. The superior results suggest that image generation pretraining is a generalist vision learner. It also shows that image generation serves as a unified and universal interface for vision tasks, similar to text generation’s role in language understanding and reasoning. We could be witnessing a major paradigm shift for computer vision, where generative vision pretraining takes a central role in building Foundational Vision Models for both generation and understanding.
Project Page: vision-banana.github.io
1 Introduction
In recent years, advanced image and video generation models [Google, 2025a, b, Black Forest Labs, 2025, ByteDance, 2026, Luma, 2026, OpenAI, 2026] have demonstrated unprecedented generation capabilities, synthesizing highly complex, high-fidelity visual context with precise semantic control. This remarkable capability for visual creation suggests that these models possess a deep, internalized comprehension of the visual world’s underlying structures, semantics, and relationships. However, leading methods on visual representation learning in general do not belong to the family of generative modeling. Instead, they include supervised discriminative learning [Krizhevsky et al., 2012, Dehghani et al., 2023, Dosovitskiy et al., 2020], contrastive learning [Chen et al., 2020b, He et al., 2020, Chen et al., 2020c, Zhai et al., 2023, Tschannen et al., 2025, Radford et al., 2021], bootstrapping [Caron et al., 2021, Grill et al., 2020], auto-encoding [He et al., 2022, Bao et al., 2021, Chen et al., 2024] among others, and their combinations [Oquab et al., 2023, Siméoni et al., 2025, Zhou et al., 2021, Cao et al., 2026]. Early efforts in generative vision pretraining [Chen et al., 2020a, Bai et al., 2024] have shown promising scaling behaviors but their effectiveness has lagged behind non-generative models.
In this paper, we investigate whether visual generative models are secretly generalist vision learners, i.e., whether models trained for image generation develop internal representations that are suitable for visual understanding tasks. To achieve this, we finetune a pretrained image generator with a small amount of computer vision data (depth estimation, surface normal estimation, segmentation, etc.). We then evaluate the resulting model on a wide variety of vision benchmarks. If the finetuned model performs at or near SOTA on these benchmarks, while retaining its image generation capabilities, then there is strong evidence that the image generator was indeed a foundation model for visual understanding – i.e., a generalist vision learner.
This is not the first paper to study the hidden understanding capabilities of generative models or use image and video generators as base models for visual understanding. Early research efforts show that generative models develop some understanding capabilities hidden in their features [Bhattad et al., 2023, Du et al., 2023, Li et al., 2023, Ranzato et al., 2011, Hjelm et al., 2018, Clark and Jaini, 2023, Baranchuk et al., 2021, Chen et al., 2016, Zhao et al., 2023, Mukhopadhyay et al., 2023, Tang et al., 2023, Zhang et al., 2023b, Li et al., 2024b, Hedlin et al., 2023, Yang and Wang, 2023]. More recent research observes that state-of-the-art image and video generators can generate visual content that look like RGB visualizations of computer vision outputs for tasks such as segmentation, depth estimation, and surface normal estimation [Zuo et al., 2025, Wiedemer et al., 2025]. However, those methods do not provide state-of-the-art results on modern benchmarks. This is partially because these models do not strictly follow the prompts to produce vision outputs in the desired formats that can be decoded back to vision outputs for computing quantitative metrics. Other researchers [He et al., 2024, 2025, Ke et al., 2024, Ye et al., 2024, Yu et al., 2024, Zhao et al., 2025, Wang et al., 2026b, Wu et al., 2025, Garcia et al., 2025, Xu et al., 2023] adapt the generation architectures by adding specialized modules and performing full-finetuning to achieve SOTA-level results on specific target tasks. Although these methods successfully leverage the understanding capabilities of the pre-trained features, they sacrifice the model’s generality across other understanding and generation tasks.
, can produce visualizations in a precise format that can enable evaluation on established benchmarks. We take an approach motivated by recent advancements in large language models (LLMs). In natural language processing (NLP), generative pretraining [Brown et al., 2020, Chowdhery et al., 2023] is performed to produce base models, often referred to as LLMs, that are good at generating text, whereas instruction-tuning [Ouyang et al., 2022, Wei et al., 2021] guides them to follow specific tasks and produce text in requested formats and stay on the task. Analogously, we position a visual generative model as a “base” model and perform instruction-tuning to align the model to produce visual output in desired formats, in accordance with the prompts, as illustrated in Fig. 1. Specifically, the model is instructed to produce RGB images that can be decoded to computer vision outputs. Such instruction prompts and decodable visualization schemes are designed to bridge and calibrate the visual generations to formats where measurable metrics for benchmarking can be applied. For example, by prompting the model to “Segment the skateboard category in pure yellow (<255, 255, 0>)”, we can easily parse the mask for skateboard by clustering pixels whose values are close to <255, 255, 0>. This strategy has three main advantages. First, it supports a wide variety of tasks with a single unified model – after instruction tuning, the weights are shared among all tasks, and only the prompt changes. Second, it requires relatively little new training data, since the instruction tuning is solely teaching the model how to format computer vision outputs as RGB. Third, it helps the model retain its original image generation capabilities, since the outputs are simply new RGB images.
| Capabilities | Benchmarks and Metrics | Vision Banana | Best Counterpart |
| 2D Understanding | Referring segmentation: RefCOCOg UMD val (cIoU ) | 73.8 | 73.4 (SAM3 Agent) |
| Referring segmentation: ReasonSeg val (gIoU ) | 79.3 | 77.0 (SAM3 Agent) | |
| Semantic segmentation: Cityscapes val (mIoU ) | 69.9 | 65.2 (SAM3) | |
| Instance segmentation: SA-Co/Gold ( ) | 47.5 | 24.6 (OWLv2) | |
| 3D Understanding | Metric depth estimation: average of 4 datasets () | 0.929 | 0.918 (Depth Anything 3) |
| Surface normal estimation: average of 4 datasets (mean angle error ) | 18.928 | 19.642 (Lotus-2) | |
| Visual Generation | Text-to-image: GenAI-Bench (win rate against the other ) | 53.5% | (Nano Banana Pro) |
| Image editing: ImgEdit (win rate against the other ) | 52.2% (Nano Banana Pro) |
We present Vision Banana
, a generalist vision model trained by performing a lightweight instruction-tuning to Nano Banana Pro on a mixture of its original image generation data and our additional vision task data. During evaluation across several benchmarks, we find that Visual Banana excels at both visual understanding and generation, as summarized in Tab. 1. On the understanding side, Vision Banana surpasses or matches state-of-the-art results on both 2D and 3D tasks. For example, it beats the highly specialized segmentation model, SAM 3 [Carion et al., 2025], on various segmentation tasks, and the 3D expert, Depth Anything 3 [Lin et al., 2025], on metric depth estimation. On the generation side, it performs on par with its base model on image generation and editing benchmarks. On GenAI-Bench [Li et al., 2024a], Vision Banana scores a win rate against its base model. On ImgEdit [Ye et al., 2025] for image editing, Vision Banana’s win rate is . Since these results are achieved with a single unified model built by a lightweight instruction-tuning on its base model, there is strong evidence that Nano Banana Pro already possessed internal representations for visual understanding, which only needed to be unlocked with instruction tuning.
The implications of this study are two-fold. First, it suggests that image generators are indeed generalist vision learners under the hood, with generative vision pretraining playing a foundational role similar to language model pretraining. Second, it suggests that image generation can serve as a universal interface for unified visual understanding, mirroring the role of text generation in language understanding and reasoning. We could be witnessing a major paradigm shift for computer vision, where generative vision pretraining takes a central role in building Foundational Vision Models for both generation and understanding.
2 Method
Instruction-tuning Nano Banana Pro
Recent image and video generators have demonstrated zero-shot capabilities in generating visualizations of visual understanding tasks [Wiedemer et al., 2025, Zuo et al., 2025]. To rigorously investigate and benchmark these capabilities, we need to align the models to generate visualizations that can be decoded back to visual task outputs for quantitative evaluation. For example, in metric depth estimation, a generated depth heatmap must be invertible back to physical depth values for quantitative assessment. Therefore, we create Vision Banana by instruction-tuning our base model, Nano Banana Pro, on a selection of vision tasks formatted in such invertible manners. Specifically, we mix vision task data into Nano Banana Pro’s own training mixture at a very low ratio. This process allows us to align the model’s emergent generative representations into measurable physical geometry and semantic labels, allowing our single generalist model to be evaluated and compared against task-specific specialists.
Mixing the vision data at a low ratio serves as a lightweight instruction-tuning strategy, ensuring that our vision task alignment does not degrade the model’s original generative priors. This strategy distinguishes our work from previous approaches that perform full fine-tuning on generative models without retaining the image generation data mixture [Gan et al., 2023, Ke et al., 2024, Zhao et al., 2025]. We validate the preservation of image generation capabilities by benchmarking Vision Banana against the base Nano Banana Pro on two tasks: text-to-image generation (GenAI-Bench [Li et al., 2024a]) and image editing (ImgEdit [Ye et al., 2025]). In human evaluations, we obtain win rates of 53.5% and 47.8% respectively, indicating that Vision Banana successfully maintains the generative power of its base model. We provide a detailed discussion of these generative capabilities in appendix˜D. Qualitative comparisons in fig.˜11 (text-to-image generation) and fig.˜12 (image editing) in the appendix confirm that the outputs remain highly similar between Vision Banana and Nano Banana Pro. These results verify that Vision Banana does not forget its generative nature.
Vision Tasks and Data.
We evaluate our framework on two fundamental categories of visual understanding: D scene understanding and D structure inference. The D suite consists of referring expression, semantic, and instance segmentation, which collectively test the model’s capability to ground natural language and segment the corresponding objects. For D understanding, we focus on monocular metric depth and surface normal estimation, which demand geometric reasoning and internal knowledge about object scales. To collect data for instruction tuning, we utilize in-house model annotations for web-crawled 2D images, as well as synthetic data from rendering engines for 3D tasks. Crucially, no training data from our evaluation benchmarks is included in the instruction-tuning mixture, ensuring that our results reflect true generalist capability.
3 Vision Banana - Generalist Vision Model from Image Generator
In this section, we present qualitative and quantitative assessments compared to task-specific specialist models. Built upon an image generator, Vision Banana achieves SOTA-level results across a broad range of visual understanding tasks, without specialized architectures or custom training losses.
| Model | mIoU |
| Non Zero-Shot Transfer | |
| SegMan-L [Fu et al., 2025b] | 84.2 |
| Zero-Shot Transfer | |
| APE-D [Shen et al., 2024] | 44.2 |
| OpenSeeD [Zhang et al., 2023a] | 47.8 |
| X-Decoder [Zou et al., 2023] | 52.0 |
| SAM 3 [Carion et al., 2025] | 65.2 |
| Vision Banana | 69.9 |
| Model | IL_MCC | ||
| Non Zero-Shot Transfer | |||
| SAM 3 [Carion et al., 2025] | 54.1 | 0.82 | 66.1 |
| SAM 3 [Carion et al., 2025] + Llama 3.2 (ft) | 61.2 | 0.86 | 70.8 |
| Zero-Shot Transfer | |||
| gDino-T [Liu et al., 2024] | 3.3 | 0.15 | 16.2 |
| LLMDet-L [Fu et al., 2025a] | 6.5 | 0.21 | 27.3 |
| Gemini 2.5 [Gemini Team, 2025] | 13.0 | 0.29 | 46.1 |
| APE-D [Shen et al., 2024] | 16.4 | 0.40 | 36.9 |
| DINO-X [Ren et al., 2024] | 21.3 | 0.38 | 55.2 |
| OWLv2 [Minderer et al., 2023] | 24.6 | 0.57 | 42.0 |
| Vision Banana + Gemini 3.1 Flash-Lite | 47.5 | 0.84 | 56.0 |
| Model | cIoU |
| Non Zero-Shot Transfer | |
| HyperSeg-Phi2-2.7B [Wei et al., 2024] | 79.4 |
| X-SAM-Phi3-3.8B [Wang et al., 2026a] | 83.8 |
| Zero-Shot Transfer | |
| HybridGL [Liu and Li, 2025] | 51.3 |
| LocalizationHeads-LLaVA-1.5-13B [Kang et al., 2025] | 67.7 |
| SAM 3 [Carion et al., 2025] + Gemini 2.5 Pro | 73.4 |
| Vision Banana | 73.8 |
| Model | gIoU |
| Non Zero-Shot Transfer | |
| X-SAM-Phi-3-3.8B [Wang et al., 2026a] | 56.6 |
| LISA-13B-LLaVA1.5 [Lai et al., 2024] | 65.0 |
| Zero-Shot Transfer | |
| SegZero-Qwen2.5-VL-7B [Liu et al., 2025] | 62.6 |
| RSVP-GPT-4o [Lu et al., 2025] | 64.7 |
| SAM 3 [Carion et al., 2025] + Gemini 2.5 Pro | 77.0 |
| Vision Banana + Gemini 2.5 Pro | 79.3 |
3.1 2D Semantic Understanding
Image segmentation stands as a cornerstone of visual understanding, traditionally requiring complex, task-specific models to classify pixels into semantic categories or object instances. Current leading methods such as the Segment Anything series [Kirillov et al., 2023, Ravi et al., 2024, Carion et al., 2025] tackle this through heavy architectural specialization and large volumes of expensive, human-annotated mask data. Vision Banana challenges this prevailing paradigm by demonstrating that SOTA segmentation can naturally emerge from image generation pretraining. Rather than training on vast amounts of meticulously crafted segmentation examples, we tap into the rich representations learned by the base image generation model. By instructing the model to generate multi-colored images of segmentation masks, we obtain dense segmentation maps from which individual masks can be decoded, therefore enabling segmentation through image generation. As detailed in Tab. 5, 5, 5 and 5, this elegant generative approach outperforms highly tuned specialist models, achieving SOTA zero-shot transfer performance on all evaluated segmentation benchmarks. We compare with other methods that have not been trained on in-domain data, i.e., the training splits of these benchmarks. We denote them as “Zero-Shot Transfer” in the tables. The usage of this term follows Segment Anything [Kirillov et al., 2023] and CLIP [Radford et al., 2021]. Non zero-shot transfer methods are marked in gray.
Semantic Segmentation.
Semantic segmentation involves classifying each pixel into a predefined category without distinguishing between individual instances. For example, the Cityscapes benchmark [Cordts et al., 2016] defines classes, including road, person, and sky. While instance and referring expression segmentation also convey semantic information, we use the term “semantic segmentation” here strictly in this instance-agnostic, category-level sense. This nature of the classical semantic segmentation task can be specified via a text prompt, and we train the model to follow such instructions. We prompt the model to generate a visualization image where each pixel is colored according to its class, as shown in Fig. 2.
Crucially, our approach is open-vocabulary: the target categories are not limited to a fixed set and can be specified dynamically in the prompt along with their corresponding color mappings. We support various prompting styles, including natural language descriptions (e.g., “the macaron cakes are represented by yellow”) and structured JSON mappings, with colors specified as named colors, hex codes, or RGB tuples. For quantitative evaluation, we post-process the generated image by assigning each pixel to the class whose target color is closest in the RGB space.
We compare Vision Banana with existing methods on the Cityscapes validation set in table˜5. During evaluation, we use the same text prompt for each example, providing the full class-to-color mapping for the 19 classes, including the ones not present in the image. As shown in table˜5, Vision Banana outperforms SAM 3 by points in mIoU and achieves the best performance among open-vocabulary models, narrowing the gap with closed-set, non-zero-shot specialists like SegMan [Fu et al., 2025b].
Instance Segmentation.
Unlike semantic segmentation, instance segmentation requires the model to distinguish between individual objects that belong to the same class. For example, if an image contains five dogs, we expect the model to produce an individual mask for each animal. This poses a unique challenge for Vision Banana: since the number of instances is unknown a priori, we cannot assign specific colors in the prompt beforehand. To address this challenge, we prompt the model with only the target class and the background color, instructing it to assign a unique, distinct color to each individual instance. We let the model dynamically assign distinct colors to different instances of that class. Qualitative examples are shown in fig.˜3. Individual instance masks can be extracted from the generated RGB images using a multi-stage clustering algorithm, detailed in appendix˜A of the appendix.
We evaluate our model on the open-vocabulary noun-phrase (NP) instance segmentation benchmark SA-Co/Gold [Carion et al., 2025]. We summarize the main results in table˜5 and provide a comprehensive category breakdown in table˜8, appendix˜B of the appendix. We also perform a qualitative evaluation in appendix˜C of the appendix. The SA-Co/Gold benchmark consists of 168k Image-NP pairs, where a large majority are negative queries (i.e., the target NP is absent from the image). While Vision Banana could theoretically handle these negative queries by generating a solid black image, we did not tune the model to generate such empty mask images. We instead defer the task of classifying the Image-NP pairs as positives or negatives to a MLLM. To do so, we prompt Gemini 3.1 Flash-Lite with “Is there an instance of <NP> visible in this image? Choose your answer from the following options: (A): Yes, (B): No.”. We then only process the positive-predicted examples with Vision Banana to generate an image of the masks. This approach is related to the "SAM 3 + Llama 3.2 (ft)" method, presented as "SAM 3 + EV" in Carion et al. [2025] where they fine-tuned Llama 3.2 to produce a presence score for SAM 3.
Under the zero-shot transfer setting, Vision Banana (paired with Gemini 3.1 Flash-Lite) achieves state-of-the-art performance, outperforming existing models including Gemini 2.5 [Gemini Team, 2025], APE-D [Shen et al., 2024], DINO-X [Ren et al., 2024], and OWLv2 [Minderer et al., 2023]. While Vision Banana still lags behind the SAM 3 specialist on SA-Co/Gold, we emphasize that unlike SAM 3, we did not include the SA-Co dataset in our training data mixture.
Referring Expression Segmentation.
Unlike traditional fixed-class segmentation, referring expression segmentation evaluates a model’s ability to segment objects described by long, free-form natural language queries. This task requires models to comprehend and reason about nuanced natural language expressions, as well as capture complex relationships between objects. As summarized in tables˜5 and 5, our model achieves state-of-the-art performance under the zero-shot transfer setting, obtaining a cIoU of on RefCOCOg UMD [Kazemzadeh et al., 2014] and a gIoU of on ReasonSeg [Lai et al., 2024]. It consistently outperforms SAM 3 Agent [Carion et al., 2025] (which pairs SAM 3 with Gemini 2.5 Pro) and other recent zero-shot methods, including HybridGL [Liu and Li, 2025], LocalizationHeads [Kang et al., 2025], SegZero [Liu et al., 2025], and RSVP [Lu et al., 2025]. On RefCOCOg, a performance gap remains compared to methods that are trained on the training split like HyperSeg [Wei et al., 2024] and X-SAM [Wang et al., 2026a].
For the complex reasoning queries in ReasonSeg, we follow standard practice by delegating the reasoning step to a multimodal LLM. Specifically, we utilize Gemini 2.5 Pro to translate the reasoning query into a descriptive reference, which then serves as the prompt for Vision Banana. We evaluate this pipeline in a single-turn inference setup, where both Gemini and Vision Banana are queried exactly once. In this setting, Vision Banana paired with Gemini 2.5 Pro outperforms several non-zero-shot methods that were trained directly on ReasonSeg, including X-SAM [Wang et al., 2026a] and LISA [Lai et al., 2024]. Qualitative results in fig.˜4 illustrate Vision Banana’s ability to ground diverse language cues, from physical actions (“stretching cat”) and atypical object roles (“toaster as a game controller”) to multilingual text on signage. This highlights a key advantage of our approach: the rich multimodal priors inherited from generative pre-training allow Vision Banana to reason about ’what’ to segment more effectively than specialized segmentation models.
Intriguingly, Vision Banana also exhibits strong cross-task transfer by demonstrating a similar grasp of referring expressions combined with standard semantic and instance segmentation tasks, despite not being explicitly trained on free-form queries for those tasks. For example, in Fig. 2(b) (right), the model understands what “patterns on the wall” is referring to. In Fig. 3(b) (right), the model successfully distinguishes crescent-shaped croissants from other variations of croissants. These findings suggest that our generative pre-training yields highly robust, transferable representations across distinct visual grounding paradigms.
3.2 3D Understanding from Monocular Images
Vision Banana demonstrates a strong ability to infer 3D structures from 2D monocular images. We evaluate this capability on two classical tasks: monocular metric depth estimation and surface normal estimation. As summarized in Tab. 1, Vision Banana achieves SOTA performance on both tasks, surpassing specialists such as Depth Anything V3 [Lin et al., 2025] and Lotus-2 [He et al., 2025].
Metric Depth Estimation.
The goal of depth estimation is to produce a depth map from a monocular image, where each pixel’s value represents the physical metric distance from the camera plane to the observed object [Eigen et al., 2014]. This is a fundamental computer vision task that benefits a wide range of applications such as robotics, augmented/virtual reality, and autonomous driving. However, depth estimation is inherently ill-posed, as 2D projections inherently discard critical 3D geometric information. Furthermore, monocular depth estimation is particularly challenging due to the absence of parallax cues available in multi-view setups, even when camera intrinsic parameters are known.
In the deep-learning era, the research community has largely framed depth estimation as a dense per-pixel supervised regression problem, employing specialized architectures and domain-specific loss functions. Most recent SOTA methods rely on camera intrinsics during training, inference, or both [Yang et al., 2024, Bochkovskii et al., 2024, Wang et al., 2025b, c, He et al., 2025, 2024, Hu et al., 2024, Cai et al., 2025, Lin et al., 2025, Piccinelli et al., 2025b, a]. While using intrinsics mitigates the inherent ambiguity of depth estimation, it also necessitates specialized model designs. In contrast, our work is predicated on the hypothesis that the mode-seeking nature of generative modeling naturally resolves training target ambiguities, thereby eliminating the need for such specialized techniques. Furthermore, the broad world knowledge acquired during pretraining endows the model with stronger priors on object sizes and distances compared to narrowly targeted models. To enable Nano Banana Pro to estimate depth in metric units, we instruct the model to output a carefully constructed false-color visualization of depth values.
To visualize depth maps as RGB images, we establish a mapping between unbounded depth values in and bounded RGB values in . Because the utility of accurate metric depth for nearby image content is generally higher than that of distant content (e.g., graspable objects matter more for robotics tasks, stereo/monodepth benchmarks usually measure accuracy terms of disparity or relative/log-depth) we “curve” metric depth prior to RGB encoding. Specifically, this is achieved by first applying the power transform of Barron [2025] to warp the depth values, and then using those curved distances to produce a false-color visualization. We constrain the power transform to and rescale it to map metric distances to normalized distances in :
| (1) |
In all experiments, we set the shape parameter to and the scale parameter to . These curved and normalized distances are then used to interpolate along a piecewise-linear function that follows the edges of the RGB cube, traversing along its edges from black to white, similarly to the first iteration of a 3D Hilbert curve. A visualization of this process is provided in Fig. 5.
This mapping from normalized distance to RGB color can be inverted by simply projecting the RGB values onto the nearest line segment and then inverting the linear interpolation along the cube’s edges. Because both the false-color visualization and the power transform are strictly invertible, their composition forms a bijection between metric depth in and RGB space in . During training, we apply this mapping to ground-truth metric depths to generate RGB training targets. At inference, we apply the inverse mapping to decode the model’s generated RGB images back into metric depth, enabling direction evaluation on standard depth benchmarks. To enhance the model’s robustness across diverse color representations, we augment our training data with alternative color maps, such as Plasma, Inferno, Viridis, and grayscale.
| DepthLM-7B [Cai et al., 2025] | Depth Any. v3 [Lin et al., 2025] | Depth Pro [Bochkovskii et al., 2024] | UniK3D [Piccinelli et al., 2025a] | MoGe-2 [Wang et al., 2025c] | Vision Banana | ||
| Camera Intrinsics | Inference | ✓ | ✓ | ||||
| Training | ✓ | ✓ | ✓ | ✓ | ✓ | ||
| Average | partial* | partial* | 0.715 | 0.823 | 0.802 | 0.882 | |
| Benchmarks | AbsRel | 0.156 | 0.144 | 0.116 | |||
| NYU | 0.915 | 0.963 | 0.961 | 0.965 | 0.961 | 0.948 | |
| [Silberman et al., 2012] | AbsRel | 0.07 | 0.074 | 0.0733 | 0.081 | ||
| iBims1 | 0.92 | 0.913 | 0.919 | 0.830 | 0.934 | ||
| [Koch et al., 2018] | AbsRel | 0.104 | 0.136 | 0.078 | |||
| ETH3D | 0.718 | 0.917 | 0.415 | 0.687 | 0.908 | 0.935 | |
| [Schops et al., 2019] | AbsRel | 0.104 | 0.327 | 0.236 | 0.104 | 0.103 | |
| DIODE-Indoor | 0.838 | 0.671 | 0.713 | 0.664 | 0.917 | ||
| [Vasiljevic et al., 2019] | AbsRel | 0.123 | 0.199 | 0.161 | 0.175 | 0.108 | |
| KITTI | 0.953 | 0.843 | 0.812 | 0.629 | 0.915 | ||
| [Uhrig et al., 2017] | AbsRel | 0.086 | 0.121 | 0.174 | 0.181 | 0.107 | |
| nuScenes | 0.865 | 0.491 | 0.840 | 0.820 | 0.643 | ||
| [Caesar et al., 2020] | AbsRel | 0.287 | 0.189 | 0.195 | 0.219 | ||
* The average of DepthLM-7B on the 4 datasets it evaluated on (NYU + iBims1 + ETH3D + nuScenes) is ; our average on the same 4 datasets is . The average of Depth-Anything V3 on the 4 datasets it evaluated on (NYU + ETH3D + DIODE + KITTI) is ; our average on the same 4 datasets is .
DepthLM is trained on nuScenes so it’s not zero-shot.
Numbers reported by Depth-Anything V3 [Lin et al., 2025].
Tab. 6 presents the empirical results of Vision Banana compared to specialist models across six major academic benchmarks. Vision Banana achieves an average accuracy of 0.882, outperforming Unik3D [Piccinelli et al., 2025a] by nearly 6 points, while achieving a 20% lower absolute relative error (AbsRel) compared to MoGe-2 [Wang et al., 2025c]. Notably, Vision Banana outperforms Depth Anything V3 [Lin et al., 2025] on average across the four datasets (NYU, ETH3D, DIODE, KITTI) on which it was evaluated ( v.s. ), demonstrating robust performance in both near-field and distant scenes. Our model is trained entirely on synthetic depth data created from simulation engines — we use zero real-world depth data, and exclude training data from any of the depth datasets we evaluate on. Note that this result is achieved without relying on camera parameters (neither intrinsics nor extrinsics) during both training or inference. By leveraging the immense geometric priors embedded in its foundation model, Vision Banana infers absolute scale solely from visual cues and object relationships, enabling zero-shot generalization to any arbitrary input image.
![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() |
| Input image | Generated depth image | Vis. view 1 | Vis. view 2 |
Qualitative inspections further validate the model’s capabilities. As illustrated in Fig. 6, Vision Banana generates highly precise depth maps that preserve crisp geometric details, even in cluttered environments like classrooms. When these 2D predictions are unprojected into 3D point clouds, they exhibit global consistency across diverse scenes, maintaining accurate planar surfaces and correct geometry. In addition to common academic benchmarks, we also conducted a “vibe test” using a casual smartphone photograph, as shown in Fig. 7. Crossed validated by depth measured on Google Maps, Vision Banana successfully produced an accurate depth estimation on this photo captured by a consumer device unseen during training.
Surface Normal Estimation.
Surface normal estimation represents another critical vision task. Surface normals, which are unit vectors with values ranging from to , serve as a critical proxy for local geometry and scene structures. Unlike the complex color mapping required for metric depth, the visualization of surface normals is intrinsically aligned with the RGB color space, allowing straightforward integration into our model.
We specifically utilize a camera-space normal formulation using the standard right-handed coordinate system (+x right, +y up, +z pointing out of the image plane). In this representation, the directional vector components map directly to RGB channels, i.e. , , :
-
Facing Left : Encoded as Pinkish Red.
-
Facing Up : Encoded as Light Green.
-
Facing the Camera : Encoded as Light Blue.
Table 7 compares Vision Banana against SOTA specialist methods on four public benchmarks. When averaged across the three indoor datasets, Vision Banana achieves the lowest mean and median angular errors. It also demonstrates competitive accuracy on outdoor scenes.
| Methods | Indoor | Indoor | Outdoor | |||||||
| Average | NYUv2 [Silberman et al., 2012] | DIODE-indoor [Vasiljevic et al., 2019] | ScanNet [Dai et al., 2017] | VKitti [Cabon et al., 2020] | ||||||
| mean | median | mean | median | mean | median | mean | median | mean | median | |
| Marigold [Ke et al., 2024] | 19.606 | 11.828 | 20.864 | 11.134 | 16.671 | 12.084 | 21.284 | 12.268 | – | – |
| DSINE [Bae and Davison, 2024] | 17.017 | 10.190 | 16.4 | 8.4 | 18.453 | 13.871 | 16.2 | 8.3 | 28.9 | 9.9 |
| StableNormal [Ye et al., 2024] | 17.168 | 10.028 | 19.707 | 10.527 | 13.701 | 9.46 | 18.098 | 10.097 | – | – |
| Lotus-2-Normal [He et al., 2025] | 16.558 | – | 16.9 | N/A | 18.575 | N/A | 14.2 | N/A | 28.894 | 9.677 |
| Vision Banana | 15.549 | 9.300 | 17.778 | 8.876 | 13.818 | 11.556 | 15.052 | 7.468 | 29.063 | 10.699 |
![]() | ![]() | ![]() |
![]() | ![]() | ![]() |
| Input image | Lotus-2-Normal | Vision Banana |
Figure 8 visually compares the output from Vision Banana with the leading external method, Lotus-2 [He et al., 2025]. Vision Banana consistently produces surface normal maps with significantly higher fidelity and finer granular details. The bottom row of Fig. 8 highlights a sample from Virtual KITTI 2 [Cabon et al., 2020]. Although Vision Banana registers slightly higher quantitative errors on this benchmark compared to Lotus-2, it yields demonstrably superior visual quality. Also note that Lotus-2 is trained on Virtual KITTI 2 for surface normal estimation, whereas Vision Banana maintains a strict zero-shot transfer protocol, having never seen the training sets of any evaluated benchmarks.
4 Discussion
Image Generators are Generalist Vision Learners.
Generative pretraining [Radford et al., 2018, 2019, Brown et al., 2020] has fundamentally transformed language understanding and reasoning. In the meantime, recent observations of emergent vision capabilities [Wiedemer et al., 2025, Zuo et al., 2025] have ignited speculation that computer vision is approaching a similar paradigm shift. By instruction-tuning a leading image generator, Nano Banana Pro, into a state-of-the-art visual generation and understanding model, we confirm that this shift is already underway. Models pretrained on large-scale image generation naturally acquire robust visual understanding capabilities. These generative priors surpass the specialized architectures and dedicated training paradigms traditionally employed by specialist vision models. We are witnessing a paradigm shift for computer vision that will be fueled by generative vision pretraining, which we believe paves the way for true Foundational Vision Models and Artificial General Intelligence from Vision (AGI-V).
Image Generation as a Universal Interface.
As a byproduct of this study, we show that image generation can serve as the universal interface for computer vision, analogous to how text generation acts as the unifying interface for many tasks embedded in natural language, including language understanding, generation, reasoning, math, coding, agentic tasks, etc.. By representing vision task outputs as RGB images, we can use natural language prompts to seamlessly instruct the model. While we are not the first to encode vision outputs as RGB [Ke et al., 2024, Zhao et al., 2025, Gan et al., 2023, Wang et al., 2023, Lu et al., 2022, 2024, Xie et al., 2024, Inclusion AI, 2025], we demonstrate that when combined with powerful pretrained visual generators, this simple design is sufficient to outperform modern domain-specific specialist models.
In addition to the unification of vision task outputs as RGB images, generative modeling naturally provides a workaround for the ambiguity in vision tasks where a single input can correspond to several modes of the output distribution. In order to prevent the collapse of the output to a blurry mean, expert discriminative models [Carion et al., 2025, Lin et al., 2025] usually resort to custom architectures and training losses. For example, the Segment Anything models [Kirillov et al., 2023, Ravi et al., 2024, Carion et al., 2025] return several segmentation masks but only apply the loss to a single one. Generative models, however, inherently learn the full data distribution, gracefully managing ambiguity by design. By eliminating the need for bespoke architectural designs, this formulation could lead to a truly unified “omni” multimodal model.
Future Work.
While Vision Banana achieves SOTA results on fundamental tasks for 2D semantic understanding and 3D understanding from monocular images, several exciting avenues remain for future exploration. First, scaling the diversity of instruction-tuned tasks may unlock further emergent cross-task generalization, similar to behaviors observed in LLMs [Wei et al., 2021]. Second, our current evaluation focuses on monocular image inputs. In the future, we can extend this framework to process multi-view inputs [Wang et al., 2025a] and video inputs [Zhang et al., 2025]. Similarly, investigating whether video generators yield even richer, temporally-aware visual representations presents a highly promising research direction. Another important next step is exploring the synergistic integration of foundational vision models with large language models to enhance cross-modality reasoning. Finally, utilizing image generators like Nano Banana Pro currently incurs a significantly higher computational overhead than running lightweight specialist models. Developing acceleration and cost-reduction strategies will be an essential hurdle to overcome for the deployment of generative vision framework.
Acknowledgment
We thank Xi Chen, Fei Xia, Kaushik Shivakumar, Abhishek Sinha, Phillip Lippe, Yilin Gao, Javier Rey, Sanghyun Woo, Renshen Wang, Wentao Yuan, Keran Rong, Rundi Wu, Manoj Kumar, Manli Shu, Francesco Piccinno, Ishita Dasgupta, Benigno Uria, Miki Rubinstein, Aäron van den Oord, Jon Shlens for their helpful discussions, advice, and technical guidance.
References
- Bae and Davison [2024] G. Bae and A. J. Davison. Rethinking inductive biases for surface normal estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9535–9545, 2024.
- Bai et al. [2024] Y. Bai, X. Geng, K. Mangalam, A. Bar, A. L. Yuille, T. Darrell, J. Malik, and A. A. Efros. Sequential modeling enables scalable learning for large vision models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22861–22872, 2024.
- Bao et al. [2021] H. Bao, L. Dong, S. Piao, and F. Wei. Beit: Bert pre-training of image transformers. arXiv preprint arXiv:2106.08254, 2021.
- Baranchuk et al. [2021] D. Baranchuk, I. Rubachev, A. Voynov, V. Khrulkov, and A. Babenko. Label-efficient semantic segmentation with diffusion models. arXiv preprint arXiv:2112.03126, 2021.
- Barron [2025] J. T. Barron. A power transform. arXiv preprint arXiv:2502.10647, 2025.
- Bhattad et al. [2023] A. Bhattad, D. McKee, D. Hoiem, and D. Forsyth. Stylegan knows normal, depth, albedo, and more. Advances in Neural Information Processing Systems, 36:73082–73103, 2023.
- Black Forest Labs [2025] Black Forest Labs. FLUX.2: Frontier Visual Intelligence. https://bfl.ai/blog/flux-2, 2025.
- Bochkovskii et al. [2024] A. Bochkovskii, A. Delaunoy, H. Germain, M. Santos, Y. Zhou, S. R. Richter, and V. Koltun. Depth pro: Sharp monocular metric depth in less than a second. arXiv preprint arXiv:2410.02073, 2024.
- Brown et al. [2020] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
- ByteDance [2026] ByteDance. Seedance 2.0. https://seed.bytedance.com/en/seedance2_0/, 2026. Accessed: 2026-03-18.
- Cabon et al. [2020] Y. Cabon, N. Murray, and M. Humenberger. Virtual kitti 2. arXiv preprint arXiv:2001.10773, 2020.
- Caesar et al. [2020] H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom. nuscenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11621–11631, 2020.
- Cai et al. [2025] Z. Cai, C.-F. Yeh, H. Xu, Z. Liu, G. Meyer, X. Lei, C. Zhao, S.-W. Li, V. Chandra, and Y. Shi. Depthlm: Metric depth from vision language models. arXiv preprint arXiv:2509.25413, 2025.
- Cao et al. [2026] B. Cao, K. Chen, K.-K. Maninis, K. Chen, A. Karpur, Y. Xia, S. Dua, T. Dabral, G. Han, B. Han, et al. Tipsv2: Advancing vision-language pretraining with enhanced patch-text alignment. arXiv preprint arXiv:2604.12012, 2026.
- Carion et al. [2025] N. Carion, L. Gustafson, Y.-T. Hu, S. Debnath, R. Hu, D. Suris, C. Ryali, K. V. Alwala, H. Khedr, A. Huang, et al. Sam 3: Segment anything with concepts. arXiv preprint arXiv:2511.16719, 2025.
- Caron et al. [2021] M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021.
- Chen et al. [2020a] M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever. Generative pretraining from pixels. In International conference on machine learning, pages 1691–1703. PMLR, 2020a.
- Chen et al. [2020b] T. Chen, S. Kornblith, M. Norouzi, and G. Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PmLR, 2020b.
- Chen et al. [2016] X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, and P. Abbeel. Infogan: Interpretable representation learning by information maximizing generative adversarial nets. Advances in neural information processing systems, 29, 2016.
- Chen et al. [2020c] X. Chen, H. Fan, R. Girshick, and K. He. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020c.
- Chen et al. [2024] X. Chen, Z. Liu, S. Xie, and K. He. Deconstructing denoising diffusion models for self-supervised learning. arXiv preprint arXiv:2401.14404, 2024.
- Chowdhery et al. [2023] A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al. Palm: Scaling language modeling with pathways. Journal of machine learning research, 24(240):1–113, 2023.
- Clark and Jaini [2023] K. Clark and P. Jaini. Text-to-image diffusion models are zero shot classifiers. Advances in Neural Information Processing Systems, 36:58921–58937, 2023.
- Cordts et al. [2016] M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele. The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3213–3223, 2016.
- Dai et al. [2017] A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5828–5839, 2017.
- Dehghani et al. [2023] M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. P. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, et al. Scaling vision transformers to 22 billion parameters. In International conference on machine learning, pages 7480–7512. PMLR, 2023.
- Dosovitskiy et al. [2020] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
- Du et al. [2023] X. Du, N. Kolkin, G. Shakhnarovich, and A. Bhattad. Generative models: What do they know? do they know things? let’s find out! arXiv preprint arXiv:2311.17137, 2023.
- Eigen et al. [2014] D. Eigen, C. Puhrsch, and R. Fergus. Depth map prediction from a single image using a multi-scale deep network. In Advances in Neural Information Processing Systems (NeurIPS), volume 27, 2014.
- Fu et al. [2025a] S. Fu, Q. Yang, Q. Mo, J. Yan, X. Wei, J. Meng, X. Xie, and W.-S. Zheng. Llmdet: Learning strong open-vocabulary object detectors under the supervision of large language models. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 14987–14997, 2025a.
- Fu et al. [2025b] Y. Fu, M. Lou, and Y. Yu. Segman: Omni-scale context modeling with state space models and local attention for semantic segmentation. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 19077–19087, 2025b.
- Gan et al. [2023] Y. Gan, S. Park, A. Schubert, A. Philippakis, and A. M. Alaa. Instructcv: Instruction-tuned text-to-image diffusion models as vision generalists. arXiv preprint arXiv:2310.00390, 2023.
- Garcia et al. [2025] G. M. Garcia, K. A. Zeid, C. Schmidt, D. De Geus, A. Hermans, and B. Leibe. Fine-tuning image-conditional diffusion models is easier than you think. In Proceedings of the Winter Conference on Applications of Computer Vision, pages 753–762, 2025.
- Gemini Team [2025] Gemini Team. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint, 2025.
- Google [2025a] Google. Introducing nano banana pro. https://blog.google/innovation-and-ai/products/nano-banana-pro/, 2025a. Accessed: 2026-03-15.
- Google [2025b] Google. Veo 3 announcement. https://blog.google/innovation-and-ai/products/generative-media-models-io-2025/, 2025b. Accessed: 2026-03-15.
- Grill et al. [2020] J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. Advances in neural information processing systems, 33:21271–21284, 2020.
- He et al. [2024] J. He, H. Li, W. Yin, Y. Liang, L. Li, K. Zhou, H. Zhang, B. Liu, and Y.-C. Chen. Lotus: Diffusion-based visual foundation model for high-quality dense prediction. arXiv preprint arXiv:2409.18124, 2024.
- He et al. [2025] J. He, H. Li, M. Sheng, and Y.-C. Chen. Lotus-2: Advancing geometric dense prediction with powerful image generative model. arXiv preprint arXiv:2512.01030, 2025.
- He et al. [2020] K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9729–9738, 2020.
- He et al. [2022] K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16000–16009, 2022.
- Hedlin et al. [2023] E. Hedlin, G. Sharma, S. Mahajan, H. Isack, A. Kar, A. Tagliasacchi, and K. M. Yi. Unsupervised semantic correspondence using stable diffusion. Advances in Neural Information Processing Systems, 36:8266–8279, 2023.
- Hjelm et al. [2018] R. D. Hjelm, A. Fedorov, S. Lavoie-Marchildon, K. Grewal, P. Bachman, A. Trischler, and Y. Bengio. Learning deep representations by mutual information estimation and maximization. arXiv preprint arXiv:1808.06670, 2018.
- Hu et al. [2024] M. Hu, W. Yin, C. Zhang, Z. Cai, X. Long, H. Chen, K. Wang, G. Yu, C. Shen, and S. Shen. Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.
- Inclusion AI [2025] Inclusion AI. Ming-flash-omni: A sparse, unified architecture for multimodal perception and generation. arXiv preprint arXiv:2510.24821, 2025.
- Kang et al. [2025] S. Kang, J. Kim, J. Kim, and S. J. Hwang. Your large vision-language model only needs a few attention heads for visual grounding. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 9339–9350, 2025.
- Kazemzadeh et al. [2014] S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg. Referitgame: Referring to objects in photographs of natural scenes. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 787–798, 2014.
- Ke et al. [2024] B. Ke, A. Obukhov, S. Huang, N. Metzger, R. C. Daudt, and K. Schindler. Repurposing diffusion-based image generators for monocular depth estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9492–9502, 2024.
- Kirillov et al. [2023] A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick. Segment anything. In IEEE/CVF International Conference on Computer Vision (ICCV), pages 4015–4026, 2023.
- Koch et al. [2018] T. Koch, L. Liebel, F. Fraundorfer, and M. Korner. Evaluation of cnn-based single-image depth estimation methods. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops, pages 0–0, 2018.
- Krizhevsky et al. [2012] A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012.
- Lai et al. [2024] X. Lai, Z. Tian, Y. Chen, Y. Li, Y. Yuan, S. Liu, and J. Jia. Lisa: Reasoning segmentation via large language model. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9579–9589, 2024.
- Li et al. [2023] A. C. Li, M. Prabhudesai, S. Duggal, E. Brown, and D. Pathak. Your diffusion model is secretly a zero-shot classifier. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2206–2217, 2023.
- Li et al. [2024a] B. Li, Z. Lin, D. Pathak, J. Li, Y. Fei, K. Wu, T. Ling, X. Xia, P. Zhang, G. Neubig, et al. Genai-bench: Evaluating and improving compositional text-to-visual generation. arXiv preprint arXiv:2406.13743, 2024a.
- Li et al. [2024b] X. Li, J. Lu, K. Han, and V. A. Prisacariu. Sd4match: Learning to prompt stable diffusion model for semantic matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 27558–27568, 2024b.
- Lin et al. [2025] H. Lin, S. Chen, J. Liew, D. Y. Chen, Z. Li, G. Shi, J. Feng, and B. Kang. Depth anything 3: Recovering the visual space from any views. arXiv preprint arXiv:2511.10647, 2025.
- Liu et al. [2024] S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su, J. Zhu, and L. Zhang. Grounding DINO: marrying DINO with grounded pre-training for open-set object detection. In ECCV (47), Lecture Notes in Computer Science, pages 38–55. Springer, 2024.
- Liu and Li [2025] T. Liu and S. Li. Hybrid global-local representation with augmented spatial guidance for zero-shot referring image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 29634–29643, 2025.
- Liu et al. [2025] Y. Liu, B. Peng, Z. Zhong, Z. Yue, F. Lu, B. Yu, and J. Jia. Seg-zero: Reasoning-chain guided segmentation via cognitive reinforcement. arXiv preprint arXiv:2503.06520, 2025.
- Lu et al. [2022] J. Lu, C. Clark, R. Zellers, R. Mottaghi, and A. Kembhavi. Unified-io: A unified model for vision, language, and multi-modal tasks. arXiv preprint arXiv:2206.08916, 2022.
- Lu et al. [2024] J. Lu, C. Clark, S. Lee, Z. Zhang, S. Khosla, R. Marten, D. Hoiem, and A. Kembhavi. Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26439–26455, 2024.
- Lu et al. [2025] Y. Lu, J. Cao, Y. Wu, B. Li, L. Tang, Y. Ji, C. Wu, J. Wu, and W. Zhu. Rsvp: Reasoning segmentation via visual prompting and multi-modal chain-of-thought. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14699–14716, 2025.
- Luma [2026] Luma. UNI-1. https://lumalabs.ai/uni-1/, 2026. Accessed: 2026-03-19.
- Minderer et al. [2023] M. Minderer, A. Gritsenko, and N. Houlsby. Scaling open-vocabulary object detection. Advances in Neural Information Processing Systems, 36:72983–73007, 2023.
- Mukhopadhyay et al. [2023] S. Mukhopadhyay, M. Gwilliam, V. Agarwal, N. Padmanabhan, A. Swaminathan, S. Hegde, T. Zhou, and A. Shrivastava. Diffusion models beat gans on image classification. arXiv preprint arXiv:2307.08702, 2023.
- OpenAI [2026] OpenAI. GPT-Image-1.5. https://openai.com/index/new-chatgpt-images-is-here/, 2026. Accessed: 2026-03-19.
- Oquab et al. [2023] M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023.
- Ouyang et al. [2022] L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730–27744, 2022.
- Piccinelli et al. [2025a] L. Piccinelli, C. Sakaridis, M. Segu, Y.-H. Yang, S. Li, W. Abbeloos, and L. Van Gool. Unik3d: Universal camera monocular 3d estimation. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 1028–1039, 2025a.
- Piccinelli et al. [2025b] L. Piccinelli, C. Sakaridis, Y.-H. Yang, M. Segu, S. Li, W. Abbeloos, and L. Van Gool. Unidepthv2: Universal monocular metric depth estimation made simpler. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025b.
- Radford et al. [2018] A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al. Improving language understanding by generative pre-training. 2018.
- Radford et al. [2019] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
- Radford et al. [2021] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PmLR, 2021.
- Ranzato et al. [2011] M. Ranzato, J. Susskind, V. Mnih, and G. Hinton. On deep generative models with applications to recognition. In CVPR 2011, pages 2857–2864. IEEE, 2011.
- Ravi et al. [2024] N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. Carion, C.-Y. Wu, R. Girshick, P. Dollár, and C. Feichtenhofer. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714, 2024.
- Ren et al. [2024] T. Ren, Y. Chen, Q. Jiang, Z. Zeng, Y. Xiong, W. Liu, Z. Ma, J. Shen, Y. Gao, X. Jiang, et al. Dino-x: A unified vision model for open-world object detection and understanding. arXiv preprint arXiv:2411.14347, 2024.
- Schops et al. [2019] T. Schops, T. Sattler, and M. Pollefeys. Bad slam: Bundle adjusted direct rgb-d slam. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 134–144, 2019.
- Shen et al. [2024] Y. Shen, C. Fu, P. Chen, M. Zhang, K. Li, X. Sun, Y. Wu, S. Lin, and R. Ji. Aligning and prompting everything all at once for universal visual perception. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13193–13203, 2024.
- Silberman et al. [2012] N. Silberman, D. Hoiem, P. Kohli, and R. Fergus. Indoor segmentation and support inference from rgbd images. In European conference on computer vision, pages 746–760. Springer, 2012.
- Siméoni et al. [2025] O. Siméoni, H. V. Vo, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V. Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoa, et al. Dinov3. arXiv preprint arXiv:2508.10104, 2025.
- Tang et al. [2023] L. Tang, M. Jia, Q. Wang, C. P. Phoo, and B. Hariharan. Emergent correspondence from image diffusion. Advances in neural information processing systems, 36:1363–1389, 2023.
- Tschannen et al. [2025] M. Tschannen, A. Gritsenko, X. Wang, M. F. Naeem, I. Alabdulmohsin, N. Parthasarathy, T. Evans, L. Beyer, Y. Xia, B. Mustafa, et al. Siglip 2: Multilingual vision-language encoders with improved semantic understanding, localization, and dense features. arXiv preprint arXiv:2502.14786, 2025.
- Uhrig et al. [2017] J. Uhrig, N. Schneider, L. Schneider, U. Franke, T. Brox, and A. Geiger. Sparsity invariant cnns. In International Conference on 3D Vision (3DV), 2017.
- Vasiljevic et al. [2019] I. Vasiljevic, N. Kolkin, S. Zhang, R. Luo, H. Wang, F. Z. Dai, A. F. Daniele, M. Mostajabi, S. Basart, M. R. Walter, and G. Shakhnarovich. DIODE: A Dense Indoor and Outdoor DEpth Dataset. CoRR, abs/1908.00463, 2019. URL http://arxiv.org/abs/1908.00463.
- Wang et al. [2026a] H. Wang, L. Qiao, Z. Jie, Z. Huang, C. Feng, Q. Zheng, L. Ma, X. Lan, and X. Liang. X-sam: From segment anything to any segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 26187–26196, 2026a.
- Wang et al. [2025a] J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny. Vggt: Visual geometry grounded transformer. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 5294–5306, 2025a.
- Wang et al. [2026b] L. Wang, A. Zanfir, E. G. Bazavan, M. Andriluka, and C. Sminchisescu. Thfm: A unified video foundation model for 4d human perception and beyond. arXiv preprint arXiv:2603.25892, 2026b.
- Wang et al. [2025b] R. Wang, S. Xu, C. Dai, J. Xiang, Y. Deng, X. Tong, and J. Yang. Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025b.
- Wang et al. [2025c] R. Wang, S. Xu, Y. Dong, Y. Deng, J. Xiang, Z. Lv, G. Sun, X. Tong, and J. Yang. Moge-2: Accurate monocular geometry with metric scale and sharp details. arXiv preprint arXiv:2507.02546, 2025c.
- Wang et al. [2023] X. Wang, W. Wang, Y. Cao, C. Shen, and T. Huang. Images speak in images: A generalist painter for in-context visual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6830–6839, 2023.
- Wei et al. [2024] C. Wei, Y. Zhong, H. Tan, Y. Liu, Z. Zhao, J. Hu, and Y. Yang. Hyperseg: Towards universal visual segmentation with large language model. arXiv preprint arXiv:2411.17606, 2024.
- Wei et al. [2021] J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021.
- Wiedemer et al. [2025] T. Wiedemer, Y. Li, P. Vicol, S. S. Gu, N. Matarese, K. Swersky, B. Kim, P. Jaini, and R. Geirhos. Video models are zero-shot learners and reasoners. arXiv preprint arXiv:2509.20328, 2025.
- Wu et al. [2025] C. Wu, J. Li, J. Zhou, J. Lin, K. Gao, K. Yan, S.-m. Yin, S. Bai, X. Xu, Y. Chen, et al. Qwen-image technical report. arXiv preprint arXiv:2508.02324, 2025.
- Xie et al. [2024] J. Xie, W. Mao, Z. Bai, D. J. Zhang, W. Wang, K. Q. Lin, Y. Gu, Z. Chen, Z. Yang, and M. Z. Shou. Show-o: One single transformer to unify multimodal understanding and generation. arXiv preprint arXiv:2408.12528, 2024.
- Xu et al. [2023] J. Xu, S. Liu, A. Vahdat, W. Byeon, X. Wang, and S. De Mello. Open-vocabulary panoptic segmentation with text-to-image diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2955–2966, 2023.
- Yang et al. [2024] L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao. Depth anything v2. Advances in Neural Information Processing Systems, 37:21875–21911, 2024.
- Yang and Wang [2023] X. Yang and X. Wang. Diffusion model as representation learner. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 18938–18949, 2023.
- Ye et al. [2024] C. Ye, L. Qiu, X. Gu, Q. Zuo, Y. Wu, Z. Dong, L. Bo, Y. Xiu, and X. Han. Stablenormal: Reducing diffusion variance for stable and sharp normal. ACM Transactions on Graphics (ToG), 43(6):1–18, 2024.
- Ye et al. [2025] Y. Ye, X. He, Z. Li, B. Lin, S. Yuan, Z. Yan, B. Hou, and L. Yuan. Imgedit: A unified image editing dataset and benchmark. arXiv preprint arXiv:2505.20275, 2025.
- Yu et al. [2024] Q. Yu, P.-T. Jiang, H. Zhang, J. Chen, B. Li, L. Zhang, and H. Lu. High-precision dichotomous image segmentation via probing diffusion capacity. arXiv preprint arXiv:2410.10105, 2024.
- Zhai et al. [2023] X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF international conference on computer vision, pages 11975–11986, 2023.
- Zhang et al. [2025] C. Zhang, G. L. Moing, S. Koppula, I. Rocco, L. Momeni, J. Xie, S. Sun, R. Sukthankar, J. K. Barral, R. Hadsell, et al. Efficiently reconstructing dynamic scenes one d4rt at a time. arXiv preprint arXiv:2512.08924, 2025.
- Zhang et al. [2023a] H. Zhang, F. Li, X. Zou, S. Liu, C. Li, J. Yang, and L. Zhang. A simple framework for open-vocabulary segmentation and detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1020–1031, 2023a.
- Zhang et al. [2023b] J. Zhang, C. Herrmann, J. Hur, L. Polania Cabrera, V. Jampani, D. Sun, and M.-H. Yang. A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence. Advances in Neural Information Processing Systems, 36:45533–45547, 2023b.
- Zhao et al. [2025] C. Zhao, Y. Sun, M. Liu, H. Zheng, M. Zhu, Z. Zhao, H. Chen, T. He, and C. Shen. Diception: A generalist diffusion model for visual perceptual tasks. arXiv preprint arXiv:2502.17157, 2025.
- Zhao et al. [2023] W. Zhao, Y. Rao, Z. Liu, B. Liu, J. Zhou, and J. Lu. Unleashing text-to-image diffusion models for visual perception. In IEEE/CVF International Conference on Computer Vision (ICCV), pages 5729–5739, 2023.
- Zhou et al. [2021] J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, and T. Kong. ibot: Image bert pre-training with online tokenizer. arXiv preprint arXiv:2111.07832, 2021.
- Zou et al. [2023] X. Zou, Z.-Y. Dou, J. Yang, Z. Gan, L. Li, C. Li, X. Dai, H. Behl, J. Wang, L. Yuan, et al. Generalized decoding for pixel, image, and language. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 15116–15127, 2023.
- Zuo et al. [2025] J. Zuo, H. Deng, H. Zhou, J. Zhu, Y. Zhang, Y. Zhang, Y. Yan, K. Huang, W. Chen, Y. Deng, R. Jin, N. Sang, and C. Gao. Is nano banana pro a low-level vision all-rounder? A comprehensive evaluation on 14 tasks and 40 datasets. arXiv preprint, 2025.
Appendix
Appendix A Multi-stage clustering algorithm for parsing instance masks
While the postprocessing of RGB images into binary masks is straightforward for referring expression segmentation and semantic segmentation, the task is more complex for instance segmentation, for which the number of masks to predict is unknown a priori. We therefore instruct the model to dynamically determine the number of masks and assign a unique color to each individual instance. However, because the output is a generated image, it is subject to high-frequency generation noise, slight color drift across a single object mask, and blending artifacts at boundaries. Standard connected-component algorithms or simple color thresholding fail in the presence of these artifacts. Our proposed algorithm acts as a robust post-processing, ensuring that continuous generative outputs are cleanly mapped back to discrete instance masks.
To parse these individual instance masks from the generated multi-colored segmentation images, we design a multi-stage connected component grouping algorithm. For any generated RGB mask image with a predefined background color , the algorithm proceeds as follows:
-
Background Initialization: We first identify and label all background pixels. A pixel is classified as background (and initialized with label ) if its color is close to the predefined background color :
(2) where denotes the RGB vector of pixel , and is the color tolerance threshold. Background pixels are excluded from subsequent steps.
-
Color Similarity Grouping: We group the remaining unlabeled pixels into initial components using a seed-based floodfill algorithm with 16-connectivity. Two adjacent pixels are merged into the same component if their RGB color distance to the component’s seed pixel satisfies:
(3) -
Noise Pruning: To filter out minor noise, any component is discarded (re-labeled as background ) if its total pixel area is smaller than a fraction of the image size:
(4) where (representing 0.02% of the image area).
-
Boundary Artifact Elimination: Generative models often produce thin colored boundary halos between distinct solid-colored objects. To prevent these from forming separate spurious components, we apply a binary erosion using a square structuring element of size . A component is pruned if its preserved area ratio after erosion is below :
(5) -
Spatially Constrained Merging: Spatially disjoint components that could represent parts of the same physical object (e.g., due to foreground occlusion) are merged if their average colors are similar () and if their combined bounding box area does not expand excessively compared to the sum of their individual areas:
(6) where is the maximum bounding box expansion factor.
Appendix B Detailed SA-Co/Gold Quantitative Evaluation
To provide a deeper understanding of our model’s instance segmentation capabilities, we report the per-category performance breakdown on the SA-Co/Gold [Carion et al., 2025] benchmark in table˜8. This benchmark evaluates models on diverse noun-phrase queries across seven distinct subsets: Metaclip, SA-1B, Crowded, Food&Drink, Sports Equip., Attributes, and Wiki-Common.
| Average | Metaclip | SA-1B | Crowded | Food&Drink | Sports Equip. | Attributes | Wiki-Common | |||||||||||||||||
| cgF1 | IL_MCC | pmF1 | cgF1 | IL_MCC | pmF1 | cgF1 | IL_MCC | pmF1 | cgF1 | IL_MCC | pmF1 | cgF1 | IL_MCC | pmF1 | cgF1 | IL_MCC | pmF1 | cgF1 | IL_MCC | pmF1 | cgF1 | IL_MCC | pmF1 | |
| Trained on SA-Co | ||||||||||||||||||||||||
| SAM 3 [Carion et al., 2025] | 54.1 | 0.82 | 66.1 | 47.3 | 0.81 | 58.6 | 53.7 | 0.86 | 62.6 | 61.1 | 0.9 | 67.7 | 53.4 | 0.79 | 67.3 | 65.5 | 0.89 | 73.8 | 54.9 | 0.76 | 72.0 | 42.5 | 0.70 | 60.9 |
| SAM 3 + Llama 3.2 (ft) [Carion et al., 2025] | 61.2 | 0.86 | 70.8 | 54.2 | 0.85 | 64.0 | 56.0 | 0.89 | 62.9 | 61.3 | 0.88 | 69.8 | 67.6 | 0.86 | 78.5 | 67.5 | 0.89 | 75.6 | 71.1 | 0.91 | 77.8 | 51.1 | 0.76 | 67.1 |
| Zero-shot | ||||||||||||||||||||||||
| gDino-T [Liu et al., 2024] | 3.3 | 0.15 | 16.2 | 2.9 | 0.21 | 13.9 | 3.1 | 0.20 | 15.4 | 0.28 | 0.08 | 3.4 | 0.96 | 0.10 | 9.8 | 1.1 | 0.10 | 11.2 | 13.8 | 0.29 | 47.3 | 0.70 | 0.06 | 12.1 |
| OWLv2 [Minderer et al., 2023] | 24.6 | 0.57 | 42.0 | 17.7 | 0.52 | 34.3 | 13.3 | 0.50 | 26.8 | 15.8 | 0.51 | 30.7 | 32.0 | 0.65 | 49.4 | 36.0 | 0.64 | 56.2 | 35.6 | 0.63 | 56.2 | 21.7 | 0.54 | 40.3 |
| LLMDet-L [Fu et al., 2025a] | 6.5 | 0.21 | 27.3 | 4.5 | 0.23 | 19.4 | 5.3 | 0.23 | 22.8 | 2.4 | 0.18 | 13.7 | 5.5 | 0.19 | 29.1 | 4.4 | 0.17 | 25.3 | 22.2 | 0.39 | 57.1 | 1.2 | 0.05 | 23.3 |
| APE-D [Shen et al., 2024] | 16.4 | 0.40 | 36.9 | 12.6 | 0.42 | 30.1 | 2.2 | 0.22 | 10.0 | 7.2 | 0.35 | 20.3 | 22.7 | 0.51 | 45.0 | 31.8 | 0.56 | 56.5 | 26.7 | 0.47 | 57.3 | 11.6 | 0.29 | 39.5 |
| DINO-X [Ren et al., 2024] | 21.3 | 0.38 | 55.2 | 17.2 | 0.35 | 49.2 | 19.7 | 0.48 | 40.9 | 12.9 | 0.34 | 37.5 | 30.1 | 0.49 | 61.7 | 28.4 | 0.41 | 69.4 | 31.0 | 0.42 | 74.0 | 9.7 | 0.18 | 53.5 |
| Gemini 2.5 [Gemini Team, 2025] | 13.0 | 0.29 | 46.1 | 9.9 | 0.29 | 33.8 | 13.1 | 0.41 | 32.1 | 8.2 | 0.27 | 30.3 | 19.6 | 0.33 | 59.5 | 15.1 | 0.28 | 53.5 | 18.8 | 0.30 | 63.1 | 6.5 | 0.13 | 50.3 |
| Vision Banana + Gemini 3.1 Flash-Lite | 47.5 | 0.84 | 56.0 | 38.0 | 0.81 | 46.7 | 33.5 | 0.84 | 39.8 | 34.4 | 0.83 | 41.2 | 57.2 | 0.88 | 64.8 | 63.0 | 0.92 | 68.2 | 58.8 | 0.86 | 68.7 | 47.4 | 0.76 | 62.3 |
We adopt the official metrics established for SA-Co evaluation:
-
Positive Micro- (): The micro-averaged score computed across all positive queries (where the target noun-phrase is present in the image), evaluating the precision and recall of the predicted masks at 10 IoU thresholds in .
-
Image-Level Matthews Correlation Coefficient (): A metric evaluating the binary classification accuracy of the model in predicting whether a noun phrase query is present or absent in the image.
-
Classification-gated (): The primary unified metric, combining and :
(7)
Under the zero-shot transfer setting, Vision Banana (paired with Gemini 3.1 Flash-Lite for positive/negative query filtering) demonstrates state-of-the-art performance on the benchmark, outperforming specialized models such as OWLv2, DINO-X, and APE-D. Arguably, our strong score can be attributed to Gemini’s discriminative capabilities, and prior works would also benefit from being paired with a MLLM. Nonetheless, Vision Banana also achieves state-of-the-art results on the metric, which directly measures open-vocabulary segmentation quality and is the focus of this work.
Appendix C Qualitative Instance Segmentation Analysis
In fig.˜9 and fig.˜10, we present a qualitative comparison of instance segmentation masks predicted by the SAM 3 specialist and our zero-shot Vision Banana model, on examples from the SA-Co/Gold [Carion et al., 2025] benchmark. Below each prediction, we report its best score across the three independent ground-truth annotations, computed as the average of the scores across the ten IoU thresholds defined by the SA-Co benchmark.
We illustrate several successful segmentation results in fig.˜9. For instance, in examples (9.a), (9.b), and (9.c), Vision Banana correctly segments 17 instances of ’manicured and pedicured nail’, 3 instances of ’incandescent lightbulb’, and 1 instance of ’hand-drawn design’, respectively. Despite producing visually accurate segmentation masks, our quantitative scores can remain low, falling significantly below those of SAM 3. We found out that this can be explained by misalignments with the SAM-annotated ground-truth masks, which heavily penalize Vision Banana performance at high IoU thresholds. Specifically, the masks produced by our model are usually slightly smaller (see example (9.a)) and finer (see example (9.c)) than the SAM-produced masks.
Compared to discriminative models like SAM 3, which suffer from ’mode-averaging’ artifacts under spatial ambiguity (see examples (9.b), (9.c) and (10.f)), Vision Banana naturally resolves such ambiguity by committing to a single, coherent mode of the target mask distribution. Furthermore, unlike SAM 3, our model consistently produces non-overlapping masks and does not leave unnatural gaps between adjacent object boundaries (see example (9.f)).
In fig.˜10, we report failure modes of our model. These primarily manifest as: (i) incorrectly merging instances together (e.g., merging wall segments in (10.a), all post boxes in example (10.c), furniture in (10.e), or top and down columns in (10.g)), and (ii) breaking cohesive “group instances” into multiple masks (e.g., segmenting the crowd into individual people in example (10.d) despite the query requesting “the crowd”). Additionally, Vision Banana occasionally misses target objects entirely, such as background walls in example (10.a) or the persons on the left in example (10.d).
This qualitative analysis reveals that our model still has clear room for improvement, particularly in highly cluttered or crowded scenes. Nonetheless, a portion of the quantitative performance gap between Vision Banana and SAM 3 is also attributable to the specific annotation biases of the SA-Co dataset, our model being evaluated zero-shot.
Original Image GT a GT b GT c SAM 3 Predictions Vision Banana
Predictions






(9.a) “pale pink and white manicured and pedicured nail“ F1=0.90 (GT a), 17 masks F1=0.72 (GT c), 17 masks






(9.b) “incandescent lightbulb“ F1=0.77 (GT a), 4 masks F1=0.60 (GT a), 3 masks






(9.c) “hand-drawn design“ F1=0.40 (GT a), 2 masks F1=0.10 (GT b), 1 mask






(9.d) “concrete floor“ F1=0.90 (GT c), 1 mask F1=0.50 (GT b), 1 mask






(9.e) “eggs Benedict“ F1=0.90 (GT a), 1 mask F1=0.00 (GT a), 1 mask






(9.f) “sardine“ F1=0.95 (GT c), 8 masks F1=0.91 (GT a), 8 masks
Original Image GT a GT b GT c SAM 3 Predictions Vision Banana
Predictions






(10.a) “a wall“ F1=0.82 (GT b), 6 masks F1=0.02 (GT c), 2 masks






(10.b) “clothing window display“ F1=0.95 (GT a), 2 masks F1=0.00 (GT a), 3 masks






(10.c) “post office box“ F1=1.00 (GT a), 28 masks F1=0.00 (GT a), 1 mask






(10.d) “the crowd“ F1=1.00 (GT c), 1 mask F1=0.00 (GT a), 24 masks






(10.e) “the furniture“ F1=0.65 (GT c), 18 masks F1=0.00 (GT a), 2 masks






(10.f) “a pot“ F1=0.40 (GT c), 1 mask F1=0.00 (GT a), 1 mask






(10.g) “the column“ F1=0.85 (GT c), 24 masks F1=0.08 (GT c), 12 masks
Appendix D Image generation and editing capabilities
Finally, we assess the generative capabilities of Vision Banana post-instruction-tuning to ensure that fine-tuning the model did not trigger catastrophic forgetting of its core generative features. Figures 11 and 12 compare Vision Banana with its base model, Nano Banana Pro, on text-to-image generation and image-editing tasks.
In Fig. 11, we present generated outputs for prompts sampled from GenAI-Bench [Li et al., 2024a]. These results demonstrate that Vision Banana continues to produce detailed and contextually accurate images with a similar quality to those of the base generative model.
Furthermore, in Fig. 12, we evaluate the two checkpoints on instruction-based image-editing tasks using prompts from ImgEdit [Ye et al., 2025]. Vision Banana shows robust adherence to detailed instructions, confirming that our instruction-tuning successfully teaches output formats for dense vision tasks without degrading the image editing capabilities learned during pre-training.
| Original image | Vision Banana | Nano Banana Pro |
![]() | ![]() | ![]() |
| Change the grassy hills in the picture to a beach with ocean waves. | ||
![]() | ![]() | ![]() |
| Remove the plant from the shelf, and resize the picture frame to be larger. | ||
![]() | ![]() | ![]() |
| Change the vehicle’s color to red. | ||
![]() | ![]() | ![]() |
| Change the background of the suit from a blank wall to a luxurious office setting that includes a wooden desk and a large window showing a cityscape. | ||





































