# 商汤开源 SenseNova-Vision-7B-MoT：单一模型覆盖所有主要视觉任务

- 来源：SenseTime (@SenseTime_AI)
- 发布时间：2026-07-13 22:41
- AIHOT 分数：55
- AIHOT 链接：https://aihot.virxact.com/items/cmrjcreil01nnbi03xikd0jds
- 原文链接：https://x.com/SenseTime_AI/status/2076678396881649805

## AI 摘要

商汤开源 SenseNova-Vision-7B-MoT，该模型将计算机视觉重构为多模态生成，无需任务专用头即可达到 SOTA，在分割与密集几何任务上超越 Google DeepMind 的 Vision Banana。具体对比：Referring segmentation（RefCOCOg cIoU）80.3 vs 73.8，Semantic segmentation（Cityscapes mIoU）71.2 vs 69.9，Depth estimation（NYUv2 δ1）98.1 vs 94.8，Surface normal（NYUv2 mean error）14.4 vs 17.8。模型支持语言或视觉提示引导，基于商汤十年 CV 积累。新任务由指令定义，而非新模型头。

## 正文

Open-sourcing SenseNova-Vision-7B-MoT - one model， all major vision tasks.

Reimagines CV as multimodal generation. Steerable by language or visual prompts. No task-specific heads - still SOTA， surpassing Google DeepMind's Vision Banana on segmentation and dense geometry.

📊 vs. Vision Banana （Google DeepMind）：
🔹Referring segmentation （RefCOCOg cIoU）： 80.3 vs 73.8
🔹Semantic segmentation （Cityscapes mIoU）： 71.2 vs 69.9
🔹Depth estimation （NYUv2 δ1）： 98.1 vs 94.8
🔹Surface normal （NYUv2 mean error）： 14.4 vs 17.8

Built on a decade of SenseTime's CV leadership - #1 in China's Vision AI market for 10 consecutive years， and named a Tech Innovator in global GenAI CV by Gartner. Vision， now native to foundation models. New tasks， defined by instruction - not by new model heads.
