Open-sourcing SenseNova-Vision-7B-MoT - one model, all major vision tasks.
Reimagines CV as multimodal generation. Steerable by language or visual prompts. No task-specific heads - still SOTA, surpassing Google DeepMind's Vision Banana on segmentation and dense geometry.
📊 vs. Vision Banana (Google DeepMind): 🔹Referring segmentation (RefCOCOg cIoU): 80.3 vs 73.8 🔹Semantic segmentation (Cityscapes mIoU): 71.2 vs 69.9 🔹Depth estimation (NYUv2 δ1): 98.1 vs 94.8 🔹Surface normal (NYUv2 mean error): 14.4 vs 17.8
Built on a decade of SenseTime's CV leadership - #1 in China's Vision AI market for 10 consecutive years, and named a Tech Innovator in global GenAI CV by Gartner. Vision, now native to foundation models. New tasks, defined by instruction - not by new model heads.