# 360 AI 研究院 MoSA：让 AI 通过视频运动学习识别物体，无需人工标注

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-07-15 00:13
- AIHOT 分数：36
- AIHOT 链接：https://aihot.virxact.com/items/cmrkvjeln01s3bi5qolfl48lt
- 原文链接：https://x.com/rohanpaul_ai/status/2077064067102253536

## AI 摘要

360 AI 研究院在 ECCV 2026 发表的 MoSA（Motion-Grounded Segment Anything）论文，提出让 AI 通过视频中物体的运动来学习识别，无需人工标注。该方法从约 10,000 小时无标签视频中自动生成了超过 2,100 万个伪标签，突破了 Meta 的 SAM 模型依赖数百万手工标注图像的限制。MoSA 属于该研究院 H1 2026 系列论文中“精准感知、精准编辑、精准生成”研究方向的一部分，旨在让 AI 的感知、编辑和生成能力更准确、可控且实用。

## 正文

AI should not just get bigger. It should also be more precise and easier to manage.

What I see in the 360 AI Research Institute's H1 2026 papers， is： see precisely， edit precisely， generate precisely.

MoSA （Motion-Grounded Segment Anything）， accepted by ECCV 2026， sits under the "See Precisely" direction.

The premise is simple， but crucial： can video teach AI to understand things without people hand-labeling every thing？

MoSA produced more than 21 million pseudo-labels from approximately 10，000 hours of unlabelled video.

Segment Anything Model is trained on notes. MoSA tests the notion that things could learn from their movement.

This is significant because normal video could be used to train AI that can see things on a large scale.

That is why MoSA is not just a paper on segmentation. This is part of a larger research goal of making AI perception， editing， and generation more accurate， controllable and useful in the real world.

### 引用推文

> Shruti：BREAKING: Chinese cybersecurity company just found a way around one of AI's most expensive problems. Meta built an AI model called SAM that could outline any ob...
