AI should not just get bigger. It should also be more precise and easier to manage.
What I see in the 360 AI Research Institute's H1 2026 papers, is: see precisely, edit precisely, generate precisely.
MoSA (Motion-Grounded Segment Anything), accepted by ECCV 2026, sits under the "See Precisely" direction.
The premise is simple, but crucial: can video teach AI to understand things without people hand-labeling every thing?
MoSA produced more than 21 million pseudo-labels from approximately 10,000 hours of unlabelled video.
Segment Anything Model is trained on notes. MoSA tests the notion that things could learn from their movement.
This is significant because normal video could be used to train AI that can see things on a large scale.
That is why MoSA is not just a paper on segmentation. This is part of a larger research goal of making AI perception, editing, and generation more accurate, controllable and useful in the real world.