# MoMo：通过时空动作分词实现机器人操作中的运动模式控制

- 来源：Apple Machine Learning Research（RSS）
- 发布时间：2026-07-30 08:00
- AIHOT 分数：44
- AIHOT 链接：https://aihot.virxact.com/items/cms863zfu009qrot05pg81kk8
- 原文链接：https://machinelearning.apple.com/research/momo-motion-mode-manipulation

## AI 摘要

机器人操作需根据任务、物体和交互环境调整动作执行方式。MoMo 提出两阶段模仿学习框架，包含时空动作分词器和行为克隆 Transformer，将任务与连续运动模式条件作为输入。在六项真实机器人操作任务中，改变该条件可稳定生成不同执行风格的动作。

## 正文

To operate effectively across diverse contexts, robots must not only perform manipulation tasks accurately but also adapt how their actions unfold to the task, object, and interaction setting. We ask whether this execution-level variation can be learned as a reusable behavioral factor shared across tasks. We present MoMo, a two-stage imitation-learning framework consisting of a spatiotemporal action tokenizer and a behavior-cloning transformer that takes task and a continuous motion-mode condition as inputs. Across six real-robot manipulation tasks, varying this condition produces steady…
