HuggingFace Daily Papers(社区热门论文)
52AI 编辑部评分,满分 100

DyPES-VLA:学习共享动力学先验与具身特定控制,实现跨具身操作

2026-08-06 08:00· 1天前
AI 导读

DyPES-VLA 提出一种跨具身视觉-语言-动作模型,通过未来预测目标训练 VLM 学习共享动力学先验,并利用具身特定的 MoE 动作头直接在原生动作空间生成控制,无需手动对齐异构动作。作为通用策略,它在 LIBERO 上达到 98.0% 成功率,在 RoboCasa-GR1 上为 59.25%,在 RoboTwin 2.0 上为 89.02%,在仿真和真实世界评估中均取得 SOTA 性能。

Abstract:Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse visual and interaction data, limiting cross-embodiment transfer. Second, they require extensive manual preprocessing to convert embodiment-specific actions into a common format. To overcome these limitations, we propose DyPES-VLA, a cross-embodiment VLA that learns shared Dynamics Priors and Embodiment-Specific control. First, we learn shared dynamics priors by training the vision-language model (VLM) with a future-prediction objective on cross-embodiment data, driving the shared query representation to capture object motion, contact, and interaction-induced scene changes. Second, an embodiment-specific Mixture-of-Experts (MoE) action head translates these shared dynamics priors into executable controls directly in each embodiment's native action space, without manually pre-aligning heterogeneous actions into a common format. This head shares attention layers to capture common temporal action structures, while its embodiment-specific feed-forward experts resolve the unique kinematic constraints and control semantics of distinct embodiments. As a generalist policy, our \ourmethod achieves state-of-the-art performance across simulation and real-world evaluations, reaching 98.0% success on LIBERO, 59.25% on RoboCasa-GR1, and 89.02% on RoboTwin~2.0.
Subjects: Robotics (cs.RO)
Cite as: arXiv:2608.06374 [cs.RO]
  (or arXiv:2608.06374v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2608.06374
arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zhide Zhong [

Thu, 6 Aug 2026 17:59:20 UTC (17,407 KB)

Access Paper:

Current browse context:

References & Citations

Bookmark

BibSonomy Reddit

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org

DyPES-VLA:学习共享动力学先验与具身特定控制,实现跨具身操作

HuggingFace Daily Papers(社区热门论文)·2026-08-06 08:00·1天前
AI 导读

DyPES-VLA 提出一种跨具身视觉-语言-动作模型,通过未来预测目标训练 VLM 学习共享动力学先验,并利用具身特定的 MoE 动作头直接在原生动作空间生成控制,无需手动对齐异构动作。作为通用策略,它在 LIBERO 上达到 98.0% 成功率,在 RoboCasa-GR1 上为 59.25%,在 RoboTwin 2.0 上为 89.02%,在仿真和真实世界评估中均取得 SOTA 性能。

原文 · 保持原样,未翻译
Abstract:Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse visual and interaction data, limiting cross-embodiment transfer. Second, they require extensive manual preprocessing to convert embodiment-specific actions into a common format. To overcome these limitations, we propose DyPES-VLA, a cross-embodiment VLA that learns shared Dynamics Priors and Embodiment-Specific control. First, we learn shared dynamics priors by training the vision-language model (VLM) with a future-prediction objective on cross-embodiment data, driving the shared query representation to capture object motion, contact, and interaction-induced scene changes. Second, an embodiment-specific Mixture-of-Experts (MoE) action head translates these shared dynamics priors into executable controls directly in each embodiment's native action space, without manually pre-aligning heterogeneous actions into a common format. This head shares attention layers to capture common temporal action structures, while its embodiment-specific feed-forward experts resolve the unique kinematic constraints and control semantics of distinct embodiments. As a generalist policy, our \ourmethod achieves state-of-the-art performance across simulation and real-world evaluations, reaching 98.0% success on LIBERO, 59.25% on RoboCasa-GR1, and 89.02% on RoboTwin~2.0.
Subjects: Robotics (cs.RO)
Cite as: arXiv:2608.06374 [cs.RO]
  (or arXiv:2608.06374v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2608.06374
arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zhide Zhong [

Thu, 6 Aug 2026 17:59:20 UTC (17,407 KB)

Access Paper:

Current browse context:

References & Citations

Bookmark

BibSonomy Reddit

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org