内容
精选全部 AI 动态AI 日报主题收藏
接入
Agent 接入
更多
关于更新日志反馈
原文
Rohan Paul@rohanpaul_ai
52
2026-07-14 07:18· 7天前
跳到正文
AI 摘要

蚂蚁集团旗下具身AI公司Robbyant发布LingBot-VA 2.0,这是一款从零训练的视频动作基础模型,专为机器人控制设计。模型采用稀疏MoE架构,约13B总参数,每token激活约1.9B参数。通过前瞻推理及系统加速,峰值异步执行频率达225Hz。高层视觉语言规划器分解长任务,底层视频动作策略处理连续运动。模型仅需10-15次演示即可适应,支持跨机器人本体迁移及零样本新任务。在RoboTwin 2.0基准上平均得分93.6,真实世界测试优于LingBot-VA和π0.5。

Most video-action robot models are a content-creation video generator with an action module attached.

LingBot-VA 2.0 from @robbyant_brain, a video-action foundation model, throws that starting point out and trains the whole stack natively for control.

And it runs closed-loop at a peak 225 Hz.

It's so important because A robot cannot move responsively when its controller pauses to imagine the next few frames. LingBot-VA 2.0 predicts during execution, then corrects using each real observation.

And it carries only about 13B video parameters while activating roughly 1.9B per token. Bigger robot models usually mean slower reactions, creating a direct conflict between intelligence and control.

LingBot-VA 2.0 is trained from scratch for robot control rather than adapted from a video generator built for content creation.

Robbyant, an embodied AI company under Ant Group, built it to learn how scenes change under actions, predict what should happen next, and turn those predictions into real-time robot movements.

Most video-action systems inherit a tokenizer and video backbone trained mainly to reproduce visual appearance. LingBot-VA 2.0 rebuilds both parts around physical control.

Its semantic visual-action tokenizer maps observations toward features from a frozen vision foundation model and learns compact latent actions from frame-to-frame changes using self-supervised inverse and forward dynamics. Unlabeled web video can therefore carry action-relevant training signals without robot action labels.

The policy is causal from the start, so every prediction can use only past observations.

Rohan Paul@rohanpaul_ai · X
52导出 Markdown
2026-07-14 07:18·7天前
在 X 看原推· x.com
AI 摘要

蚂蚁集团旗下具身AI公司Robbyant发布LingBot-VA 2.0,这是一款从零训练的视频动作基础模型,专为机器人控制设计。模型采用稀疏MoE架构,约13B总参数,每token激活约1.9B参数。通过前瞻推理及系统加速,峰值异步执行频率达225Hz。高层视觉语言规划器分解长任务,底层视频动作策略处理连续运动。模型仅需10-15次演示即可适应,支持跨机器人本体迁移及零样本新任务。在RoboTwin 2.0基准上平均得分93.6,真实世界测试优于LingBot-VA和π0.5。

Most video-action robot models are a content-creation video generator with an action module attached.

LingBot-VA 2.0 from @robbyant_brain, a video-action foundation model, throws that starting point out and trains the whole stack natively for control.

And it runs closed-loop at a peak 225 Hz.

It's so important because A robot cannot move responsively when its controller pauses to imagine the next few frames. LingBot-VA 2.0 predicts during execution, then corrects using each real observation.

Its sparse Mixture-of-Experts video backbone has about 13B total parameters, while about 1.9B are active per token, keeping the compute lower during each step. A high-level vision-language planner breaks long tasks into smaller instructions, while the low-level video-action policy handles continuous movement.

Foresight Reasoning predicts future visual states while the robot is already acting, then replaces imagined states with every new real observation. Combined with few-step distillation and systems acceleration, the paper reports a peak asynchronous execution frequency of 225 Hz.

The model adapts from 10-15 demonstrations, transfers across robot embodiments, and handles some new tasks zero-shot.

In the paper's own evaluations, it reaches 93.6 average on RoboTwin 2.0 and reports stronger real-world results than LingBot-VA and π0.5 across the tested tasks.

🧵 1.

具身智能模型发布
在 X 查看原推导出 Markdown
同一事件 · 1 家报道
  • 7月11日精选蚂蚁集团 Robbyant 发布 LingBot-VA 2.0,首个原生具身基础模型MarkTechPost(RSS)

And it carries only about 13B video parameters while activating roughly 1.9B per token. Bigger robot models usually mean slower reactions, creating a direct conflict between intelligence and control.

LingBot-VA 2.0 is trained from scratch for robot control rather than adapted from a video generator built for content creation.

Robbyant, an embodied AI company under Ant Group, built it to learn how scenes change under actions, predict what should happen next, and turn those predictions into real-time robot movements.

Most video-action systems inherit a tokenizer and video backbone trained mainly to reproduce visual appearance. LingBot-VA 2.0 rebuilds both parts around physical control.

Its semantic visual-action tokenizer maps observations toward features from a frozen vision foundation model and learns compact latent actions from frame-to-frame changes using self-supervised inverse and forward dynamics. Unlabeled web video can therefore carry action-relevant training signals without robot action labels.

The policy is causal from the start, so every prediction can use only past observations.

Its sparse Mixture-of-Experts video backbone has about 13B total parameters, while about 1.9B are active per token, keeping the compute lower during each step. A high-level vision-language planner breaks long tasks into smaller instructions, while the low-level video-action policy handles continuous movement.

Foresight Reasoning predicts future visual states while the robot is already acting, then replaces imagined states with every new real observation. Combined with few-step distillation and systems acceleration, the paper reports a peak asynchronous execution frequency of 225 Hz.

The model adapts from 10-15 demonstrations, transfers across robot embodiments, and handles some new tasks zero-shot.

In the paper's own evaluations, it reaches 93.6 average on RoboTwin 2.0 and reports stronger real-world results than LingBot-VA and π0.5 across the tested tasks.

🧵 1.

具身智能模型发布
在 X 查看原推x.com
同一事件 · 1 家报道点击查看
  • 7月11日精选蚂蚁集团 Robbyant 发布 LingBot-VA 2.0,首个原生具身基础模型MarkTechPost(RSS)