Topic · 主题全部主题 →

具身智能

AI 走进物理世界:人形机器人、具身基础模型与真实环境操作能力的进展追踪。

2,340条收录
119条精选

精选归档 · 第 4 页

6180 条 · 共 119

6月16日

星期二 · 1 条
12:39
Qwen:Blog Retrieval(API)精选
AI 评分 73/100
Qwen-Robot Suite:面向物理世界智能的基础模型套件

Qwen 发布三款基础模型——Qwen-RobotNav、Qwen-RobotManip 和 Qwen-RobotWorld。Nav 通过可控观测协议统一指令跟随、点/物体目标导航、目标追踪和自动驾驶五类任务,在 VLN-CE RxR 上达 76.5% SR,HM3Dv2 物体目标导航(仅 RGB)75.6% SR,EVT-Bench 追踪率 90.0%,NAVSIM 91.4 PDMS。Manip 利用规范状态-动作空间对超 38,100 小时异构开源机器人数据进行跨本体训练。World 通过自然语言动作接口协同训练 20 余种本体,预测操控、驾驶和导航的物理未来。三者共同将通用智能转化为物理行动。

另有 5 家信源报道Qwen:Blog Retrieval(API)公众号:通义实验室(千问)Hacker News 热门(buzzing.cc 中文翻译)X:通义千问 / Qwen (@Alibaba_Qwen)MarkTechPost(RSS)
推荐理由:Qwen把导航、操作和世界模型打包成机器人基础套件,用统一语言接口串起来,这是具身智能从单点突破走向系统化的信号,虽然离真实世界还有距离,但方向很清晰。

6月15日

星期一 · 1 条
22:23
The Verge:AI(RSS)精选
AI 评分 70/100
Skydio CEO Adam Bry:硅谷不应为无人机使用画红线

Skydio是美国最大的无人机制造商,主攻公共安全、军事、能源、基建巡检等企业市场。CEO Adam Bry表示,特朗普政府去年底禁止中国产无人机后,廉价消费级无人机几乎消失,Skydio产品成为主要替代方案。公司认为无人机正从工具转向自主基础设施——通过机库、远程操控和软件整合实现规模化应用,AI在其中扮演关键角色。访谈还涉及Skydio与军方合作的态度,以及自主技术如何带动公司扩张。


推荐理由:Adam Bry 的立场很鲜明,硅谷不该替前线士兵做决定。这是军工 AI 伦理争议中的一个不避讳声音,做相关产品的人值得听。

6月12日

星期五 · 3 条
11:00
HuggingFace Daily Papers(社区热门论文)精选
AI 评分 75/100
WEAVER:一种更优、更快、更长的机器人操作世界模型

WEAVER是一种多视图世界模型架构,通过流匹配损失训练预测未来潜变量和奖励值,满足保真度、一致性和效率三个要求。在机器人操作任务上,WEAVER在政策评估中与真实成功率的相关系数ρ=0.870,在π₀.₅基础模型基础上实现政策改进成功率提升38%,测试时规划成功率提升14%,且速度比先前世界模型快5–10倍。在分布外场景下表现也优于先前世界模型。代码、模型和视频已开源。


推荐理由:世界模型在机器人操控上第一次同时跑通了「高保真、长时一致、高推理效率」这三个硬指标,真机实验把成功率拉高38%,代码模型全开源,搞具身智能的值得认真读。
05:29
Rohan Paul@rohanpaul_ai精选
AI 评分 83/100
Jeff Bezos 在 CNBC 披露 Prometheus 愿景:构建人工通用工程师,融资 120 亿美元估值 410 亿美元Jeff Bezos on CNBC explains revealed what Prometheus is building.Today his new company Prometheus announced a $12B funding round at a valuation of $41B .Prometheus trying to build an artificial general engineer that can help design and manufacture physical products like engines, medical devices, and electronics.So the target areas are hard physical products like jet engines, chips, bridges, medical devices, consumer electronics, aerospace systems, vehicles, and drug design, where design cycles can take years because every idea has to survive physics, materials, cost, testing, and factory limits.Bezos’ jet-engine example explains it well: asking for the same engine with 10% more thrust can become a 10-year engineering program, and Prometheus wants to shrink that “dream-build” cycle by 10x or more.The $6.2B launch funding gave Prometheus a massive starting base, and the new raise says the company likely needs far more compute, talent, and industrial data before it can prove the product.Their $41B valuation shows that frontier AI is becoming less a software race than a compute procurement race.A company with no broadly shipped product can raise $12 billion at a $41 billion valuation because investors are not only funding a model, they are prepaying for the machines that might make the model possible.The scarce asset is no longer just talent or algorithms, but clustered GPUs, power contracts, cooling, networking, and the operational skill to keep expensive silicon busy.They are proof that demand is arriving faster than infrastructure can be built, and that every frontier funding round quietly turns into a future claim on power, racks, GPUs, and uptime.Jeff Bezos 在 CNBC 披露其新公司 Prometheus 的愿景:构建人工通用工程师,设计制造喷气发动机、芯片、医疗设备等硬物理产品,将传统数年设计周期缩短 10 倍以上。公司宣布完成 120 亿美元融资,估值 410 亿美元。初始启动资金 62 亿美元,新一轮融资表明公司需要更多算力、人才和工业数据才能验证产品。410 亿美元估值表明,前沿 AI 已从软件竞赛变为计算采购竞赛--投资者实质在为可能实现模型所需的机器预付费。
另有 2 家信源报道TechCrunch:AI(RSS)X:Kim (@kimmonismus)
推荐理由:这不是又一家AI初创,而是直接宣告算力即护城河的开端。Bezos的12B融资对创业者和投资人都是一本摊开的说明书,得读。

6月10日

星期三 · 1 条
17:42
Huawei Cloud@HuaweiCloud1精选
AI 评分 69/100
华为云发布全球首个端到端具身AI平台CloudRoboUnveiling the world’s first end-to-end embodied AI development platform!Huawei Cloud CloudRobo streamlines the entire embodied AI development lifecycle—from data and models to deployment and integration—backed by a secure, trusted PB-scale data foundation.At #INSPIRE2026, the National and Local Co-built Humanoid Robotics Innovation Center, Yijiahe Technology, and Shanghai Jiao Tong University highlighted CloudRobo’s core strengths: • Dual evaluation system for data & models • Fast assembly of active force control models • Robots connected to cloud in hours, models deployed in minutesReady to shape the embodied AI future? Join the Embodied AI Zone today.👉 Learn more: https://tinyurl.com/3zda3sr3 #HuaweiCloud #CloudRobo #EmbodiedAI #Robotics华为云推出全球首个端到端具身AI开发平台CloudRobo,覆盖从数据、模型到部署、集成的全生命周期,基于PB级可信数据底座。在INSPIRE2026上,国家地方共建人形机器人创新中心、Yijiahe Technology、上海交通大学展示了其核心能力:数据与模型双评估系统、主动力控模型快速组装、机器人小时级上云、模型分钟级部署。

推荐理由:具身智能开发链条太长,华为云这个平台把数据、模型、部署打通了,对机器人创业团队来说可能是个加速器,但实际效果还得看落地案例。

6月9日

星期二 · 3 条
21:00
公众号:火山引擎精选
AI 评分 69/100
全新汽车品牌AIVA发布,火山引擎助力打造AI汽车新体验

由赛力斯、宁德时代等多方产业资本组建的AI出行品牌AIVA正式发布。火山引擎提供豆包大模型、智能座舱等技术服务。概念车AIVA Origin Concept亮相,首款量产车AIVA ME7将于2026年内亮相,全系覆盖20万元以上市场。AIVA提出“AI定义汽车”路径,让汽车成为具身AI生命体。火山引擎副总裁表示,人与汽车的关系将实现交互、智能、感受三方面根本转变。未来双方将围绕AI交互、智能体验、情感陪伴深度共创。


推荐理由:AIVA把「先有AI再有车」当作造车逻辑,火山引擎直接下场定义汽车AI体验,这是豆包大模型从软件跑到物理世界的第一次大规模试水,做具身智能和车载产品的人该仔细看看。
09:21
IT之家(RSS)精选
AI 评分 70/100
两部门:到2026年底人形机器人等重点产品完成应用验证并常态部署

工信部、国资委6月8日联合发布通知,目标到2026年底,人形机器人等重点产品在代表性场景完成应用验证并开启常态部署,形成百个以上高价值场景,万台级规模落地。要求各省级地区选取不少于20个场景单元(覆盖两类领域),央企不少于10个。围绕打造实景实训空间、组建创新应用联合体、攻关作业技能、加强验证部署、强化要素保障、凝练经验等六大任务展开,鼓励“人形机器人即服务”等商业创新。


推荐理由:工信部和国资委联合发文,目标2026年底人形机器人万台规模落地,这不是画饼,是实打实的场景清单和验证要求,做机器人的同行该逐条对照了。
08:00
HuggingFace Daily Papers(社区热门论文)精选
AI 评分 78/100
Embodied-R1.5:通过具身基础模型演化物理智能

Embodied-R1.5是一个统一具身基础模型,将具身认知、任务规划、纠错与指向能力整合在单一架构中。基于三条自动化数据构建流水线,团队搭建超过150亿模型token的数据系统,并设计多任务平衡强化学习方案以缓解异构任务冲突。其Planner-Grounder-Corrector闭环框架使模型能在长周期任务中自主执行并自我纠正。仅8B参数的Embodied-R1.5在24个具身VLM基准中的16个上达到SOTA,超越Gemini-Robotics-ER-1.5与GPT-5.4,并可微调为VLA,在4个操作任务基准上领先π_{0.5}等模型。零样本真实机器人实验验证了其指令遵循、可操作物体判别、铰接物体操控与长周期复杂任务中的泛化能力。模型权重、数据集、训练代码及评估框架EmbodiedEvalKit已开源。


推荐理由:仅8B参数就在24项具身视觉语言基准上赢过GPT-5.4和Gemini-Robotics,还把模型权重、训练代码全开源了,做具身智能的团队不跟进就是犯罪。

6月8日

星期一 · 1 条
14:20
IT之家(RSS)精选
AI 评分 73/100
全球首个:高德发布3D原生城市世界模型ABot-Earth0.5

阿里巴巴旗下高德发布全球首个3D原生城市世界模型ABot-Earth0.5,已建成覆盖190多个国家和地区的3D地图。用户输入卫星图或文字描述,10分钟即可在消费级GPU上生成公里级3D城市,输出可编辑3DGS格式,可直接导入Unity等引擎。制图成本为传统百分之一,效率提升约千倍,可为具身智能、低空经济、应急救援等提供支撑。目前已开放内测,可前往abot-earth.amap.com提交申请。


推荐理由:第一个把分钟级 3D 城市重建拉进消费级 GPU 的世界模型,成本打到了传统方案的百分之一,对具身智能和低空经济是底层能力补全,值得内测试试。

6月5日

星期五 · 1 条
22:15
IT之家(RSS)精选
AI 评分 72/100
开源鸿蒙 OpenHarmony 具身智能版本 EmbodiedAI 1.0.1 发布

6月5日,开源鸿蒙具身智能PMC(筹)发布EmbodiedAI 1.0.1版本。该版本聚焦机器人控制与智能体应用,升级导航规划、运动控制、仿真开发、硬件适配等核心能力,兼容ROS生态、机器人模拟器及多种本体形态。集成开源鸿蒙原生模拟器、MuJoCo、Gazebo三大仿真环境,打通从代码开发到真机验证的全流程链路。人形机器人、四足机器狗、商用服务机器人等已完成适配验证。目前具身智能方向已组建18个专项SIG工作组,版本源码已正式开放。


推荐理由:开源鸿蒙的具身智能框架终于从概念走向工程交付,EmbodiedAI 1.0.1 打通仿真和真机,对于不想被ROS绑架的机器人团队是个新选择。

6月4日

星期四 · 1 条
03:20
Fei-Fei Li@drfeifei精选
AI 评分 78/100
世界模型的功能分类http://x.com/i/article/2062244283940544512A Functional Taxonomy of World Models“The world is everything that is the case.” — Ludwig Wittgenstein, Tractatus Logico-Philosophicus, 1921The world is not made of words.In an earlier essay, we argued that spatial intelligence is AI’s next frontier and that world models are the path to it. Here, the World Labs team and I want to go one level deeper: of the many things now being built and called ‘world models,’ which functional pieces actually compose that capacity — and what is each one for?Language models have given machines an extraordinary command of concepts, vocabulary, and reasoning, but the physical world, virtual or real, runs on a different substrate. Where language models learn the statistical structure of text, world models learn the statistical structure of space and time: how light falls on a surface, how a garden looks from an angle no camera has captured, how objects respond to force and follow the laws of physics.That makes “world model” one of the most important and most overloaded terms in AI today. Computer vision, robotics, reinforcement learning, and generative AI each claim to be building world models, and each means something quite different. A video model that produces gorgeous but physically impossible flames, a language model improvising a playable game, and a physics engine that faithfully simulates combustion all go by the same name.The ancient Greeks could never agree on what the world was made of, whether fire, water, or indivisible atoms, because “world” was never a single thing. It was always a stand-in for whatever totality a given thinker needed to reason about. AI has inherited the same problem, at exactly the moment when the field needs precision.The loop beneath the taxonomyCutting through that confusion starts with a diagram older than any of the technology in question. Reinforcement learning textbooks, including the canonical Sutton and Barto, have used a version of the same picture for decades to describe how an agent interacts with a world. The formal name for this picture is the partially observable Markov decision process, or POMDP, and the original definition of the term “world model” belongs to that tradition.An agent, which can be a person, a robot, or a software system, takes actions. Those actions affect the state of the world. The agent never sees the state directly. What reaches the agent are observations: the photons that fall on a retina, the readings from a sensor, and the pixels in a video frame. New observations inform new actions, and the loop continues.The word “state” needs unpacking, because the meaning shifts from field to field. This is not the chemist’s state, the difference between solid, liquid, and gas. This is the physicist’s and roboticist’s state: a complete description of what is happening in the world at a given moment, including every object, every position, every velocity, every property. State is the underlying reality of the world; complete in principle, but never directly visible to any agent inside it. Observations are an agent’s partial view of that reality. Actions are what the agent does in response.This loop — agent to action to state to observation and back — is the structure that gave the modern term “world model” its technical meaning. The phrase itself is older, traced to Kenneth Craik’s 1943 proposal that minds reason by running “small-scale models” of reality, and carried into neural networks by the late 1980s and early 1990s. And the loop also explains what people mean by the term today. The different things now being called world models are in fact different projections of this same loop. Each one outputs a different piece of it.Three functions of a world modelThe first kind of world model is a renderer. A renderer outputs observations in the form of pixels meant for human eyes, and the quality that matters most is visual fidelity. A video model that turns a text prompt into a cinematic drone shot is a renderer. So is an interactive system like Google’s Genie 3, or World Labs’ own RTFM, where the model generates frames in real time conditioned on user input. The model carries no explicit understanding of three-dimensional structure. It produces what a viewer would see, not what is. The buildings in the drone shot may look flawless from above, but try to drive through the city below and they fall apart.The second kind is a simulator. A simulator outputs state: a geometrically, physically or dynamically faithful representation of the world that humans and computer programs can both compute on and interact with. Where the renderer’s contract is purely visual, the simulator’s contract is structural, demanding geometry that holds up under inspection, physics that respects Newton’s laws, and dynamics that behave the way the world needs to behave given the laws of physics. A simulator serves two consumers at once. Human professionals such as architects, designers, filmmakers, and game developers need accuracy beyond visual plausibility. Computer programs such as reinforcement learning agents, robot controllers, and autonomous vehicles use simulators as training grounds where they can interact with the world at scale, testing scenarios that would be dangerous, expensive, or impossible to run in reality.The third kind is a planner. A planner outputs actions. Given an observation and a goal, a planner answers the question of what the agent should do next. This is, in many ways, the inverse of the renderer. Where a renderer takes actions as input and produces observations, a planner takes observations as input and produces actions, closing the perception-action loop. Vision-Language-Action models, model-based systems, and the new wave of World Action Models are all attempts at planners: systems that can decide what a robot should do in an unstructured world.These three categories describe most of what is actually shipping today, and the distinction between them is useful in practice. The categories are not, however, fundamentally separate. The same underlying knowledge of how the world works—geometry, physics, dynamics—sits beneath all of them. A model that can render a cup from any angle ought, in principle, to be able to simulate what happens when the cup is pushed and plan a hand to pick the cup up. Increasingly, the most interesting research deliberately blurs the boundaries between the three.Why simulation is the linchpinOf the three categories, the simulator gets the least public attention, and is the most consequential of the three. This essay addresses this asymmetry.The renderer is by far the most commercially mature. A number of image- or text-to-video products are expanding in the consumer or enterprise markets rapidly. Google’s Nano Banana model has put renderer-quality image generation in the hands of potentially hundreds of millions of users. The technology is real, and the markets are real. Yet renderers optimize for visual plausibility rather than physical accuracy, and that ceiling matters. Their outputs are beautiful, but they cannot be trusted to design a building or train a robot.The planner is the most intriguing and the most nascent, closely connected to the rapidly evolving field of robotic learning. The field has produced robotic demos in the last two years that look impressive in videos, but candor is required about what those demos actually show. Almost all have been confined to heavily constrained laboratory setups, with narrow object sets and short task horizons. None have been validated at the complexity, variability, or duration that real-world deployment demands. The gap between a compelling demo reel and a robot that reliably works in a kitchen, a warehouse, or an operating room remains vast. The commercial bets are nonetheless substantial. A wave of well-funded entrants is racing to ship general-purpose planning systems, while the largest infrastructure players are positioning planning atop broader simulation stacks. A robot that can plan is a robot that can work, and the entire industry is racing to be the one that gets there first.Simulation is the bridge between the two. If language is an abstraction of the world and pixels are a projection of it, then geometry, physics, and dynamics are the world itself. A simulator must work at that level: the structural backbone from which both visual appearance (for renderers) and action consequences (for planners) can be derived.A model that masters simulation can project its understanding into pixels for human consumption, and into action predictions for embodied agents. A model that masters only rendering, or only planning, cannot do either. The commercial surface area is enormous. NVIDIA’s Omniverse alone targets what the company estimates as more than a trillion dollars of addressable market in factories, warehouses, supply chains, and digital twins. Robotics training, autonomous vehicle testing, architectural visualization, engineering, and drug discovery all depend on something simulation-shaped.The hardest open problems in the field live there too. Three-dimensional data with explicit geometry, material properties, and physical annotations is orders of magnitude scarcer than the internet video that renderers train on. The sim-to-real gap, which is the difference between how things behave in simulation and how they behave in reality, persists. Generative simulators introduce a new risk on top of that: AI-generated geometry can look correct while containing self-intersections or wrong scale that produce nonsensical physics. Multi-physics simulation at scale, where rigid bodies, deformable objects, fluids, and cloth all interact, remains orders of magnitude more expensive than single-domain simulation.At World Labs, Marble is our first move into this territory. It takes multimodal prompts (text, image, video, or spatial sketch) and generates explorable 3D environments, outputting Gaussian splats for visual exploration alongside collision meshes a physics engine can operate on. But Marble is only the first chapter of a much longer arc being written across the field as the lines between rendering, simulation, and planning begin to collapse.Where the boundaries are collapsing and what comes nextBut more is to come. The most important pattern in the field right now is that the three categories are starting to blend into one another. The shared insight is that the knowledge required to render a world, simulate it, and act in it is largely the same. Continuing the earlier example, a model that truly understands how a cup sits on a table (its geometry, material properties, response to force, etc.) should be able to render that cup from any angle, simulate what happens when the cup is pushed, and plan for a hand to pick the cup up. The three categories are three projections of a single underlying understanding.For example: a small but growing number of recent work from various robotics labs have demonstrated that—at least conceptually—a pretrained video renderer can be used as the backbone for joint world-and-action prediction, suggesting a bridge between the renderer and the planner by letting one model imagine what will happen and what to do. World Labs’ Marble already outputs Gaussian splats and collision meshes from a single model, dissolving the boundary between the renderer and the simulator. Every level is moving from passive output to interactive system, with renderers becoming action-conditioned, simulators generating worlds that are more controllable and editable, and planners deliberating rather than just reacting.The logical endpoint is a unified world model: one foundation model that can render photorealistic views, produce physically accurate structure, and plan action sequences, switching between output modalities depending on what the downstream consumer needs. We will still face a number of daunting challenges. The data picture is uneven, with renderers awash in internet video while simulators and planners face acute shortages of 3D assets and robot demonstrations. Optimizing for visual beauty can sacrifice the precision a robot or a high-fidelity simulation needs. Reconciling these tensions inside a single architecture is the defining open problem in world model research today, and this is what World Labs sets out to do as we continue to evolve Marble.The direction, however, is clear. The same bet the field has been making since the late 1980s — that a sufficiently rich model of the world is all that any agent needs to see worlds, build them, and act in them — is the bet now driving an entire generation of research. What gives that “big bet” weight is the convergence already underway: three threads, each already driving and shaping multi-billion-dollar industries on its own, that began as separate research programs are starting to behave like one. Taken together, as the boundaries between them collapse, they will reshape something larger: the relationship between machine intelligence and the physical world it inhabits - the long arc of spatial intelligence.Language gave machines a way to talk about that world. World models are how machines will finally come to understand, imagine, reason and interact with it.World Labs团队与李飞飞发文,梳理"世界模型"这一被滥用的术语。对比语言模型学习文本统计,世界模型学习空间与时间统计(如光照、物理规律)。基于部分可观马尔可夫决策过程(POMDP)框架,智能体通过动作影响世界状态,观测是部分视图。当前被称为"世界模型"的不同系统本质上是同一循环的不同投影:第一类为渲染器,输出给人眼看的像素,以视觉保真度为核心。文章着重于概念分层,未给出具体模型名、参数或基准分数。
推荐理由:李飞飞亲手给纷乱的「世界模型」下了个三分类——渲染、模拟、规划,而且点破模拟才是根基。做机器人、空间智能的人,这篇是今年的坐标系。

6月3日

星期三 · 1 条
01:40
HuggingFace Daily Papers(社区热门论文)精选
AI 评分 71/100
AFUN: 迈向功能理解的可供性基础模型

AFUN是一个用于功能理解的可供性基础模型。它从单个RGB-D观察和语言任务描述出发,能同时预测任务条件的功能掩码(where)和3D接触后运动曲线(how)。为实现开放世界泛化,该研究构建了一个大规模标准化数据管道,整合了机器人、人类、仿真与真实扫描数据。评估结果显示,AFUN在可供性分割任务上,于4个基准的8个测试集中平均gIoU/cIoU指标分别大幅领先基线模型+23.9/+26.3;在接触点预测上,命中率比最佳基线高出12.7%–61.3%;在3D运动预测上也取得最佳性能。该模型无需针对特定机器人实体进行微调即可直接部署。


推荐理由:在 affordance 基础模型方向做出一步,跨 8 个测试集大幅超越基线,并可直接部署到真实机器人,对具身智能的通用化是个值得关注的信号。

6月2日

星期二 · 1 条
01:03
Hacker News 热门(buzzing.cc 中文翻译)精选
AI 评分 75/100
英伟达 Cosmos 3

英伟达发布了 Cosmos 3,这是一个用于物理 AI 推理的世界和行动模型。该信息来源于英伟达开发者博客,发布日期为 2026 年 6 月 1 日。

另有 8 家信源报道IT之家(RSS)Hugging Face:Blog(RSS)LMSYS:Blog(Chatbot Arena 团队)HuggingFace Daily Papers(社区热门论文)X:Artificial Analysis (@ArtificialAnlys)X:Kim (@kimmonismus)X:Perplexity (@perplexity_ai)X:Satya Nadella (@satyanadella)
推荐理由:Cosmos 3 把物理推理、世界生成和行动生成塞进一个开源模型,从机器人到自动驾驶都能用,英伟达这次是真的想定义物理 AI 的训练范式。

6月1日

星期一 · 2 条
09:28
IT之家(RSS)精选
AI 评分 72/100
全球首次:MWC26 上海将举办"人形机器人点球大战",宇树科技等 8 支队伍参赛、参演

全球首次“人形机器人点球大战”将于2026年6月在MWC上海举行。8支中国顶尖具身智能战队将进行自主对抗,无需人工操控或预设脚本。赛事旨在集中展示人形机器人在动态平衡、精准控制与自主决策等方面的技术突破。


推荐理由:全球首次人形机器人点球大赛,不再是论文指标或仿真跑分,而是把动态平衡、自主决策塞进一场体育规则,具身智能的进展此刻比任何展台都诚实。
00:13
Sam Altman@sama精选
AI 评分 83/100
OpenAI正式进军机器人领域并启动招聘OpenAI Robotics is hiring, looking for exceptional full-stack hardware, ops, systems, and ML engineers to help us program and manufacture robots that are useful for society.AI should be able to help people in the physical world. In the short term, we are focused on robots to support skilled workers to build our future infrastructure; in the long term, we imagine everyone having a personal robot doing anything they need.Our world simulation research program, led by Aditya Ramesh (@model_mechanic), has evolved over the past year into OpenAI Robotics. Progress is rapid, and based on a foundation of co-design between robotics hardware and ML research.If you love working hands-on across the robotics stack and want to build the future, please consider joining us. Send an email with your background and evidence of exceptional accomplishment to: robotics-recruiting@openai.comOpenAI宣布成立OpenAI Robotics团队,并开始招聘全栈硬件、系统及ML工程师,以编程和制造能服务社会的机器人。该项目由Aditya Ramesh领导,其世界模拟研究计划已演变为机器人研究,强调硬件与ML研究的协同设计。短期目标是支持技术工人构建未来基础设施,长期愿景是为每个人提供个人机器人。另有 2 家信源报道IT之家(RSS)X:Emad Mostaque (@EMostaque)
推荐理由:OpenAI 正式踩进物理世界,从软件杀到硬件,这步迟早要来。短期说辅助工人,长期说人人都一个机器人,野心和风险一样大。

5月31日

星期日 · 2 条
08:00
HuggingFace Daily Papers(社区热门论文)精选
AI 评分 70/100
τ_0-WM:用于机器人操控的统一视频-动作世界模型

τ_0-World Model (τ_0-WM) 是一个统一的视频-动作世界模型,旨在机器人执行动作前预测并评估其未来后果。模型基于共享的视频扩散主干网络构建,提供两个接口:一个联合预测未来视觉潜在表示与连续动作块的视频动作模型,以及一个能将动作序列展开为多视角未来并预测任务进度分数的动作条件视频模拟器。τ_0-WM 使用约27,300小时的多元数据训练,包括真实机器人遥操作、UMI风格交互、自我中心人类视频等。推理时,模型通过测试时计算采样动作候选,并利用去噪一致性和基于模拟器的修正来筛选低质量动作,在长时程和精细机器人操控任务上表现出优于相关基准的性能。


推荐理由:机器人操作领域的大一统尝试,把视频预测和动作生成放在一个扩散模型里,还用27万小时数据训练,做具身智能的可以看看这个架构。
08:00
HuggingFace Daily Papers(社区热门论文)精选
AI 评分 70/100
定位何处:基础模型能否通过主动探索达到目标视角

研究提出目标视角复现任务(TVR)与模拟基准TVRBench,评估基础模型在3D环境中主动调整视角以匹配目标图像的能力。当前最优开源与闭源模型成功率仅7.8%和12.0%,瓶颈在于处理多轮视觉历史及需要平移而非旋转时的性能下降。通过构建统一的后训练框架,视觉动作SFT将9B开源模型成功率提升至50.8%,多轮GRPO进一步达到51.4%,为训练主动感知与行动的模型提供了基准。代码与模型已开源。


推荐理由:主动探索视角是具身智能的关键短板,这篇论文用一个新基准把问题量化了——目前最强的模型也只能对上12%的目标。他们同时放出了训练框架和代码,做空间智能的可以直接拿来跑。

5月29日

星期五 · 2 条
23:13
Qwen:Blog Retrieval(API)精选
AI 评分 66/100
Qwen-VLA:从理解世界到付诸行动

通义千问推出通用视觉-语言-动作模型Qwen-VLA,基于Qwen多模态骨干,将视觉感知、语言理解与空间推理扩展至连续动作生成和轨迹预测。训练分四阶段:文本到动作预训练(T2A)、持续预训练(CPT)、监督微调(SFT)和强化学习(RL)。在LIBERO上达97.9%,Simpler-WidowX达73.7%,RoboTwin-Easy/Hard达86.1%/87.2%,匹配或超越专精模型。数据涵盖超10,000小时公共机器人轨迹、1,000+小时内部真实轨迹及800万+合成仿真轨迹。


推荐理由:Qwen-VLA 把机器人操作、导航和跨实体控制统一进一个模型,在多个基准上打平甚至超越专用模型,这是通用具身智能的一个重要信号,但离实际可用还有距离。
22:53
公众号:通义实验室(千问)精选
AI 评分 61/100
Qwen-VLA:迈向通用具身智能的统一动作框架

通义实验室提出Qwen-VLA,以Qwen3.5-4B视觉语言主干与1.15B参数DiT动作解码器构建统一视觉-语言-动作模型。通过文本到动作DiT预训练和本体感知提示,将操作、导航与轨迹预测统一在同一框架下,支持11种机器人平台。在5个仿真基准中,单一通用模型在3个上超越最佳专用模型;ALOHA真机in-domain成功率83.6%,OOD泛化76.9%,分别超越π₀.₅超35和40个百分点;DOMINO动态操作零样本达26.6%;VLN-CE导航R2R和RxR分别达57.5%和59.6%,均超越专用模型。


推荐理由:通义把操作、导航和轨迹预测塞进一个脑子,在11种机器人上通用,这是具身智能从'专家'走向'通才'的关键一步,做机器人的值得翻翻论文。