# PlayWorld：用智能体玩家在长程目标下评测世界模型

- 来源：HuggingFace Daily Papers（社区热门论文）
- 发布时间：2026-08-13 08:00
- AIHOT 分数：54
- AIHOT 链接：https://aihot.virxact.com/items/cmssd3ass03trrod0gzy9h2vo
- 原文链接：https://arxiv.org/abs/2608.13552

## AI 摘要

香港大学与快手可灵团队推出 PlayWorld 基准，用多模态 Agent Player 模拟人类玩家，在 171 个带指定目标的场景中与世界模型交互，以长程目标而非固定动作轨迹进行跨模型公平对比。评测覆盖几何一致性、交互保真度、视野外演化与洞察演化四个核心维度，并纳入视频质量与可控性基础指标。对九个前沿世界模型的实验显示，现有模型在长程交互目标上仍不可靠，尤其在空间一致性与持续状态演化方面。

## 正文

Kaixin Ding

Xi Chen

Minghong Cai

Affiliation:

Zhiyuan Xu

Affiliation:

[5pt] The University of Hong Kong

Yiyang Wang

Affiliation:

[5pt] The University of Hong Kong

Yuxiang Lu

Affiliation:

[5pt] The University of Hong Kong

Junyi Li

Affiliation:

[5pt] The University of Hong Kong

Shuyang Chen

Affiliation:

Zhejiang University[5pt]

Leaderboard

Affiliation:

Kling Team, Kuaishou Technology

Xin Tao

Affiliation:

Kling Team, Kuaishou Technology

Affiliation:

Kling Team, Kuaishou Technology

Hengshuang Zhao

Abstract

Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically evaluates a world model by pursuing long-horizon objectives through interaction. For example, a user may turn around 360 degrees to see whether the environment remains consistent, or walk into the water and inspect whether realistic water ripples are generated. The action sequence required to achieve the same objective may vary substantially between models, making fixed action-conditioned evaluation unsuitable for cross-model comparison. To address this, we employ multi-modal Agent Players to interact with world models toward specified long-horizon objectives. Building on this paradigm, we introduce PlayWorld, a benchmark providing 171 scenarios, each with a specified objective. To evaluate performance thoroughly, we assess models along four core dimensions: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution. In addition, we incorporate basic ability metrics for video quality and controllability. Experiments across nine state-of-the-art world models reveal that current models remain unreliable on long-horizon interactive objectives, particularly in maintaining spatial consistency and persistent state evolution. Code and data are available at https://kxding.github.io/project/PlayWorld/

Work done at Kling Team, Kuaishou Technology.

Project Lead. Corresponding author.

Figure 1: PlayWorld evaluates world models from the perspective of a human player. We define long-horizon objectives such as “turning around 360 degrees” or “walking into the water”, and use an Agent Player to interact with the world models. We then evaluate them from multiple perspectives, including geometry consistency, interaction fidelity, and persistent state evolution.

1 Introduction

General video generation models have advanced rapidly in recent years, achieving remarkable visual quality and photorealism (22; 35; 28). Video world models (3; 10; 1) go a step further by simulating interactive environments that users can explore while maintaining causally consistent evolution across space and time. This turns them into open worlds that users can freely interact with.

As a rapidly growing body of work pursues this goal (47; 17; 13; 38; 9; 46; 31), evaluation has become an increasingly important challenge. General video-generation benchmarks such as VBench (16) assess aesthetics, temporal coherence, and text alignment, but do not measure whether a model responds correctly to user actions. World-model benchmarks further evaluate 3D consistency, memory consistency, interaction, and out-of-sight state evolution (7; 18; 5; 41; 39). However, these benchmarks typically rely on customized trajectories predefined for individual evaluation cases, as summarized in Tab. 1.

This evaluation paradigm does not fully match how human players judge interactive world models. In practice, a human user usually evaluates a world model by pursuing a high-level, long-horizon objective rather than by checking whether a predefined action trajectory is followed. For example, a customized trajectory may specify three → commands with the intent of rotating the camera by 360∘ and returning to the initial view of Marina Bay Sands, as illustrated in Fig. 1. Because action granularity varies across models, the same three commands may complete a full rotation in one rollout but only a partial turn in another. Marina Bay Sands may therefore leave and re-enter the view in some rollouts but fail to reappear in others, making the resulting geometry consistency scores incomparable. Such customized trajectories also cannot reliably complete more complex action patterns, such as orbiting a sculpture while maintaining a fixed viewing direction (Fig. 4), because the required action duration and stopping point differ across models. Moreover, web-based closed-source world models such as Genie 3 (10) and HappyOyster (1) are accessible only through a web interface and require direct human interaction, which is labor-intensive and prone to bias. These challenges call for an automated evaluation protocol that can interact with world models like a human player.

To address this challenge, we introduce PlayWorld, which evaluates world models using long-horizon objectives. PlayWorld uses Agent Player, a multi-modal agent that simulates a human player. Each evaluation case provides a basic action sequence, which serves as a shared initial reference across world models. Given this reference, the Agent Player observes the generated frames and action history at each step and adaptively adjusts execution through Keep, Stop, Extend, Correct, or End decisions. These decisions adapt the number and duration of actions to each model’s response, enabling fairer comparison under the same objectives. PlayWorld further introduces a VQA rubric verifier for open-ended interactive scenarios, covering four core world-model capabilities: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution. The overall evaluation framework assesses model behavior over rollouts lasting approximately 10 to 60 seconds using continuous W/A/S/D, ↑/↓/←/→, and WAIT controls. The benchmark contains 171 human-annotated cases evaluated across nine representative world models, revealing the capability boundaries of current systems.

2 Related Work

媒体内容 · 前往原文查看

Table 1: Comparison of world-model evaluation benchmarks. We compare input formats, closed-loop adaptation, long-horizon revisitation, evaluated capabilities, and rollout duration. Objective denotes a scene-grounded high-level goal, while and denote camera and action inputs. For PlayWorld, the Objective is accompanied by a basic action sequence that serves only as the Agent Player’s initial reference; State Evolution includes visible and out-of-sight evolution. ✓, ✗, and ∼ indicate full, no, and partial coverage, respectively.

Benchmark Input Closed-Loop Adaptation Long-horizon Revisit Geometry Consistency State Evolution Interaction Fidelity Time Scale / Video

WorldScore (7) Text + Image + ✗ ✗ ✓ ✗ ✗ ∼2–10 s

WorldMark (41) Image + ✗ ∼ ✓ ✗ ✗ 20 / 40 / 60 s

WBench (44) Text + Image + / ✗ ∼ ✓ ∼ ✓ 10–45 s

WorldRoamBench (40) Image + ✗ ∼ ✓ ∼ ✓ 10–60 s

MemoBench (5) Text + Image + ✗ ✗ ✓ ✓ ✗ 4–6 s

Omni-WorldBench (39) Text + Image + ✗ ✗ ∼ ✓ ✓ 3–6 s

PlayWorld (Ours) Text + Image + Objective ✓ ✓ ✓ ✓ ✓ 10–60 s

World models. Video world models are rapidly developing across diverse application domains (3; 13). In autonomous driving, world models predict how traffic scenes change in response to ego motion, vehicle controls, and surrounding agents. GAIA-1 (14), DriveDreamer (37), DrivingWorld (15), and Vista (8) generate future observations for scenario simulation, planning, and policy evaluation. In robotics, world models instead predict future visual observations conditioned on actions, context, or task instructions. IRASim (45), Cosmos (21), RoboScape (29), and LVP (4) use these predictions to support robot learning, planning, and control. Beyond these domain-specific applications, general-purpose interactive world models generate visual environments continuously in response to user actions rather than producing a fixed video from a text or image prompt. Representative systems include Matrix-Game (13; 38), YUME (19; 20), WorldPlay (31), LingBot-World (27), LingBot-World2 (9), SANA-WM (46), HY-World (34), HY-World2 (33), and Hunyuan-GameCraft (17; 32). These models progressively generate the environment in response to movement keys, mouse or gamepad controls, camera trajectories, or text instructions. Closed-source systems such as Genie 3 (10) and HappyOyster (1) provide similar interaction through user-facing web interfaces, although their implementation details remain unavailable. These systems use different architectures and control formats but share the goal of generating worlds that respond continuously and coherently to user actions. Open-ended exploration and adventure scenarios place users in direct control of the generated environment. We therefore consider them a fundamental setting.

Evaluation of world models. Existing world-model benchmarks evaluate several complementary capabilities. VBench (16) and T2V-CompBench (30) measure perceptual quality, temporal stability, and compositional alignment, but do not assess responses to user actions. WorldSimBench (24) evaluates perceptual quality and downstream embodied-control performance, without adapting evaluation controls according to generated observations. WorldScore (7) and WorldMark (41) assess camera control and multi-view geometry under predefined camera trajectories or action sequences. Their results may conflate geometric inconsistency with trajectory failure when identical controls produce different movement magnitudes across models. MIND (43), MemoBench (5), and STEVO-Bench (18) evaluate memory and out-of-sight evolution through revisitation, visible–disappear–reappear processes, and lookaway controls. These protocols follow fixed schedules that cannot account for model-specific action granularity and response speed. Omni-WorldBench (39) broadens evaluation to interaction, physics, and memory, while WBench (44) and WorldRoamBench (40) introduce predefined multi-turn navigation patterns. However, these patterns cannot be flexibly composed into diverse long-horizon trajectories or adjusted online when the generated observation deviates from the expected state. Overall, existing benchmarks predominantly drive world models with predefined low-level controls. Because models differ in action granularity and response dynamics, the same controls may fail to reach comparable evaluation states. PlayWorld specifies a shared long-horizon objective and uses a multi-modal Agent Player to adapt action execution online. This closed-loop protocol preserves a consistent evaluation intent across models while supporting more diverse long-horizon interaction.

3 PlayWorld

To better simulate how human players evaluate world models, PlayWorld introduces a multi-modal Agent Player that interacts with each evaluated model and adaptively adjusts action execution. The PlayWorld benchmark organizes long-horizon objectives across diverse scenario types. Its VQA rubric verifier assesses four core world-model capabilities: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution. We also combine basic quality metrics for complementary evaluation.

3.1 Agent Player Design

The Agent Player consists of a replaceable multi-modal agent model (e.g., Claude or Gemini) and an agent interface. The agent model observes generated frames and execution history, then decides how to adapt action execution toward the long-horizon objective. The agent interface executes these decisions on the evaluated world model, captures the resulting observation, and returns it to the agent model, forming a closed interaction loop.

Agent model decision. A straightforward way to instantiate the Agent Player is to let the agent plan the trajectory from scratch and interact with the world model without any human-provided action input. However, because each decision requires the agent model to reason over the current observation, fully online planning introduces substantial decision latency and slows action execution. It may also produce highly different trajectories across world models with different capabilities, making cross-model comparison less consistent. We therefore use a human-annotated basic action sequence as a shared initial reference. This preserves a consistent evaluation intent across models, reduces the burden of open-ended sequential planning, and still allows observation-conditioned adaptation. As shown on the left of Fig. 1, the Agent Player receives the long-horizon objective and the shared basic action sequence. At each interaction step, the agent model observes the generated frame, the previously executed actions, the scene description, and the objective. Based on this context, it decides whether to keep, stop, extend, correct, or end the current rollout. Keep retains the planned action, continuing the current hold or advancing to the next action as scheduled. Stop terminates the active action early when the intended visual state has been reached. Extend increases the action hold duration when additional movement is required. Correct revises or skips the next planned action when the observed state no longer matches the reference sequence. End terminates the evaluation case only when the Agent Player determines that the long-horizon objective and its required observation have been completed. For interaction fidelity cases involving an obstacle, reaching the obstacle does not trigger End; the Agent Player continues the forward action so that the rollout reveals whether the subject collides with the obstacle realistically or penetrates it.

Agent interface execution. The agent interface directly interacts with the evaluated world model. It receives an initial frame and a human-annotated basic action sequence expressed using W/A/S/D, ↑/↓/←/→, and WAIT, and translates these actions into the world model’s native controls. At each step, it executes the active action, captures the resulting generated frame, returns the frame to the agent model, and applies the returned decision to update subsequent execution. This process continues in a closed loop until the Agent Player returns End or the 40-step interaction budget is exhausted. For web-based world models, the agent interface uses browser automation to dispatch controls and capture generated frames directly from the model interface. Further details are provided in the supplementary.

Figure 2: Benchmark construction pipeline of PlayWorld. Starting from diverse initial worlds, annotators define long-horizon objectives based on scenario-specific attributes, write basic action sequences, and construct sample-specific VQA rubrics. The resulting benchmark spans over 170 human-annotated cases and 50 action patterns over 10–60-second rollouts, with more than 820 questions applied to over 1,400 interactive videos.

3.2 PlayWorld Benchmark Overview

Benchmark construction. To support long-horizon evaluation, we construct the PlayWorld benchmark around three core elements: an initial world, a scene-grounded objective, and a basic action sequence. As shown in Fig. 2, we first collect candidate images from Pexels (23) and Google Images (12), covering natural, urban, and fantasy environments with various subjects such as humans, animals, and vehicles. Annotators then screen the images for clear scene structure and select first-person or third-person views according to the target capability and scene content. For each selected image, Gemini 3.1 Pro (11) generates an environment caption, which is verified and revised by human annotators. Annotators further define a long-horizon objective based on scenario-specific attributes, specify a visually verifiable completion condition, and write a basic action sequence. Finally, annotators construct sample-specific VQA rubrics for the target evaluation dimension. All cases are manually curated to ensure that the expected outcome is observable and suitable for dimension-specific evaluation.

Benchmark composition. Fig. 3 summarizes the composition of PlayWorld. The benchmark comprises over 170 human-annotated cases covering 50 action patterns, ranging from a single action to combinations of up to five actions and producing 10–60-second rollouts. Evaluating nine world models yields over 1,400 interactive videos, which are assessed using more than 820 task-conditioned VQA questions. The cases span four evaluation dimensions and cover six primary subject types. The VQA questions are organized into four evidence families: identity and appearance, physics and dynamics, temporal evolution, and spatial and trajectory reasoning. Fig. 4 provides representative examples for each dimension. Geometry consistency cases emphasize revisitable visual anchors such as lettering and buildings with distinctive shapes; interaction fidelity cases include walking into water, fire, cliffs, or grass; out-of-sight evolution cases emphasize continued changes across disappear–reappear intervals, such as a person peeling an apple, making dessert, or painting a picture; and insight evolution cases focus on long visible processes without an explicit action trajectory, such as shoppers queuing for checkout or a chef plating food.

媒体内容 · 前往原文查看

Figure 3: Composition of PlayWorld. The benchmark comprises over 170 human-annotated cases across diverse scenario categories. The cases span four evaluation dimensions and diverse subject types, with dimension-specific VQA questions tailored to their evaluation criteria.

Figure 4: Representative cases in PlayWorld. Each case combines an initial world, a scene-grounded long-horizon objective, and a sample-specific VQA rubric. The expected trajectory is visualized alongside each case; for insight evolution, the camera remains stationary and observes the scene for 60 seconds.

3.3 VQA Rubric Verifier

Automated metrics primarily capture pixel-level fidelity and low-level perceptual quality, but often fail to measure consistency under large viewpoint changes, assess interaction plausibility, or determine whether long-horizon objectives are completed. We therefore design a VQA rubric verifier, where Gemini 3.1 Pro (11) answers sample-specific Yes/No rubric questions for each generated rollout. Each rubric assigns greater importance to questions that are more relevant to the target capability and excludes questions that are not applicable to the case. Because instruction-following failures can make memory evaluation unreliable, as discussed in MemoBench (5), the verifier uses dimension-specific validation as a gate before rubric scoring. The weighted Yes/No outcomes are aggregated and converted into a case-level score on a 1–5 scale (↑: higher is better), as reported in Tab. 2. The evaluation covers four dimensions:

Geometry consistency. To quantitatively assess long-horizon geometry consistency, the rubric considers whether the environment preserves its structure during camera movement and revisitation. We first verify Trajectory Validity by determining whether the rollout follows the intended path and reaches the required viewpoint. Once the rollout passes this validation, its sample-specific VQA rubric covers two aspects. Scene Identity checks the preservation of object identity, count, texture, material, and color, such as whether a wall mural preserves its color and shape after a 360-degree turn or whether a revisited sculpture remains unchanged. Spatial Consistency considers relative object positions and the overall scene layout, for example whether a red car that is initially to the right of a green car appears in the correct left-right order after the camera crosses the street and looks back.

Interaction fidelity. For interaction fidelity, we examine whether the intended interaction occurs and elicits a physically plausible response. Before rubric scoring, Subject and Reachability determines whether the controlled subject responds to the actions and reaches the intended interaction region. A completely static subject fails this validation because the intended interaction is never tested. The subsequent rubric covers Contact and Collision, Motion and Causality, and Visual Response. For instance, it asks whether the subject collides with a wall rather than passing through it and whether its motion follows plausible kinematics, such as maintaining grounded footsteps without floating. It also checks whether entering water produces appropriate physical feedback, such as visible ripples, splashes, or changes in apparent water depth around the subject.

Out-of-sight evolution. Out-of-sight evolution focuses on whether a target preserves its identity and continues to change while unobserved. We first use Trajectory Validity to determine whether the target leaves the field of view and subsequently reappears as required. For eligible rollouts, Reappearance Consistency asks whether the revealed target is the same entity observed before it disappeared. Hidden-State Evolution determines whether the target state progresses plausibly during the unobserved interval, such as whether an in-progress drawing gains additional content while it is out of sight. Physical Causality checks whether the revealed state can reasonably result from that hidden evolution, such as a person continuing to move upward on an escalator, blue ink spreading through clear water, or a cup becoming fuller as water is poured into it.

Insight evolution. Insight evolution focuses on continuously observed phenomena without a fixed action trajectory. The rubric evaluates a visible subject or process from three aspects. Identity and State Progression checks whether the relevant person or object remains present and whether its state progresses over time, rather than suddenly disappearing. Motion and Physical Plausibility identifies static behavior, repetitive but non-progressive actions, abrupt transitions, and implausible dynamics. For example, if a person is eating one of three buns, the rubric checks whether the bun being eaten gradually changes and whether the number or appearance of the remaining buns evolves consistently. Meanwhile, Temporal Scene Consistency detects unintended changes to the surrounding environment. Because this dimension uses stationary observation, it does not require trajectory validation.

媒体内容 · 前往原文查看

Table 2: Rubric-based VQA evaluation across four dimensions on the PlayWorld benchmark. Scores range from 1 to 5 (↑: higher is better). Overall is the unweighted mean across the four evaluation dimensions. Bold: best; underline: second best.

Model Geometry Consistency Interaction Fidelity Insight Evolution Out-of-sight Evolution Overall

Closed-source Models

Genie 3 2.74 2.40 1.51 1.81 2.12

HappyOyster 2.54 2.15 1.47 1.54 1.92

Open-source Models

LingBot-World 2.11 2.23 1.33 1.43 1.78

LingBot-World2 2.04 2.13 1.95 1.16 1.82

HY-World2 2.14 2.06 1.13 1.09 1.61

SANA-WM 1.72 1.89 1.13 1.16 1.48

Hunyuan-GameCraft-2 1.62 1.52 1.21 1.31 1.42

HY-WorldPlay 1.12 1.63 1.01 1.08 1.21

Matrix-Game-3.0 1.30 1.25 1.00 1.00 1.14

媒体内容 · 前往原文查看

Table 3: Trajectory-validation pass rates on the PlayWorld benchmark. The insight evolution dimension uses stationary observation and therefore does not require trajectory validation and is marked as “–”. Overall is the unweighted mean of the pass rates for geometry consistency, interaction fidelity, and out-of-sight evolution.

Model Geometry Consistency Interaction Fidelity Insight Evolution Out-of-sight Evolution Overall

Closed-source Models

Genie 3 77.1% 93.5% – 90.7% 87.1%

HappyOyster 63.8% 93.5% – 81.4% 79.6%

Open-source Models

LingBot-World 58.3% 76.0% – 83.7% 72.7%

LingBot-World2 58.3% 85.1% – 93.0% 78.8%

HY-World2 47.9% 70.0% – 37.2% 51.7%

SANA-WM 62.5% 88.0% – 90.7% 80.4%

Hunyuan-GameCraft-2 25.0% 78.0% – 81.4% 61.5%

HY-WorldPlay 14.6% 80.0% – 30.2% 41.6%

Matrix-Game-3.0 37.5% 88.0% – 79.1% 68.2%

3.4 Basic Ability Evaluation

We further use automatic metrics to test video quality and controllability. We choose Aesthetic Quality, Imaging Quality, Motion Smoothness, and Temporal Flickering from VBench (16), Temporal Consistency from Omni-WorldBench (39), Depth Stability from MemoBench (5), and Subject Consistency from HyDRA (6). These video quality metrics do not require ground-truth videos and mainly capture frame-level or inter-frame signals. All video quality metrics are normalized to [0,1] and reported as percentages. We then evaluate Action Controllability using Translation Pass Rate and Rotation Pass Rate, following WorldMark (41). Camera poses are estimated by VGGT (36) from uniformly sampled frames, while target trajectories are defined by the executed actions, including online adjustments made by the Agent Player. A rollout passes the translation or rotation criterion when the corresponding error falls below an acceptable threshold. Both pass rates are computed over rollouts with valid estimates and reported as percentages, where higher values indicate better controllability. Tab. 6 reports all nine higher-is-better metrics together with a Basic Ability Score, which converts their average rank into a percentage.

4 Experiments

4.1 Models and Evaluation Protocol

Evaluation models. We evaluate nine representative world models. Five models are evaluated through their web-based interfaces to reproduce the user-facing interaction setting: Genie 3 (10), LingBot-World (27), LingBot-World2 (9), HY-World2 (33), and HappyOyster (1). Four models are executed locally through chunk-wise generation: SANA-WM (46), Hunyuan-GameCraft-2 (32), HY-WorldPlay (31), and Matrix-Game-3.0 (38). All models are controlled by the same Agent Player, whose agent interface converts the decisions into each model’s native input format. Implementation details and generation configurations are provided in the supplementary material.

Figure 5: Representative failures across four evaluation dimensions. Current video world models can follow camera or action controls, but still struggle to simulate coherent and realistic world dynamics.

Evaluation protocol. For each evaluation case, all models receive the same initial world, long-horizon objective, and basic action sequence. The same Agent Player observes the generated frames and adjusts action execution for every model. We evaluate the resulting rollouts, which last approximately 10–60 seconds, using the rubric verifier in Sec. 3.3 and the basic ability metrics in Sec. 3.4. The rubric verifier reports geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution, while the basic ability evaluation reports Video Quality and Action Controllability. We also report dimension-specific validation pass rates in Tab. 3. Specifically, geometry consistency is validated using Trajectory Validity, interaction fidelity using Subject and Reachability, and out-of-sight evolution using Trajectory Validity. These three validation criteria are evaluated independently. If a rollout does not pass the corresponding validation criterion, its score for that dimension is set to the minimum value of 1. The insight evolution dimension uses stationary observation and therefore does not require trajectory validation.

4.2 Agent Player Analysis

Control strategy. Existing benchmarks typically execute a predefined trajectory for each evaluation case. We retain a basic action sequence as a shared initial reference but adapt its execution online to account for differences in action granularity across world models. Directly executing the basic action sequence without online adaptation (Preset Only) may cause the same controls to overshoot or undershoot the intended state. Agent Only instead predicts every action online from the latest generated observation and previously executed actions, but planning each action from scratch increases the sequential inference burden and slows trajectory execution. Preset + Agent combines both sources of guidance: it uses the basic action sequence as a reference while adapting its execution to the generated observations and model-specific control granularity. We compare these three strategies on the same 25 cases for Genie 3 and HappyOyster.

As shown in Tab. 4, Preset + Agent achieves the highest Trajectory Score and Human Preference on both world models, confirming the benefit of combining a shared action prior with online adaptation. Its Agent-modified Action Ratio remains within 10–20% on both models, showing that the Agent makes targeted modifications. By adapting action execution online, PlayWorld reaches the intended objective more reliably than executing a predefined trajectory alone.

Agent model selection. We further investigate the effect of multi-modal model choice on the trajectories produced by the Agent Player. We compare three agent models on the same Genie 3 cases using PlayWorld’s control strategy. As shown in Tab. 5, agent model choice produces only minor differences in trajectory quality.

We attribute this limited sensitivity to the relatively constrained decision task, in which each agent model adapts actions based on the current visual state and basic action sequence rather than planning an open-ended trajectory from scratch. Claude Sonnet achieves the highest Trajectory Score, whereas Claude Haiku achieves the highest Human Preference and the lowest measured Decision Latency. Because latency depends strongly on provider infrastructure and network conditions, the measured values may vary across deployment environments. Considering the overall quality–efficiency trade-off, we use Claude Haiku, and the agent model can be readily upgraded as stronger multi-modal models emerge.

Implementation details. For each comparison, we aggregate judgments from five independent raters with video-generation experience. Trajectory Score is assigned by Gemini 3.1 Pro on a 0–2 scale. Human Preference aggregates explicit pairwise preference judgments, assigning full credit to wins and half credit to ties. For the control-strategy comparison, Preset Only and Agent Only are each compared against Preset + Agent, whose rate pools comparisons against both baselines. For the agent model comparison, each agent model pools comparisons against the other agent models. The comparison order and left–right placement of the paired trajectories are randomized in the questionnaire. Majority Agreement is the percentage of pairwise comparisons for which more than half of the raters select the same outcome. Agent-modified Action Ratio is the macro-averaged per-case fraction of actions modified or inserted by the Agent Player. Decision Latency, reported only for the agent model comparison, is the average end-to-end waiting time per agent-model call.

媒体内容 · 前往原文查看

Table 4: Comparison of trajectory control strategies. Preset Only executes the basic action sequence without online adaptation. Agent Only selects all actions online from generated observations and executed-action history. Preset + Agent adapts action execution online using the basic action sequence as a reference.

Setting Trajectory Score ↑ Human Preference ↑ Majority Agreement Agent-modified Action Ratio

Genie 3

Preset Only 0.92 39.6% 88.0% 0.0%

Agent Only 0.88 29.2% 84.0% 100.0%

Preset + Agent 1.08 65.6% 86.0% 12.0%

HappyOyster

Preset Only 1.00 40.8% 76.0% 0.0%

Agent Only 0.68 24.4% 92.0% 100.0%

Preset + Agent 1.12 67.4% 84.0% 14.9%

媒体内容 · 前往原文查看

Table 5: Comparison of agent models. We conduct the comparison on Genie 3, with Preset + Agent control strategy. We recommend Claude Haiku as the agent model.

Agent Model Trajectory Score ↑ Human Preference ↑ Majority Agreement Agent-modified Action Ratio Decision Latency (s/call) ↓

Claude Haiku 1.08 57.8% 88.9% 12.0% 3.83

Claude Sonnet 1.24 45.3% 94.1% 12.4% 6.21

Gemini 3.1 Pro 1.08 46.5% 94.1% 12.6% 4.36

4.3 Observations in World Models

Sustained world evolution remains the primary bottleneck. Tab. 2 shows that Genie 3 achieves the highest geometry consistency, interaction fidelity, out-of-sight evolution, and overall scores, while LingBot-World2 leads insight evolution and HappyOyster ranks second overall. However, out-of-sight evolution and insight evolution consistently receive lower scores across models. This suggests that current world models are better at preserving visible structure or producing immediate responses than sustaining semantic state changes over long horizons. In practice, during continuous observation, an ongoing process may remain static, reset, or evolve without preserving its causal order. After leaving the field of view, a target may reappear unchanged, return at an incorrect location, or be replaced by another entity. Together, these results indicate that persistent state evolution across time and occlusion remains a central challenge for current world models.

Long-horizon revisitation exposes global spatial inconsistency. Fig. 5 shows that failures that appear minor in individual frames can accumulate into structural inconsistencies over a complete rollout. Across multiple orbit and revisit cases, salient landmarks are repeatedly generated at new viewpoints rather than anchored to unique spatial locations. For example, when the camera orbits 360 degrees around the Taj Mahal, the model may repeatedly regenerate the monument at different viewpoints, creating multiple inconsistent copies of the same landmark. Although each frame remains visually plausible, the complete sequence violates spatial uniqueness and global scene topology. This failure is consistent with current models relying heavily on local appearance continuity and short-range motion cues, without reliably maintaining a persistent global representation of the generated 3D world.

Interaction fidelity remains limited beyond simple collision responses. Current models handle simple interactions reasonably well. In third-person cases, several leading models can stop a character at an obstacle boundary, whereas first-person movement may still pass directly through solid obstacles, as shown in Fig. 5. More complex interactions, such as walking into water, often fail to produce the expected physical or visual response. In addition, third-person models do not consistently bind input actions to the controlled character, resulting in delayed, insufficient, or unrelated motion.

媒体内容 · 前往原文查看

Table 6: Basic ability of video world models on the PlayWorld benchmark. Video Quality includes Aesthetic Quality, Imaging Quality, Motion Smoothness, Temporal Flickering, Temporal Consistency, Depth Stability, and Subject Consistency. Action Controllability includes Translation Pass Rate and Rotation Pass Rate. Each pass rate is computed over rollouts with a valid estimate. Basic Ability Score converts the average rank across all nine metrics into a higher-is-better percentage. All values are reported as percentages (%). Bold: best; underline: second best.

Model Video Quality Action Controllability Basic Ability Score ↑

Aes. ↑ Img. ↑ Mot. ↑ Flick. ↑ Temp. ↑ Depth ↑ Subj. ↑ Translation ↑ Rotation ↑

Closed-source Models

Genie 3 52.00 75.22 99.00 97.79 98.64 88.70 85.60 64.1 50.6 72.2

HappyOyster 49.66 73.66 99.46 99.02 99.52 91.80 88.40 58.0 48.5 76.4

Open-source Models

LingBot-World 51.12 72.04 97.96 96.27 98.02 92.20 83.50 64.8 44.7 51.4

LingBot-World2 51.47 73.98 97.84 95.64 96.88 91.40 82.30 67.9 45.6 50.0

HY-World2 47.43 67.79 99.12 99.46 91.17 90.70 89.80 40.8 42.2 45.8

SANA-WM 51.75 72.59 98.92 96.48 98.20 84.40 85.90 78.6 42.2 56.9

Hunyuan-GameCraft-2 50.14 67.82 98.25 95.24 97.97 90.40 83.70 89.6 29.4 40.3

HY-WorldPlay 47.71 61.53 98.89 95.18 95.45 91.00 87.40 75.8 44.2 38.9

Matrix-Game-3.0 44.10 66.03 98.90 96.92 95.23 87.30 65.70 42.4 27.7 18.1

Trajectory control and world-model capability can diverge. Tab. 3 reports whether rollouts follow the trajectories required by the objectives. The results show that strong trajectory control does not necessarily imply strong world-model ability. For example, SANA-WM obtains a relatively high overall validation pass rate, suggesting that its trajectory control often reaches the intended objective region. Nevertheless, its rubric scores remain modest. This gap indicates that SANA-WM can often follow the required trajectory, but still struggles to preserve memory, maintain spatial consistency, or generate physically plausible interactions.

Automatic metrics only test the basic ability of interactive world models. As shown in Tab. 6, most evaluated models perform well on Video Quality, while their Action Controllability remains less precise. Moreover, passing the action-control thresholds does not guarantee that a rollout reaches the objective-specific target state or completes its long-horizon objective. Depth Stability and Subject Consistency may also remain high when a model produces little or no camera motion, even though the intended evaluation objective is not completed. Although HappyOyster achieves the highest Basic Ability Score, this does not imply that it has the strongest long-horizon world-model capability.

Figure 6: Alignment between human preferences and VQA metrics.

4.4 Human Validation

We conduct a human evaluation to examine whether the VQA metrics align with human perception. Participants provide 600 valid pairwise judgments on video pairs evenly sampled across the four dimensions. We aggregate pairwise wins for each model within each dimension and normalize the resulting preference scores. Each point in Fig. 6 represents an evaluated model; the red line shows a linear fit, and the shaded region denotes its confidence band. Spearman’s ρ measures rank agreement across the nine models for each evaluation dimension and the Overall score. The consistently positive correlations show that the VQA metrics closely preserve human preference rankings. Further details are provided in Appendix D.

5 Conclusion

We have presented PlayWorld, a benchmark that uses multi-modal Agent Players to simulate how human users evaluate interactive video world models. With 171 human-annotated cases and an end-to-end automated evaluation system, PlayWorld supports consistent evaluation under shared long-horizon objectives. Together, PlayWorld establishes a practical basis for more precise evaluation of video world models.

Acknowledgments

We thank Jing He, Zhekai Chen, Wenwang Huang, Junjie Wang, Xianzhe Fan, Chenhao Ye, Tianxing Wu, and Junwei Luo for fruitful discussions.

References

Alibaba Cloud Community (2026) Alibaba Cloud Community Alibaba launches happyoyster, a world model product for real-time immersive creation and interaction. Note: https://www.alibabacloud.com/blog/alibaba-launches-happyoyster-a-world-model-product-for-real-time-immersive-creation-and-interaction_603048Alibaba Cloud Community blog post Cited by: §1, §1, §2, §4.1.

Anthropic (2025) Anthropic Claude haiku 4.5. Note: https://platform.claude.com/docs/en/about-claude/models/overviewClaude model documentation, model ID: claude-haiku-4-5-20251001 Cited by: Appendix A.

Bruce et al. (2024) J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y. Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, Y. Aytar, S. Bechtle, F. Behbahani, S. Chan, N. Heess, L. Gonzalez, S. Osindero, S. Ozair, S. Reed, J. Zhang, K. Zolna, J. Clune, N. de Freitas, S. Singh, and T. Rocktäschel Genie: Generative Interactive Environments. arXiv. External Links: 2402.15391, Document Cited by: §1, §2.

Chen et al. (2026a) B. Chen, T. Zhang, H. Geng, C. Zhang, P. Li, K. Song, W. T. Freeman, J. Malik, P. Abbeel, R. Tedrake, V. Sitzmann, and Y. Du Large Video Planner Enables Generalizable Robot Control. arXiv. External Links: 2512.15840, Document Cited by: §2.

Chen et al. (2026b) H. Chen, K. Zhou, H. Hua, K. Zhang, J. Qian, W. Ma, H. Chen, C. Liu, Y. Zhao, X. Wang, W. Li, A. Yuille, P. P. Liang, and Y. Du MemoBench: Benchmarking World Modeling in Dynamically Changing Environments. arXiv preprint arXiv:2606.27537. Cited by: Appendix A, §1, Table 1, §2, §3.3, §3.4.

Chen et al. (2026c) K. Chen, D. Liang, X. Zhou, Y. Ding, X. Liu, P. Wan, and X. Bai Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models. arXiv. External Links: Document Cited by: Appendix A, §3.4.

Duan et al. (2025) H. Duan, H. Yu, S. Chen, L. Fei-Fei, and J. Wu WorldScore: A Unified Evaluation Benchmark for World Generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 27713–27724. Cited by: §1, Table 1, §2.

Gao et al. (2024) S. Gao, J. Yang, L. Chen, K. Chitta, Y. Qiu, A. Geiger, J. Zhang, and H. Li Vista: a generalizable driving world model with high fidelity and versatile controllability. arXiv preprint arXiv:2405.17398. Cited by: §2.

Gao et al. (2026) Z. Gao, Q. Wang, J. Zhu, J. Chen, Z. Liu, Q. Bai, J. Wang, Y. Yuan, H. Wang, Y. Lu, K. L. Cheng, H. Zhang, J. Gao, T. Feng, Y. Liu, Y. Yao, Y. Xu, X. Zhu, Y. Shen, and H. Ouyang Infinite worlds with versatile interactions. External Links: 2607.07534, Document, Link Cited by: §1, §2, §4.1.

Google DeepMind (2025) Google DeepMind Genie 3: A new frontier for world models. Note: https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/ Cited by: §1, §1, §2, §4.1.

Google (2026a) Google Gemini 3.1 pro. Note: https://ai.google.dev/gemini-api/docs/modelsGemini API model documentation Cited by: Appendix A, §3.2, §3.3.

Google (2026b) Google Google images. Note: https://images.google.com/Accessed July 2026 Cited by: §3.2.

He et al. (2025) X. He, C. Peng, Z. Liu, B. Wang, Y. Zhang, Q. Cui, F. Kang, B. Jiang, M. An, Y. Ren, B. Xu, H. Guo, K. Gong, S. Wu, W. Li, X. Song, Y. Liu, Y. Li, and Y. Zhou Matrix-Game 2.0: An open-source real-time and streaming interactive world model. arXiv. External Links: 2508.13009, Document Cited by: §1, §2.

Hu et al. (2023) A. Hu, L. Russell, H. Yeo, Z. Murez, G. Fedoseev, A. Kendall, J. Shotton, and G. Corrado GAIA-1: a generative world model for autonomous driving. arXiv preprint arXiv:2309.17080. Cited by: §2.

Hu et al. (2024) X. Hu, W. Yin, M. Jia, J. Deng, X. Guo, Q. Zhang, X. Long, and P. Tan DrivingWorld: constructing world model for autonomous driving via video GPT. arXiv preprint arXiv:2412.19505. Cited by: §2.

Huang et al. (2024) Z. Huang, Y. He, J. Yu, F. Zhang, C. Si, Y. Jiang, Y. Zhang, T. Wu, Q. Jin, N. Chanpaisit, Y. Wang, X. Chen, L. Wang, D. Lin, Y. Qiao, and Z. Liu VBench: Comprehensive Benchmark Suite for Video Generative Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 21807–21818. Cited by: Appendix A, §1, §2, §3.4.

Li et al. (2025) J. Li, J. Tang, Z. Xu, L. Wu, Y. Zhou, S. Shao, T. Yu, Z. Cao, and Q. Lu Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition. arXiv. External Links: 2506.17201, Document Cited by: §1, §2.

Ma et al. (2026) Z. Ma, M. Liufu, and G. Gkioxari Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models. arXiv. External Links: 2603.13215, Document Cited by: §1, §2.

Mao et al. (2025a) X. Mao, Z. Li, C. Li, X. Xu, K. Ying, T. He, J. Pang, Y. Qiao, and K. Zhang Yume-1.5: A Text-Controlled Interactive World Generation Model. arXiv. External Links: 2512.22096, Document Cited by: §2.

Mao et al. (2025b) X. Mao, S. Lin, Z. Li, C. Li, W. Peng, T. He, J. Pang, M. Chi, Y. Qiao, and K. Zhang Yume: An Interactive World Generation Model. arXiv. External Links: 2507.17744, Document Cited by: §2.

NVIDIA et al. (2025) NVIDIA, N. Agarwal, A. Ali, M. Bala, Y. Balaji, E. Barker, T. Cai, P. Chattopadhyay, Y. Chen, Y. Cui, Y. Ding, D. Dworakowski, J. Fan, M. Fenzi, F. Ferroni, S. Fidler, D. Fox, S. Ge, Y. Ge, J. Gu, S. Gururani, E. He, J. Huang, J. Huffman, P. Jannaty, J. Jin, S. W. Kim, G. Klár, G. Lam, S. Lan, L. Leal-Taixe, A. Li, Z. Li, C. Lin, T. Lin, H. Ling, M. Liu, X. Liu, A. Luo, Q. Ma, H. Mao, K. Mo, A. Mousavian, S. Nah, S. Niverty, D. Page, D. Paschalidou, Z. Patel, L. Pavao, M. Ramezanali, F. Reda, X. Ren, V. R. N. Sabavat, E. Schmerling, S. Shi, B. Stefaniak, S. Tang, L. Tchapmi, P. Tredak, W. Tseng, J. Varghese, H. Wang, H. Wang, H. Wang, T. Wang, F. Wei, X. Wei, J. Z. Wu, J. Xu, W. Yang, L. Yen-Chen, X. Zeng, Y. Zeng, J. Zhang, Q. Zhang, Y. Zhang, Q. Zhao, and A. Zolkowski Cosmos World Foundation Model Platform for Physical AI. arXiv. External Links: 2501.03575, Document Cited by: §2.

OpenAI (2024) OpenAI Video generation models as world simulators. Note: https://openai.com/index/video-generation-models-as-world-simulators/ Cited by: §1.

Pexels (2026) Pexels Pexels: free stock photos and videos. Note: https://www.pexels.com/Accessed: 2026-05-03 Cited by: §3.2.

Qin et al. (2024) Y. Qin, Z. Shi, J. Yu, X. Wang, E. Zhou, L. Li, Z. Yin, X. Liu, L. Sheng, J. Shao, L. Bai, W. Ouyang, and R. Zhang WorldSimBench: Towards Video Generation Models as World Simulators. arXiv. External Links: 2410.18072, Document Cited by: §2.

Radford et al. (2021) A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, M. Meila and T. Zhang (Eds.), Proceedings of Machine Learning Research, pp. 8748–8763. External Links: Link Cited by: Appendix A.

Redmon et al. (2016) J. Redmon, S. K. Divvala, R. B. Girshick, and A. Farhadi You only look once: unified, real-time object detection. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pp. 779–788. External Links: Link, Document Cited by: Appendix A.

Robbyant Team et al. (2026) Robbyant Team, Z. Gao, Q. Wang, Y. Zeng, J. Zhu, K. L. Cheng, Y. Li, H. Wang, Y. Xu, S. Ma, Y. Chen, J. Liu, Y. Cheng, Y. Yao, J. Zhu, Y. Meng, K. Zheng, Q. Bai, J. Chen, Z. Shen, Y. Yu, X. Zhu, Y. Shen, and H. Ouyang Advancing Open-source World Models. arXiv. External Links: 2601.20540, Document Cited by: §2, §4.1.

Seedance et al. (2026) T. Seedance, D. Chen, L. Chen, X. Chen, Y. Chen, Z. Chen, Z. Chen, F. Cheng, T. Cheng, Y. Cheng, M. Chi, X. Chi, J. Cong, Q. Cui, F. Ding, Q. Dong, Y. Du, H. Duanmu, J. Fan, J. Fang, J. Fang, Z. Fang, C. Feng, Y. Gao, D. Gu, D. Guo, H. Guo, Q. Guo, B. Hao, H. Hao, H. He, J. He, Q. He, T. Hoang, H. Hu, R. Hu, Y. Hu, J. Huang, W. Huang, Z. Huang, Z. Huang, J. Jin, M. Jing, A. Kim, S. Lao, Y. Leng, B. Li, G. Li, H. Li, H. Li, J. Li, M. Li, X. Li, X. Li, Y. Li, Y. Li, Y. Li, Y. Li, C. Liang, H. Liang, J. Liang, Y. Liang, W. Liao, J. H. Lien, S. Lin, X. Lin, F. Ling, Y. Ling, F. Liu, J. Liu, J. Liu, J. Liu, S. Liu, S. Liu, W. Liu, X. Liu, Z. Liu, R. Lu, L. Lyu, J. Ma, T. Ma, X. Nie, J. Ning, J. Pan, X. Pan, R. Peng, X. Qu, Y. Ren, Y. Shen, G. Shi, L. Shi, Y. Song, F. Sun, L. Sun, R. Sun, W. Tang, B. Tao, Z. Tao, D. Wang, F. Wang, H. Wang, K. Wang, Q. Wang, R. Wang, S. Wang, S. Wang, W. Wang, X. Wang, Y. Wang, Y. Wang, Y. Wang, Y. Wang, Z. Wang, Z. Wang, G. Wei, M. Wei, D. Wu, G. Wu, H. Wu, H. Wu, J. Wu, J. Wu, R. Wu, S. Wu, X. Wu, X. Wu, Y. Wu, R. Xia, X. Xia, X. Xiao, S. Xu, B. Yang, J. Yang, R. Yang, T. Yang, Y. Yang, Z. Yang, Z. Yang, F. Ye, B. Yi, X. Yin, Y. You, L. Yuan, W. Zeng, X. Zeng, Y. Zeng, S. Zhai, Z. Zhai, B. Zhang, C. Zhang, H. Zhang, J. Zhang, M. Zhang, P. Zhang, S. Zhang, X. Zhang, X. Zhang, X. Zhang, X. Zhang, Y. Zhang, Z. Zhang, H. Zhao, H. Zhao, L. Zhao, Y. Zhao, G. Zheng, J. Zheng, X. Zheng, Z. Zheng, K. Zhu, and F. Zuo Seedance 2.0: Advancing Video Generation for World Complexity. arXiv. External Links: Document Cited by: §1.

Shang et al. (2025) Y. Shang, X. Zhang, Y. Tang, L. Jin, C. Gao, W. Wu, and Y. Li RoboScape: physics-informed embodied world model. arXiv preprint arXiv:2506.23135. Cited by: §2.

Sun et al. (2025a) K. Sun, K. Huang, X. Liu, Y. Wu, Z. Xu, Z. Li, and X. Liu T2V-CompBench: a comprehensive benchmark for compositional text-to-video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8406–8416. External Links: Document Cited by: §2.

Sun et al. (2025b) W. Sun, H. Zhang, H. Wang, J. Wu, Z. Wang, Z. Wang, Y. Wang, J. Zhang, T. Wang, and C. Guo WorldPlay: towards long-term geometric consistency for real-time interactive world modeling. External Links: 2512.14614, Document, Link Cited by: §1, §2, §4.1.

Tang et al. (2026) J. Tang, J. Liu, J. Li, L. Wu, H. Yang, P. Zhao, S. Gong, X. Yuan, S. Shao, L. Zhang, and Q. Lu Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model. arXiv. External Links: 2511.23429, Document Cited by: §2, §4.1.

Team HY-World et al. (2026) Team HY-World, C. Cao, X. Zuo, Z. Wang, Y. Zhang, J. Wu, Z. Liu, Y. Gong, Y. Liu, B. Yuan, et al. HY-World 2.0: a multi-modal world model for reconstructing, generating, and simulating 3d worlds. External Links: 2604.14268, Document, Link Cited by: §2, §4.1.

Tencent Hunyuan (2025) Tencent Hunyuan HY-World 1.5: a systematic framework for interactive world modeling with real-time latency and geometric consistency. Note: Technical reporthttps://3d-models.hunyuan.tencent.com/world/world1_5/HYWorld_1.5_Tech_Report.pdf Cited by: §2.

Wan et al. (2025) T. Wan, A. Wang, B. Ai, B. Wen, C. Mao, C. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, J. Zeng, J. Wang, J. Zhang, J. Zhou, J. Wang, J. Chen, K. Zhu, K. Zhao, K. Yan, L. Huang, M. Feng, N. Zhang, P. Li, P. Wu, R. Chu, R. Feng, S. Zhang, S. Sun, T. Fang, T. Wang, T. Gui, T. Weng, T. Shen, W. Lin, W. Wang, W. Wang, W. Zhou, W. Wang, W. Shen, W. Yu, X. Shi, X. Huang, X. Xu, Y. Kou, Y. Lv, Y. Li, Y. Liu, Y. Wang, Y. Zhang, Y. Huang, Y. Li, Y. Wu, Y. Liu, Y. Pan, Y. Zheng, Y. Hong, Y. Shi, Y. Feng, Z. Jiang, Z. Han, Z. Wu, and Z. Liu Wan: Open and Advanced Large-Scale Video Generative Models. arXiv. External Links: 2503.20314, Document Cited by: §1.

Wang et al. (2025) J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotný VGGT: visual geometry grounded transformer. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025, Nashville, TN, USA, June 11-15, 2025, pp. 5294–5306. External Links: Link, Document Cited by: Appendix A, §3.4.

Wang et al. (2023) X. Wang, Z. Zhu, G. Huang, X. Chen, J. Zhu, and J. Lu DriveDreamer: towards real-world-driven world models for autonomous driving. arXiv preprint arXiv:2309.09777. Cited by: §2.

Wang et al. (2026) Z. Wang, Z. Liu, J. Li, K. Huang, B. Xu, F. Kang, M. An, P. Wang, B. Jiang, Y. Wei, Y. Xietian, J. Pei, L. Hu, B. Jiang, H. Xue, Z. Wang, H. Sun, W. Li, W. Ouyang, X. He, Y. Liu, Y. Li, and Y. Zhou Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory. arXiv. External Links: 2604.08995, Document Cited by: §1, §2, §4.1.

Wu et al. (2026) M. Wu, Z. Cai, F. Zhao, X. Feng, R. Dang, B. Song, R. Tian, J. Zhu, J. Lei, H. Dou, J. Tang, L. Sun, J. Wu, X. Chu, Z. Liu, and K. Huang Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models. arXiv. External Links: 2603.22212, Document Cited by: Appendix A, §1, Table 1, §2, §3.4.

Xu et al. (2026a) T. Xu, J. Sui, Z. Gao, K. Shi, W. Yang, Z. Liu, Z. Sun, M. Sun, H. Pan, F. Jiang, M. Xu, Q. Fan, Y. Gao, Y. Li, and B. Chen WorldRoamBench: an open-world benchmark for long-horizon stability of interactive world models. arXiv preprint arXiv:2606.31672. Cited by: Table 1, §2.

Xu et al. (2026b) X. Xu, Z. Lin, K. He, Y. Feng, X. Mao, Y. Yin, Y. Ge, and K. Zhang WorldMark: a unified benchmark suite for interactive video world models. arXiv preprint arXiv:2604.21686. Cited by: Appendix A, §1, Table 1, §2, §3.4.

Yang et al. (2024) L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao Depth anything v2. In Advances in Neural Information Processing Systems, Vol. 37, pp. 21875–21911. External Links: Document Cited by: Appendix A.

Ye et al. (2026) Y. Ye, X. Lu, Y. Jiang, Y. Gu, R. Zhao, Q. Liang, J. Pan, F. Zhang, W. Wu, and A. J. Wang MIND: Benchmarking Memory Consistency and Action Control in World Models. arXiv. External Links: 2602.08025, Document Cited by: §2.

Ying et al. (2026) K. Ying, H. Hu, S. Ren, J. Li, F. Chen, Z. Wang, X. Cao, X. Cai, and H. Ding WBench: a comprehensive multi-turn benchmark for interactive video world model evaluation. arXiv preprint arXiv:2605.25874. Cited by: Table 1, §2.

Zhu et al. (2024) F. Zhu, H. Wu, S. Guo, Y. Liu, C. Cheang, and T. Kong IRASim: a fine-grained world model for robot manipulation. arXiv preprint arXiv:2406.14540. Cited by: §2.

Zhu et al. (2026a) H. Zhu, H. Liu, Y. Zhao, T. Ye, J. Chen, J. Yu, T. He, S. Han, and E. Xie SANA-WM: efficient minute-scale world modeling with hybrid linear diffusion transformer. External Links: 2605.15178, Document, Link Cited by: §1, §2, §4.1.

Zhu et al. (2026b) Y. Zhu, J. Feng, W. Zheng, Y. Gao, X. Tao, P. Wan, J. Zhou, and J. Lu Astra: General Interactive World Model with Autoregressive Denoising. arXiv. External Links: 2512.08931, Document Cited by: §1.

Appendix A Implementation Details

Closed-loop execution. Every evaluation case provides all models with the same initial frame, long-horizon objective, and human-annotated basic action sequence. The agent interface executes the current action and returns the latest generated frame to the Agent Player, which observes the frame, the remaining basic action sequence, and the recent action history before returning Keep, Stop, Extend, Correct, or End. This closed loop continues until the Agent Player returns End or the 40-step interaction budget is exhausted, producing rollouts of approximately 10–60 seconds. We use Claude Haiku 4.5 (2) as the agent model for all reported benchmark experiments.

Model execution. Genie 3, LingBot-World, LingBot-World2, HY-World2, and HappyOyster are evaluated through their user-facing web interfaces, whereas SANA-WM, Hunyuan-GameCraft-2, HY-WorldPlay, and Matrix-Game-3.0 are evaluated locally through chunk-wise generation and model-specific adapters. All settings use the same action definitions and record the executed actions and generated observations. HY-World2 requires an input image that supports global scene modeling. Cases whose initial images do not satisfy this requirement are skipped and excluded from its score aggregation. Hunyuan-GameCraft-2 requires a non-empty action input for every generated chunk; consequently, insight evolution cases that require prolonged passive observation or WAIT-only evolution cannot always be evaluated under their intended protocol.

Basic ability metrics. Video Quality includes seven higher-is-better metrics. Aesthetic Quality, Imaging Quality, Motion Smoothness, and Temporal Flickering follow VBench (16); Temporal Consistency follows Omni-WorldBench (39); Depth Stability uses Depth Anything V2 (42) following MemoBench (5); and Subject Consistency is adapted from HyDRA’s DSCctx (6). Subject Consistency detects dynamic subjects with YOLO (26), extracts CLIP (25) features from the subject crops, and measures their cross-window similarity. When no corresponding dynamic subject is detected in both windows, Subject Consistency is reported as N/A.

For Action Controllability, VGGT (36) estimates camera poses from uniformly sampled frames following WorldMark (41). The executed action sequence, including Agent Player adjustments, defines the target trajectory. A rollout passes Translation when its target-normalized Translation Error is below 0.3 and passes Rotation when its mean geodesic Rotation Error is below 45∘. Each pass rate is computed only over rollouts with a valid estimate for the corresponding metric. Basic Ability Score ranks models independently on the seven Video Quality metrics and two Action Controllability pass rates, averages the nine ranks, and converts the result into a higher-is-better percentage.

VQA rubric verifier. Gemini 3.1 Pro (11) serves as the VQA rubric verifier. For geometry consistency, interaction fidelity, and out-of-sight evolution, dimension-specific validation first checks whether the rollout follows the trajectory required by the objective. Only valid rollouts proceed to rubric scoring; invalid rollouts receive the minimum dimension score of 1. Insight evolution uses stationary observation and therefore does not require trajectory validation.

To balance temporal coverage and spatial detail, the verifier receives two contact-sheet streams. The primary stream samples the rollout at 10 FPS, resizes each frame to 384×216, and groups every 25 frames into a 5×5 grid. The detail stream samples at 0.5 FPS, resizes each frame to 800×450, and groups every four frames into a 2×2 grid. Gemini receives both streams together with the case objective and its sample-specific VQA rubric. For each applicable question, it returns a binary answer aj∈{0,1} and supporting evidence. Regular questions receive weight 1, the question expressing the defining requirement receives weight 2, and N/A questions receive weight 0. The dimension score is

Rd=∑j∈Adwj​aj∑j∈Adwj,Sd=1+4​Rd,

where Ad contains only applicable questions. The Overall score is the unweighted mean of geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution.

Radar-chart scales. In Fig. 1, the four rubric-based dimensions—geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution—are displayed on a range of 1–3, while Basic Ability is displayed on a range of 0–1. The axes therefore visualize performance within their respective ranges and should not be interpreted as sharing the same absolute scale.

Appendix B VQA Scoring Robustness

Gemini’s outputs may vary across repeated calls even when the prompt and visual inputs are fixed. Repeating rubric verification for every rollout is costly at the full benchmark scale, so benchmark evaluations commonly rely on a single scoring pass under a fixed configuration. We likewise use one VQA scoring pass per rollout for the reported results and conduct an additional independent pass to assess robustness. Across the nine models, the mean per-model sample variance between the two scoring passes is 0.0112. This low variance indicates that a single fixed scoring pass provides stable aggregate results, although the two-pass analysis is not a precise estimate of scoring uncertainty. When the evaluation budget permits, we recommend averaging multiple independent VQA scoring passes to further reduce verifier variance.

Appendix C Agent Interface for Web-Based Models

Web-based world models do not expose a common local inference API. The agent interface therefore connects to an authenticated browser session, initializes the evaluation case, dispatches model-specific controls, monitors generation, and captures the resulting observation. The captured frame is returned to the Agent Player, and the interface executes the next control after Keep, Stop, Extend, or Correct, or terminates the episode after End. For reproducibility, it records the task identifier, initial frame, basic action sequence, executed controls, intermediate frames, Agent Player decisions, timing, termination condition, and final video.

Browser interaction. For each case, the agent interface opens the model interface, supplies the initial-world condition, selects the required perspective when this option is available, and waits until the generated world becomes interactive. During execution, controls are dispatched through the browser interface and the rendered world is captured after each step. The Agent Player decides only how the world should be controlled, while login, page navigation, submission, rendering-state checks, and output collection remain deterministic interface operations.

Example execution. Fig. 7 shows one HappyOyster case from caption entry to termination. The screenshots are direct crops from the recorded browser session; browser chrome, system controls, and assistant overlays are excluded without altering the generated world content.

Figure 7: Agent interface execution on a web-based world model. (a) The agent interface enters the initial-world condition; (b) waits until the generated world becomes interactive; (c)–(d) dispatches controls and captures observations for the Agent Player’s online decisions; and (e) terminates the episode while preserving the final video and execution record.

Appendix D Human Annotation and Validation

Benchmark annotation. Human annotators screen candidate initial worlds, verify Gemini-generated captions, assign the viewpoint, write long-horizon objectives and basic action sequences, and construct sample-specific VQA rubrics. Annotators use observable evidence and remove questions whose answers cannot be determined from the generated rollout.

Human-alignment study. Five participants independently evaluate the same 120 video pairs, comprising 30 comparisons for each of geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution. This yields 600 valid pairwise judgments, with no missing responses. For each comparison, participants view two videos generated for the same evaluation case, with model identities concealed using anonymous codes. They select the left video, the right video, or Tie when the two are difficult to distinguish.

Human preference assigns full credit to a win and half credit to a tie. Overall human preference pools comparisons across the four equally sampled dimensions. For visualization, we normalize the preference rates within each dimension so that the scores of the nine evaluated models sum to 100%.

媒体内容 · 前往原文查看

Table 7: Inter-rater agreement in the human-alignment study. All results are computed from the same 120 pairwise items and 600 judgments used in Fig. 6. Unanimous Agreement requires all five raters to select the same outcome, whereas Majority Agreement requires at least three raters to agree.

Dimension Items Judgments Unanimous Agreement Majority Agreement Fleiss’ κ

Geometry Consistency 30 150 33.3% 96.7% 0.461

Interaction Fidelity 30 150 16.7% 96.7% 0.323

Out-of-sight Evolution 30 150 23.3% 100.0% 0.419

Insight Evolution 30 150 43.3% 90.0% 0.483

Overall 120 600 29.2% 95.8% 0.434

At least three of the five raters agree on 95.8% of all pairwise items. Fleiss’ κ is 0.434 overall and remains positive across all four dimensions, supporting the reliability of the aggregated human preference rankings. Figure 6 compares the normalized human preference scores with the corresponding rubric-based VQA scores. The Spearman correlations are ρ=0.933 for Overall, ρ=0.983 for geometry consistency, ρ=0.933 for interaction fidelity, ρ=0.812 for out-of-sight evolution, and ρ=0.745 for insight evolution; all are statistically significant (p<0.05).
