Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world-model variant converts heterogeneous navigation inputs into calibrated metric 6-DoF camera trajectories and injects them through a geometry-aware conditioning pathway, enabling controller-agnostic interaction across flexible viewpoints. To support efficient long-horizon generation, we transform a bidirectional audio-visual backbone into a causal few-step generator using progressive teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI-Echo-1.5 achieves improvements over existing long-video baselines in cross-shot consistency, visual quality, text alignment, and speech fidelity. Its world-model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long-horizon persistence on SANA-WM-Bench. Together, these results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.
JoyAI-Echo-1.5:面向持久故事与交互世界的长时程音视频生成
AI 导读
JoyAI-Echo-1.5 发布,这是一个统一的音视频生成系统,包含长视频与世界模型两个变体。长视频变体通过可组合的跨镜头记忆与说话人线索,实现跨镜头的角色外观和语音身份保持;世界模型变体将导航输入转换为 6-DoF 相机轨迹,支持控制器无关的交互。其世界模型变体在 WBench 上以 81.7 的平均分排名第一。
HuggingFace Daily Papers(社区热门论文)
58
AI 编辑部评分,满分 100JoyAI-Echo-1.5:面向持久故事与交互世界的长时程音视频生成
JoyAI-Echo-1.5 发布,这是一个统一的音视频生成系统,包含长视频与世界模型两个变体。长视频变体通过可组合的跨镜头记忆与说话人线索,实现跨镜头的角色外观和语音身份保持;世界模型变体将导航输入转换为 6-DoF 相机轨迹,支持控制器无关的交互。其世界模型变体在 WBench 上以 81.7 的平均分排名第一。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org