# V-RAE 将视觉基础模型表征用于视频生成

- 来源：Saining Xie (@sainingxie)
- 发布时间：2026-08-27 05:08
- AIHOT 分数：33
- AIHOT 链接：https://aihot.virxact.com/items/cmtalznil01qrroamntkdegou
- 原文链接：https://x.com/sainingxie/status/2092720774268301417

## AI 摘要

谢赛宁转发介绍 V-RAE，该模型直接采用冻结视觉基础模型（DINOv3、SigLIP2、EUPE、V-JEPA 2.1）的表征作为视频生成潜空间，而非传统 VAE 潜空间。在匹配设置下，V-RAE 在 Kinetics-600 上达 2.13 rFVD，UCF101 上 117.86 gFVD，收敛速度比 VAE 潜空间快 6 倍，并引入 tFVD 评估时序平滑度。

## 正文

great to see RAE extending to video!

### 引用推文

> Minghui Guo：🚀 What if video generators could build on representations that already understand the visual world? We are excited to introduce V-RAE: Rethinking Video Latent ...
