SenseTime开源了基于NEO-Unify架构的多模态图像生成模型SenseNova-U1。该架构完全摒弃了传统视觉编码器和VAE,原生地将理解、推理和生成统一为一个系统。该系列模型(8B和A3B参数)在开源模型中效率领先,以紧凑尺寸提供商业级性能与出色成本效益。其特色功能包括原生生成图文交织内容,适用于制作指南等实用场景;并擅长高密度信息渲染,能生成知识插图、海报、PPT和漫画等丰富结构的布局。模型已在Hugging Face和GitHub等平台开源。
SenseTime open-sourced SenseNova-U1, a multimodal image generation model built on NEO-Unify!
This architecture drops the visual encoder and VAE entirely. It generates images natively as one system that can handle understanding, reasoning, and generation processes.
@SenseTime_AI 🤖