We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Surprisingly, we find that the converged diffusion loss scales with the amount of structured language in the prompt. To quantify structured language, we adapt two complementary measures: a white-box likelihood metric (GPG) and a black-box attribute metric (ED). Across controlled training runs, the converged diffusion loss decreases approximately linearly with GPG and follows a power law with ED. Guided by these scaling properties, we improve \emph{diffusability} by constructing structured prompts with semantic and geometric annotations derived from images, and improve \emph{promptability} by training a prompter through supervised fine-tuning, cold-start, and verifier-gated on-policy distillation. The resulting system outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world-knowledge benchmark, while matching or surpassing the strongest closed-weight models on most evaluations.
视觉生成中文本条件作用的缩放特性研究
AI 导读
研究发现,视觉生成中扩散损失随提示词中结构化语言量呈规律缩放:与GPG近似线性下降,与ED呈幂律关系。基于此,通过构建含语义与几何标注的结构化提示提升可扩散性,并训练提示器提升可提示性。最终系统在几乎所有组合、推理与世界知识基准上超越开源模型,并在多数评测中匹敌或超过最强闭源模型。
HuggingFace Daily Papers(社区热门论文)
53
AI 编辑部评分,满分 100视觉生成中文本条件作用的缩放特性研究
研究发现,视觉生成中扩散损失随提示词中结构化语言量呈规律缩放:与GPG近似线性下降,与ED呈幂律关系。基于此,通过构建含语义与几何标注的结构化提示提升可扩散性,并训练提示器提升可提示性。最终系统在几乎所有组合、推理与世界知识基准上超越开源模型,并在多数评测中匹敌或超过最强闭源模型。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org