面向通用图像生成的能力中心化数据设计:从语料库到协同演进能力

HuggingFace Daily Papers(社区热门论文)·2026-08-18 08:00·6天前
AI 导读

一项研究提出能力驱动的数据基础设施,将能力专属监督构建与能力对齐的课程调度相结合,通过三个数据引擎分别构建文本-图像 grounding、图像间变换和图像-知识关联的互补关系监督。该框架构建了4.4亿图像T2I语料库、1.2亿编辑对和超2700万图像-实体对,并据此从零训练了3B和6B两种规模的多模态扩散模型,在CPI-Bench上进行了定量评估。

HuggingFace Daily Papers(社区热门论文)
50AI 编辑部评分,满分 100

面向通用图像生成的能力中心化数据设计:从语料库到协同演进能力

2026-08-18 08:00· 6天前
AI 导读

一项研究提出能力驱动的数据基础设施,将能力专属监督构建与能力对齐的课程调度相结合,通过三个数据引擎分别构建文本-图像 grounding、图像间变换和图像-知识关联的互补关系监督。该框架构建了4.4亿图像T2I语料库、1.2亿编辑对和超2700万图像-实体对,并据此从零训练了3B和6B两种规模的多模态扩散模型,在CPI-Bench上进行了定量评估。

Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a capability-driven data infrastructure that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org