Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as an addressed token, predicts a structured identity--layout plan, and renders the plan as a visual condition. Its key objective, Layout-Grounded ID Loss, uses annotated face regions to supervise the intended identities directly, avoiding unstable embedding-based face matching; ID Representation Forcing additionally trains a prediction for each identity before image synthesis. On an identity-disjoint benchmark, WithEveryone achieves the highest target-context identity similarity, improving face similarity from 0.462 for GPT-Image-2 to 0.499, while reducing copy-paste artifacts from 0.169 to 0.055. It further covers 97.3% of the requested identities with a duplicate rate of only 2.8%. These results show that explicit identity--layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference-face copying.
WithEveryone:面向群像生成的统一规划与身份锚定框架
AI 导读
WithEveryone 提出统一框架,支持生成包含至多十个参考身份的群像。其核心目标 Layout-Grounded ID Loss 利用标注人脸区域直接监督目标身份,避免不稳定的嵌入匹配;在身份不相交基准上,目标上下文身份相似度从 GPT-Image-2 的 0.462 提升至 0.499,复制粘贴伪影从 0.169 降至 0.055,身份覆盖率达 97.3%,重复率仅 2.8%。
HuggingFace Daily Papers(社区热门论文)
51
AI 编辑部评分,满分 100WithEveryone:面向群像生成的统一规划与身份锚定框架
WithEveryone 提出统一框架,支持生成包含至多十个参考身份的群像。其核心目标 Layout-Grounded ID Loss 利用标注人脸区域直接监督目标身份,避免不稳定的嵌入匹配;在身份不相交基准上,目标上下文身份相似度从 GPT-Image-2 的 0.462 提升至 0.499,复制粘贴伪影从 0.169 降至 0.055,身份覆盖率达 97.3%,重复率仅 2.8%。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org