Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor (f), Disentangle (g), Apply (oplus), and Compose (C). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement (g) and attribute binding (oplus) rather than scene-level composition (C), with even the best model scoring only 0.74 on attribute fidelity. Project page: https://amuseum-whr.github.io/TraceBench
TRACE-Bench:分解与诊断多参考图像生成
AI 导读
TRACE-Bench 将多参考图像生成任务形式化为 Anchor、Disentangle、Apply、Compose 四种原子操作,并据此构建约 1,600 个评测案例(覆盖 1–8 个操作槽位,源自 631 个公式模板与约 4,000 张参考图)。对 9 个领先模型的评测显示,主要瓶颈在于解耦与属性绑定而非场景级组合,最佳模型属性保真度得分仅 0.74。
HuggingFace Daily Papers(社区热门论文)
54
AI 编辑部评分,满分 100TRACE-Bench:分解与诊断多参考图像生成
TRACE-Bench 将多参考图像生成任务形式化为 Anchor、Disentangle、Apply、Compose 四种原子操作,并据此构建约 1,600 个评测案例(覆盖 1–8 个操作槽位,源自 631 个公式模板与约 4,000 张参考图)。对 9 个领先模型的评测显示,主要瓶颈在于解耦与属性绑定而非场景级组合,最佳模型属性保真度得分仅 0.74。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org