# MPIE-Bench：多人物交互编辑的解剖合理性评测基准

- 来源：HuggingFace Daily Papers（社区热门论文）
- 发布时间：2026-07-30 08:00
- AIHOT 分数：46
- AIHOT 链接：https://aihot.virxact.com/items/cms8epxbs0aukrot06jof8d4d
- 原文链接：https://arxiv.org/abs/2607.27616

## AI 摘要

MPIE-Bench 是一个包含 2,500 个样本的基准，覆盖 405 个场景、14 类交互和四种接触密度（C0-C3），用于评测多人物交互编辑中的解剖与几何问题。配套的 MPIE-Eval 通过冻结的公开多人物网格重建，从解剖完整性和交互接触几何两个新维度打分。在十个编辑模型上，网格解剖得分最高 0.65、交互得分最高 0.72，而 VLM 清单评分超过 0.95，显示现有评测存在明显饱和。

## 正文

Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.
