# CPI-Bench：面向真实世界图像编辑的综合、实用与智能基准

- 来源：HuggingFace Daily Papers（社区热门论文）
- 发布时间：2026-08-14 08:00
- AIHOT 分数：49
- AIHOT 链接：https://aihot.virxact.com/items/cmswpn90p03juroztdmrludav
- 原文链接：https://arxiv.org/abs/2608.14546

## AI 摘要

CPI-Bench 提出一个面向真实世界图像编辑的综合、实用与智能基准，包含三个子集：覆盖多图像编辑的 CPI-General-Bench、聚焦高频真实用户场景的 CPI-Practical-Bench，以及评估高要求推理编辑的 CPI-Intelligent-Bench。

## 正文

With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, we propose CPI-Bench, a Comprehensive, Practical andIntelligent benchmark for real-world image editing. CPI-Bench comprises three core subsets: CPI-General-Bench, which comprehensively covers diverse editing tasks and pioneers the inclusion of multi-image editing evaluation; CPI-Practical-Bench, which focuses on high-frequency real-user application scenarios; and CPI-Intelligent-Bench, which is dedicated to evaluating capabilities in highly demanding reasoning-based editing. Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models. It provides a comprehensive and reliable quantification of gaps in general editing capabilities, practical deployment efficacy, and advanced reasoning-based editing, offering invaluable guidance for the future optimization of image editing models. Crucially, our ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating it faithfully captures the preferences and perceptual judgments of human evaluators, serving as a robust proxy for real-world user experience.
