Artificial Analysis 更新 Image Editing Arena 榜单,MAI-Image-2.6-Preview 居首

Artificial Analysis · @ArtificialAnlys · X·2026-09-03 04:20·5小时前
AI 导读

Artificial Analysis 更新其 Image Editing Arena,新增 Identity-Preserving Edits、UI/UX Design 等编辑任务测试范围,按 7 类编辑动作和 10 个真实用例基于人类偏好排名,新榜单已上线并开放投票。

Artificial Analysis@ArtificialAnlys
52AI 编辑部评分,满分 100

Artificial Analysis 更新 Image Editing Arena 榜单,MAI-Image-2.6-Preview 居首

2026-09-03 04:20· 5小时前
AI 导读

Artificial Analysis 更新其 Image Editing Arena,新增 Identity-Preserving Edits、UI/UX Design 等编辑任务测试范围,按 7 类编辑动作和 10 个真实用例基于人类偏好排名,新榜单已上线并开放投票。

We have updated the Artificial Analysis Image Editing Arena to expand the range of editing tasks we test for, from Enhancement & Restoration to Identity-Preserving Edits, and from Marketing & Advertising to UI/UX Design. Our updated evaluation measures model performance on complex edit tasks, and assesses not just which model is best overall, but which model is best for specific editing needs. The new leaderboard is live and voting is open.

Image Editing models are advancing fast. AI now fits into different stages of image creation workflows, from generation through to post-production, and editing is no longer a side feature. We treat Image Editing and Reference to Image as separate benchmarks because they sit at different points in that workflow: Image Editing covers post-production, changing an image you already have; while our upcoming Reference to Image benchmark covers generating novel images from reference images.

For Image Editing, we test a model's ability to make specific changes to an image while keeping everything else unchanged. This includes complex edit instructions that chain multiple different asks, as frontier models have largely saturated single-instruction edits.

Different edit requests call for different models. Relighting a cinematic scene is a different problem from reworking the design of a marketing asset. We rank models on human preference across 7 editing actions, such as Object-Level Edit, Identity-Preserving Edit, and Enhancement & Restoration, and 10 real-world use cases, such as Marketing & Advertising, UI/UX Design, and Live-Action Film. The overall benchmark samples evenly across both.

Initial insights from an in-depth analysis of the 10 highest ranking models on the Artificial Analysis Image Editing Leaderboard:

➤ MAI-Image-2.6-Preview leads the overall leaderboard and 4 of the 7 editing action boards: Scene & Style Edit, Text or Symbol Edits, Reasoning-Based Edit, and Enhancement & Restoration, where it is tied #1 with MAI-Image-2.5. It excels at restyling, relighting, and retouching images.

➤ GPT Image 2 (high) ranks #2 overall but #1 on Object-Level Edit and Composition & Framing. It is the strongest at precise local edits and spatial reframing. It is weaker at whole-image transformations that must keep the image's content intact, ranking #7 on Scene & Style Edit and #6 on Enhancement & Restoration.

➤ Seedream 5.0 Pro is the character and identity specialist, ranking #1 on Identity-Preserving Edit.

➤ MAI-Image-2.5-Flash is the value pick of the top 10, at $20 per 1,000 images against $211 for GPT Image 2 (high) and $90 for Seedream 5.0 Pro.

See below for the editing action and use case breakdowns 🧵

来源:Artificial Analysis· x.com