OpenRouter 推出图像基准测试,覆盖全部 39 款图像模型

OpenRouter:Announcements(RSS)·2026-08-21 08:00·3天前·Brian Thomas
AI 导读

OpenRouter 发布 Visual Image Benchmarks,帮助用户快速评估其提供的全部 39 款图像模型(截至 2026 年 8 月)的实际能力。基准测试用七类挑战性提示词区分模型差异,包括不可能场景、计数、文本、空间关系、否定指令、编辑和一致性,并以网格形式展示结果,支持按价格和生成时间排序。OpenRouter 计划未来将该工具扩展至视频和音频模态。

OpenRouter:Announcements(RSS)
55AI 编辑部评分,满分 100

OpenRouter 推出图像基准测试,覆盖全部 39 款图像模型

2026-08-21 08:00· 3天前· Brian Thomas
AI 导读

OpenRouter 发布 Visual Image Benchmarks,帮助用户快速评估其提供的全部 39 款图像模型(截至 2026 年 8 月)的实际能力。基准测试用七类挑战性提示词区分模型差异,包括不可能场景、计数、文本、空间关系、否定指令、编辑和一致性,并以网格形式展示结果,支持按价格和生成时间排序。OpenRouter 计划未来将该工具扩展至视频和音频模态。

Image Benchmarks: See the Capabilities of Every Model

Unlike choosing a text model, where we have a vast array of LLM benchmarks, picking an image model can feel arbitrary. Output samples tend to be curated eye candy and LLM-as-a-judge evals can’t yet capture the details a human would notice instantly. While arena scores help, they evaluate which outputs people prefer rather than what a model can actually do.

Today we’re launching Visual Image Benchmarks to help you quickly evaluate the capabilities of all the image models we offer (39 as of August ‘26). We’ve selected a set of challenging prompts designed to differentiate the capabilities of models, and show every result in a grid with sorting for both price and generation time.

视频 · 前往原文观看

Prompts designed to test the boundaries of image model capabilities

Each challenge is designed to differentiate the capabilities of image models. We’ve initially grouped them into seven families:

  • Improbable scenes. A wine glass filled level with the rim, umbrellas that are closed. Training data is full of the ordinary version of both.
  • Counting. Three fingers, specific numbers of cards and dice.
  • Text. One long exact string on a poster, and several languages in the same frame.
  • Spatial relations. Occlusion and mirror reflections.
  • Negation. A zebra with no stripes, a Times Square with no advertising.
  • Editing. Minimal diffs, object removal, person removal, all from a reference image.
  • Consistency. Holding a product or four reference subjects steady across a new scene.

The prompts are written so you can instantly evaluate them visually. For example, can a model follow the instruction to fully fill a wine glass?

Full Wine Glass challenge sorted by cost, with the prompt above the grid and each wine glass output labeled with model name, price, and generation time

More visual evals and additional modalities coming soon

It’s similarly challenging to evaluate video and audio models to understand their quality and capabilities. We intend to expand this tool out across modalities as well as keep it up to date with all the new image models we add.

Check out our image benchmarks today! If you have any prompts that could challenge the next round of models, share it with us in #feedback on our Discord.

Generating images via the OpenRouter API and Chat

Once you’ve evaluated models via these benchmarks, try them on your own prompts through the image generation API or Chat to see how they perform on your own content.

来源:OpenRouter:Announcements(RSS)· openrouter.ai