# Qwen-Image-3.0 发布：支持 4.5k token 输入与 10px 小字渲染

- 来源：Hacker News 热门（buzzing.cc 中文翻译）
- 作者：ilreb
- 发布时间：2026-07-21 18:38
- AIHOT 分数：67
- AIHOT 链接：https://aihot.virxact.com/items/cmrujbb8y0i98bi07lewhdpe2
- 原文链接：https://qwen.ai/blog?id=qwen-image-3.0

## AI 摘要

阿里通义千问推出第三代基础图像生成模型 Qwen-Image-3.0，核心能力聚焦“真实”，支持最长 4.5k token 的指令输入，可一次性生成包含 9 个复杂信息图的 3×3 网格布局。模型能精确渲染小至 10px 的文字，并原生支持 12 种语言，可模拟网页、游戏、直播等主流界面，在学术论文、报纸等密集排版场景中保持可读性。

## 正文

Qwen-Image-3.0: Rich Content, Authentic Details, Deep Knowledge

QWEN CHAT We are launching Qwen-Image-3.0, the third-generation foundational image generation model in the Qwen-Image series. If the keyword for Qwen-Image-1.0 was “Precision”, and the keywords for Qwen-Image-2.0 were “Precision, Variety, Completeness, Beauty, and Authenticity”, then the core of Qwen-Image-3.0 comes down to a single word — “Real” (实).

This “Real” is embodied across three dimensions:

Rich Content: Supports up to 4.5k token input, effortlessly generating complex layouts such as newspapers, storyboards, and exam papers.

Authentic Details: Supports precise rendering of text as small as 10px, vividly reproducing details like pores and hair strands with lifelike, micro-level depiction.

Deep Knowledge: Supports native rendering of 12 languages, simulates mainstream interfaces such as web pages, games, and livestreams, and draws on rich world knowledge.

In a word, Qwen-Image-3.0 is not just pursuing “good-looking” — it is pursuing “useful”, making image generation a truly deployable productivity tool.

Rich Content#

Let’s start with the image below, generated by Qwen-Image-3.0:

As you can see, Qwen-Image-3.0 can accurately render a math slide, including spatial relationships, mathematical symbols, theorem descriptions, and other rich visual content. This content is laid out reasonably, with proper relative positioning, and looks rich in information.

However, this is not the true strength of Qwen-Image-3.0. In fact, this is only “1/9” of its real capability, because this image is actually just one cell of a complex 3×3 grid generated by Qwen-Image-3.0. Let’s look at the original image:

That’s right — the entire image above was generated by Qwen-Image-3.0 in a single pass, rather than being stitched together from multiple images. The difficulty of this image lies in the fact that each cell is a complex infographic; to precisely describe the full 3×3 grid takes a full 3.7k tokens.

These 3.7k tokens must fully depict a tunnel safety comic, a spatial geometry lesson, a stylistic analysis of “Chu Shi Biao” (Memorial on Dispatching the Troops), physics projectile motion, a biology parasitology explainer, a medical diagram of right-side chest pain, the Sylow theorems of group theory, a bank internal-control management infographic, and a cell DNA structure comparison — each cell containing precise Chinese and English text, formulas, charts, cartoon characters, and more.

And yet this remains effortless for Qwen-Image-3.0 — because Qwen-Image-3.0 raises the acceptable instruction length to 4.5k tokens, which means the model can understand and render extremely complex, information-dense visual layouts.

This is an important characteristic of the “Rich Content” we mentioned: content can expand horizontally. Horizontal expansion reflects the model’s strength in semantic juxtaposition and spatial control — the ability to lay out multiple concepts in an orderly fashion within a single image and render them without mutual interference.

Beyond horizontal expansion, depth is another important characteristic of Rich Content.

Horizontal expansion tests “how many parallel elements can be placed on a single canvas,” while depth tests the model’s semantic deconstruction and logical nesting — whether it can render multiple nested interfaces layer by layer within a single image. The following example uses a single instruction to display, from outer to inner: a VSCode programming interface → a Qwen Chat interface → a Wechat interface → a pour-over coffee poster. Each layer preserves the authentic style and details of its respective UI, forming a “picture-in-picture-in-picture” visual depth.

The two examples above illustrate the meaning of “Rich Content” along both the horizontal and vertical dimensions.

Authentic Details#

If “Rich Content” addresses the question of “how much to draw,” then “Authentic Details” addresses the question of “how finely to draw.” Qwen-Image-3.0 reaches a new height in the rendering precision of micro-level details: 10px small text is clearly legible, pores and hair strands are rendered in fine detail, and skin texture approaches photographic realism. Let’s start with the precise rendering of small text.

Below is a knowledge infographic about whale sharks, containing a large amount of text and illustrations. Qwen-Image-3.0 is able to accurately render every region.

Academic papers are the ultimate stress test for small-text rendering — dense LaTeX formulas, subscripts and superscripts, Greek letters, and theorem numbering, where not a single symbol can go wrong.

The model renders a full page of an academic paper in the field of algebraic geometry, including multiple lines of complex formula derivations. LaTeX typesetting elements such as superscripts, subscripts, curly braces, fraction bars, and multi-line alignment are all accurately presented, maintaining excellent readability even at small font sizes.

Qwen-Image-3.0 can also generate fine text on realistic paper. The example below is a newspaper generated by Qwen-Image-3.0, in which the model not only accurately generates dense text but also simulates the authentic look of a newspaper.

In editing tasks, we can also generate fine text. For example, in the case below, the model produces annotations with a realistic style.

The model overlays realistic red handwritten annotations onto the book page — underlines, wavy lines, circles, arrows, and short comments — with natural, fluent handwriting that perfectly simulates the style of a high school student’s class notes.

Beyond the fine reproduction of text and layout, “Authentic Details” also stands out in texture depiction. Below are two portrait photography examples in which the model captures extremely delicate textures.

Beyond portraits, the model can also depict the delicate textures of other objects.

In editing tasks, we can also generate images with rich details.

Given a damaged or incomplete traditional painting, the model can restore the missing parts while faithfully maintaining the original artistic style and brushwork.

The model completes the restoration of the eagle-combat painting with brushwork consistent with the original, preserving the ink-wash gradients, feather texture, and compositional balance while removing mold spots and signs of damage.

Deep Knowledge#

“Rich Content” answers “how complex can it draw,” “Authentic Details” answers “how lifelike can it draw,” and “Deep Knowledge” answers “how broadly can it draw.” Qwen-Image-3.0 possesses rendering capabilities covering 12 languages, multiple fonts, 100+ artistic styles, and a variety of UI interfaces — all backed by the model’s deep understanding of world knowledge.

In the three examples below, the model accurately renders Japanese, Korean, and Spanish respectively.

Beyond accurate rendering of multiple languages, the model also possesses rich world knowledge. In particular, the model can generate various realistic UI interfaces.

We can also leverage the model’s powerful world knowledge to create infographics. In the example below, we generate a complex infographic based on a real image.

While preserving the main subject of the original insect photograph, the model adds professional elements such as taxonomic information, morphological annotations, magnified detail views, and a scale bar, producing a research figure ready for direct use in academic publication.

In addition to the world knowledge the model already possesses, it can also connect to the internet to retrieve the latest world knowledge. For example, we can ask the model to generate a weather forecast image for Hangzhou on July 21.

The model can also find specific IP figures and create based on them. For instance, we can generate an image of Qi Baishi and Van Gogh introducing Qwen-Image-3.0 in a livestream room.

Conclusion#

From the “Precision” of Qwen-Image-1.0, to the “Precision, Variety, Completeness, Beauty, and Authenticity” of Qwen-Image-2.0, and now to the “Real” of Qwen-Image-3.0 — the goal we have always pursued is to move image generation from “usable” to “practical,” and from “good-looking” to “useful.”

Supported by its three core features — “Rich Content, Authentic Details, and Deep Knowledge” — Qwen-Image-3.0 achieves significant breakthroughs in high-value productivity scenarios such as newspaper PDFs, short-drama storyboards, and complex UI interfaces. We believe that as the capabilities of image generation models continue to improve, they will unlock genuine productivity value in even more fields, including design, content creation, education, and e-commerce.

That concludes the main highlights of this update. We hope you enjoy using Qwen-Image-3.0!
