QWEN CHAT 我们正式发布通义千问图像生成系列第三代基础模型——Qwen-Image-3.0。如果说 Qwen-Image-1.0 的关键词是"精准",Qwen-Image-2.0 的关键词是"精准、多样、完整、美观、真实",那么 Qwen-Image-3.0 的核心可以归结为一个字——"实"。
这个"实"体现在三个维度:
- 内容丰富:支持高达 4.5k 模型 token 输入,轻松生成报纸、分镜脚本、试卷等复杂版式。
- 细节真实:支持精确渲染小至 10px 的文字,生动再现毛孔、发丝等细节,呈现逼真的微观级刻画。
- 知识深厚:原生支持 12 种语言渲染,模拟网页、游戏、直播等主流界面,并依托丰富的世界知识进行创作。
总之,Qwen-Image-3.0 追求的不仅是"好看",更是"好用",让图像生成成为真正可落地的生产力工具。
内容丰富#
先来看下面这张由 Qwen-Image-3.0 生成的图片:
可以看到,Qwen-Image-3.0 能够准确渲染一张数学幻灯片,包括空间关系、数学符号、定理描述等丰富的视觉内容。这些内容布局合理,相对位置恰当,信息量十足。
然而,这并非 Qwen-Image-3.0 的真正实力。实际上,这仅仅是其真实能力的"九分之一",因为这张图只是 Qwen-Image-3.0 生成的复杂 3×3 网格中的一个单元格。我们来看原始图片:
没错——上面整张图片是由 Qwen-Image-3.0 一次性生成的,而非多张图片拼接而成。这张图片的难点在于,每个单元格都是一张复杂的信息图;要精确描述整个 3×3 网格,需要整整 3.7k 模型 token。
这 3.7k 个模型 token 必须完整描绘一幅隧道安全漫画、一堂空间几何课、一篇《出师表》的风格分析、物理抛体运动、生物学寄生虫学讲解、一张右侧胸痛的医学示意图、群论的西罗定理、银行内控管理信息图,以及细胞 DNA 结构对比——每个单元格都包含精确的中英文文本、公式、图表、卡通角色等内容。
然而这对 Qwen-Image-3.0 来说依然毫不费力——因为 Qwen-Image-3.0 将可接受的指令长度提升到了 4.5k 个模型 token,这意味着模型能够理解并渲染极其复杂、信息密集的视觉布局。
这就是我们所说的“丰富内容”的一个重要特征:内容可以横向扩展。横向扩展体现了模型在语义并列与空间控制方面的能力——即在一张图片中有序地排布多个概念,并使它们互不干扰地呈现出来。
除了横向扩展,深度是“丰富内容”的另一个重要特征。
横向扩展考验的是“一张画布上能放置多少个并行元素”,而深度则考验模型的语义解构与逻辑嵌套能力——能否在一张图片中逐层渲染多个嵌套界面。下面的示例通过一条指令,从外到内依次展示了:一个 VSCode 编程界面 → 一个 Qwen Chat 界面 → 一个微信界面 → 一张手冲咖啡海报。每一层都保留了各自界面的真实风格与细节,形成了“画中画中画”的视觉深度。
以上两个示例分别从横向和纵向两个维度说明了“丰富内容”的含义。
真实细节#
如果说“丰富内容”解决的是“画多少”的问题,那么“真实细节”解决的则是“画多细”的问题。Qwen-Image-3.0 在微观细节的渲染精度上达到了新高度:10 像素的小字清晰可辨,毛孔和发丝得到精细呈现,皮肤质感接近照片级真实感。我们先从精细小字的渲染说起。
下图是一张关于鲸鲨的知识信息图,包含大量文字和插图。Qwen-Image-3.0 能够准确渲染每一个区域。
学术论文是对小字渲染的终极考验——密集的 LaTeX 公式、上下标、希腊字母和定理编号,任何一个符号都不能出错。
该模型渲染了一整页代数几何领域的学术论文,包含多行复杂的公式推导。上标、下标、花括号、分数线以及多行对齐等 LaTeX 排版元素均被准确呈现,即使在小字号下也保持了极佳的可读性。
Qwen-Image-3.0 还能在逼真的纸张上生成精细文字。下例是一份由 Qwen-Image-3.0 生成的报纸,模型不仅准确生成了密集的文字,还模拟了报纸的真实外观。
在编辑任务中,我们同样可以生成精细文字。例如在下述案例中,模型生成了具有真实风格的批注。


模型在书页上叠加了逼真的红色手写批注——下划线、波浪线、圆圈、箭头和简短评语——笔迹自然流畅,完美模拟了高中生课堂笔记的风格。
除了文字和版面的精细还原,“真实细节”在纹理刻画方面同样表现突出。以下是两幅人像摄影示例,模型捕捉到了极为细腻的质感。

除了人像,模型还能描绘其他物体的细腻纹理。



在编辑任务中,我们同样可以生成细节丰富的图像。




对于受损或不完整的传统绘画,该模型能够修复缺失部分,同时忠实保留原作的艺术风格与笔触。


该模型完成了这幅鹰战图的修复,笔触与原作保持一致,保留了水墨渐变、羽毛纹理和构图平衡,同时去除了霉斑和破损痕迹。
深度知识#
“丰富内容”回答的是“能画得多复杂”,“真实细节”回答的是“能画得多逼真”,而“深度知识”回答的是“能画得多广泛”。Qwen-Image-3.0 具备覆盖 12 种语言、多种字体、100 多种艺术风格以及各类 UI 界面的渲染能力——这一切都依托于模型对世界知识的深度理解。
在以下三个示例中,模型分别准确渲染了日语、韩语和西班牙语。

除了能准确渲染多种语言,该模型还拥有丰富的世界知识。特别是,模型可以生成各种逼真的 UI 界面。
我们还可以利用模型强大的世界知识来创建信息图。在下面的示例中,我们基于一张真实图片生成了一个复杂的信息图。


在保留原始昆虫照片主要主体的同时,模型添加了分类信息、形态学标注、放大细节图和比例尺等专业元素,生成了一张可直接用于学术发表的研究用图。
除了模型自身已具备的世界知识,它还可以连接互联网检索最新的世界知识。例如,我们可以要求模型生成一张杭州 7 月 21 日的天气预报图。
模型还能找到特定的 IP 形象并据此进行创作。例如,我们可以生成一张齐白石和梵高在直播间介绍 Qwen-Image-3.0 的图片。

结论#
从 Qwen-Image-1.0 的“精准”,到 Qwen-Image-2.0 的“精准、多样、完整、美观、真实”,再到如今 Qwen-Image-3.0 的“真实”——我们始终追求的目标,是让图像生成从“可用”走向“实用”,从“好看”走向“好用”。
在“内容丰富、细节真实、知识深入”这三大核心能力的支撑下,Qwen-Image-3.0 在报纸 PDF、短剧分镜、复杂 UI 界面等高价值生产力场景中取得了重大突破。我们相信,随着图像生成模型能力的持续提升,它们将在设计、内容创作、教育和电商等更多领域释放真正的生产力价值。
以上就是本次更新的主要亮点。希望您能喜欢使用 Qwen-Image-3.0!
QWEN CHAT We are launching Qwen-Image-3.0, the third-generation foundational image generation model in the Qwen-Image series. If the keyword for Qwen-Image-1.0 was “Precision”, and the keywords for Qwen-Image-2.0 were “Precision, Variety, Completeness, Beauty, and Authenticity”, then the core of Qwen-Image-3.0 comes down to a single word — “Real” (实).
This “Real” is embodied across three dimensions:
- Rich Content: Supports up to 4.5k token input, effortlessly generating complex layouts such as newspapers, storyboards, and exam papers.
- Authentic Details: Supports precise rendering of text as small as 10px, vividly reproducing details like pores and hair strands with lifelike, micro-level depiction.
- Deep Knowledge: Supports native rendering of 12 languages, simulates mainstream interfaces such as web pages, games, and livestreams, and draws on rich world knowledge.
In a word, Qwen-Image-3.0 is not just pursuing “good-looking” — it is pursuing “useful”, making image generation a truly deployable productivity tool.
Rich Content#
Let’s start with the image below, generated by Qwen-Image-3.0:
As you can see, Qwen-Image-3.0 can accurately render a math slide, including spatial relationships, mathematical symbols, theorem descriptions, and other rich visual content. This content is laid out reasonably, with proper relative positioning, and looks rich in information.
However, this is not the true strength of Qwen-Image-3.0. In fact, this is only “1/9” of its real capability, because this image is actually just one cell of a complex 3×3 grid generated by Qwen-Image-3.0. Let’s look at the original image:
That’s right — the entire image above was generated by Qwen-Image-3.0 in a single pass, rather than being stitched together from multiple images. The difficulty of this image lies in the fact that each cell is a complex infographic; to precisely describe the full 3×3 grid takes a full 3.7k tokens.
These 3.7k tokens must fully depict a tunnel safety comic, a spatial geometry lesson, a stylistic analysis of “Chu Shi Biao” (Memorial on Dispatching the Troops), physics projectile motion, a biology parasitology explainer, a medical diagram of right-side chest pain, the Sylow theorems of group theory, a bank internal-control management infographic, and a cell DNA structure comparison — each cell containing precise Chinese and English text, formulas, charts, cartoon characters, and more.
And yet this remains effortless for Qwen-Image-3.0 — because Qwen-Image-3.0 raises the acceptable instruction length to 4.5k tokens, which means the model can understand and render extremely complex, information-dense visual layouts.
This is an important characteristic of the “Rich Content” we mentioned: content can expand horizontally. Horizontal expansion reflects the model’s strength in semantic juxtaposition and spatial control — the ability to lay out multiple concepts in an orderly fashion within a single image and render them without mutual interference.
Beyond horizontal expansion, depth is another important characteristic of Rich Content.
Horizontal expansion tests “how many parallel elements can be placed on a single canvas,” while depth tests the model’s semantic deconstruction and logical nesting — whether it can render multiple nested interfaces layer by layer within a single image. The following example uses a single instruction to display, from outer to inner: a VSCode programming interface → a Qwen Chat interface → a Wechat interface → a pour-over coffee poster. Each layer preserves the authentic style and details of its respective UI, forming a “picture-in-picture-in-picture” visual depth.
The two examples above illustrate the meaning of “Rich Content” along both the horizontal and vertical dimensions.
Authentic Details#
If “Rich Content” addresses the question of “how much to draw,” then “Authentic Details” addresses the question of “how finely to draw.” Qwen-Image-3.0 reaches a new height in the rendering precision of micro-level details: 10px small text is clearly legible, pores and hair strands are rendered in fine detail, and skin texture approaches photographic realism. Let’s start with the precise rendering of small text.
Below is a knowledge infographic about whale sharks, containing a large amount of text and illustrations. Qwen-Image-3.0 is able to accurately render every region.
Academic papers are the ultimate stress test for small-text rendering — dense LaTeX formulas, subscripts and superscripts, Greek letters, and theorem numbering, where not a single symbol can go wrong.
The model renders a full page of an academic paper in the field of algebraic geometry, including multiple lines of complex formula derivations. LaTeX typesetting elements such as superscripts, subscripts, curly braces, fraction bars, and multi-line alignment are all accurately presented, maintaining excellent readability even at small font sizes.
Qwen-Image-3.0 can also generate fine text on realistic paper. The example below is a newspaper generated by Qwen-Image-3.0, in which the model not only accurately generates dense text but also simulates the authentic look of a newspaper.
In editing tasks, we can also generate fine text. For example, in the case below, the model produces annotations with a realistic style.


The model overlays realistic red handwritten annotations onto the book page — underlines, wavy lines, circles, arrows, and short comments — with natural, fluent handwriting that perfectly simulates the style of a high school student’s class notes.
Beyond the fine reproduction of text and layout, “Authentic Details” also stands out in texture depiction. Below are two portrait photography examples in which the model captures extremely delicate textures.

Beyond portraits, the model can also depict the delicate textures of other objects.



In editing tasks, we can also generate images with rich details.




Given a damaged or incomplete traditional painting, the model can restore the missing parts while faithfully maintaining the original artistic style and brushwork.


The model completes the restoration of the eagle-combat painting with brushwork consistent with the original, preserving the ink-wash gradients, feather texture, and compositional balance while removing mold spots and signs of damage.
Deep Knowledge#
“Rich Content” answers “how complex can it draw,” “Authentic Details” answers “how lifelike can it draw,” and “Deep Knowledge” answers “how broadly can it draw.” Qwen-Image-3.0 possesses rendering capabilities covering 12 languages, multiple fonts, 100+ artistic styles, and a variety of UI interfaces — all backed by the model’s deep understanding of world knowledge.
In the three examples below, the model accurately renders Japanese, Korean, and Spanish respectively. 

Beyond accurate rendering of multiple languages, the model also possesses rich world knowledge. In particular, the model can generate various realistic UI interfaces. 
We can also leverage the model’s powerful world knowledge to create infographics. In the example below, we generate a complex infographic based on a real image.


While preserving the main subject of the original insect photograph, the model adds professional elements such as taxonomic information, morphological annotations, magnified detail views, and a scale bar, producing a research figure ready for direct use in academic publication.
In addition to the world knowledge the model already possesses, it can also connect to the internet to retrieve the latest world knowledge. For example, we can ask the model to generate a weather forecast image for Hangzhou on July 21.
The model can also find specific IP figures and create based on them. For instance, we can generate an image of Qi Baishi and Van Gogh introducing Qwen-Image-3.0 in a livestream room.

Conclusion#
From the “Precision” of Qwen-Image-1.0, to the “Precision, Variety, Completeness, Beauty, and Authenticity” of Qwen-Image-2.0, and now to the “Real” of Qwen-Image-3.0 — the goal we have always pursued is to move image generation from “usable” to “practical,” and from “good-looking” to “useful.”
Supported by its three core features — “Rich Content, Authentic Details, and Deep Knowledge” — Qwen-Image-3.0 achieves significant breakthroughs in high-value productivity scenarios such as newspaper PDFs, short-drama storyboards, and complex UI interfaces. We believe that as the capabilities of image generation models continue to improve, they will unlock genuine productivity value in even more fields, including design, content creation, education, and e-commerce.
That concludes the main highlights of this update. We hope you enjoy using Qwen-Image-3.0!