摘要
在本工作中,我们推出了 Qwen3.5-Omni,这是 Qwen-Omni 模型系列的最新进展。相较于前代模型,Qwen3.5-Omni 实现了重大进化,参数量扩展至数千亿级别,并支持 256k 上下文长度。通过利用包含异构文本-图像对以及超过 1 亿小时音视频内容的海量数据集,该模型展现出强大的全模态能力。Qwen3.5-Omni-Plus 在 215 项音频及音视频理解、推理与交互子任务及基准测试中取得了 SOTA 结果,在关键音频任务上超越 Gemini-3.1 Pro,并在综合音视频理解方面与之持平。在架构上,Qwen3.5-Omni 为 Thinker 和 Talker 均采用了混合注意力专家混合(MoE)框架,实现了高效的长序列推理。该模型支持复杂的交互,能够理解超过 10 小时的音频内容以及 400 秒的 720P 视频(以 1 FPS 处理)。为了解决流式语音合成中因文本与语音 tokenizer 编码效率差异而固有的不稳定和不自然问题,我们引入了 ARIA(自适应速率交错对齐)。ARIA 动态对齐文本与语音单元,在将延迟影响降至最低的同时,显著提升了对话语音的稳定性和韵律。此外,Qwen3.5-Omni 拓展了语言边界,支持 10 种语言的多语言理解与语音生成,并具备类人的情感细腻度。除预设语音外,该模型还能通过用户提供的样本实现零样本语音定制。最后,Qwen3.5-Omni 展现出卓越的音视频定位能力,能够生成具有精确时间同步和自动场景分割的脚本级结构化字幕。值得注意的是,我们观察到全模态模型中出现了一项新能力:直接基于音视频指令进行编码,我们称之为音视频氛围编码(Audio-Visual Vibe Coding)。Qwen3.5-Omni 可通过 API 公开访问。
1 引言
人类与世界的交互本质上是全模态且智能体的,涉及视觉、听觉和语言信息的整合,并通过文本、语音以及目标导向的工具介导行为来产生回应,从而促进与其他生物的信息交换并展现智能。基于大模型在文本(Brown 等人,2020;OpenAI,2023;Gemini Team,2024;Anthropic,2023b;a;2024;Bai 等人,2023a;Yang 等人,2024a;2025a;Touvron 等人,2023;Dubey 等人,2024)、视觉(Li 等人,2023;Liu 等人,2023;Zhu 等人,2023;Bai 等人,2023b;2025a)和音频(Chu 等人,2023;2024)领域的理解与推理能力的快速进步,能够跨所有模态联合处理与生成的原生全模态系统已引起广泛关注(OpenAI,2024;Comanici 等人,2025;Xu 等人,2025a;b)。然而,现有模型主要运行在被动感知-响应范式内,在可扩展的智能体行为、实时交互、自主工具使用以及跨模态推理方面的能力有限,而这些能力正是实际部署所必需的前提条件。
在本报告中,我们介绍 Qwen3.5-Omni,这是通义千问最新一代的全模态大语言模型,支持对文本、图像、音频以及视听内容的理解。Qwen3.5-Omni 以全模态方式在大量文本、视觉数据以及超过 1 亿小时的视听数据上进行了原生预训练,被设计为一个原生全模态智能体模型:它不仅能够跨所有模态进行感知和推理,还能执行操作——自主调用网页搜索、执行复杂函数调用、生成语音输出,并进行实时流式交互。该模型系列包含 Plus 和 Flash 两个变体,均为指令模型,支持 256k token 的长上下文输入。
Qwen3.5-Omni 基于 Qwen2.5-Omni(Xu 等人,2025a)中引入的思考者-说话者架构,并在 Qwen3-Omni(Xu 等人,2025b)的基础上引入了五项关键技术升级:(1)思考者和说话者均采用混合注意力混合专家(MoE)设计,实现了高效推理;(2)支持长达 256k token 的长上下文建模,可处理超过 10 小时的音频以及超过 400 秒、以 1 FPS 采样的 720P 视听内容;(3)在语音生成方面,采用多码本编解码表示,实现了单帧即时合成;(4)说话者引入了 ARIA 技术,该技术在流式解码过程中动态对齐文本和语音单元,显著提升了自然度和鲁棒性;(5)多语言训练大幅扩展,覆盖 113 种语言和方言的语音识别以及 36 种语言的语音合成。
在这些技术进步的推动下,Qwen3.5-Omni 相较于 Qwen3-Omni 带来了三项重大新能力:(1) 可控的音视频描述,能够生成可控、详细且结构化的描述,以及剧本级别的精细描述,包括自动分段、时间戳标注,以及对角色及其与音频关系的详细描述;(2) 全面的实时交互,包括通过原生话轮转换意图识别实现的语义打断、对音量、语速和情感的端到端语音控制,以及基于用户提供的样本进行语音克隆;(3) 原生的全模态智能体行为,包括自主网页搜索、复杂的函数调用调用,以及音视频氛围编程——这是一种涌现能力,模型能够直接从音视频指令生成可执行代码,使模型无需外部编排即可响应实时查询。
关键在于,Qwen3.5-Omni 在文本和视觉模态上保持了最先进的性能,且相对于同等规模的单模型 Qwen 版本没有性能下降。在涵盖音视频基准、音频基准、ASR 基准、特定语言的语音到文本翻译任务以及特定语言 ASR 任务的 215 项音频与音视频理解、推理和交互子任务及基准测试中,Qwen3.5-Omni-Plus 取得了最先进的结果,在通用音频理解、推理、识别、翻译和对话方面超越了 Gemini-3.1 Pro,而其整体音视频理解能力达到了 Gemini-3.1 Pro 的水平。
2 架构
2.1 概览
如图2所示,Qwen3.5-Omni 继续采用 Thinker-Talker 架构(Xu 等人,2025a)。与 Qwen3-Omni(Xu 等人,2025b)相比,Qwen3.5-Omni 在可扩展性、对齐和实时交互方面引入了若干关键改进:
-
整体主干网络采用混合专家模型(MoE)设计,提升了可扩展性,同时在多模态理解和生成任务中更好地平衡了能力与效率。
-
Thinker 分别通过视觉编码器和 AuT 接收视觉和音频信号。音频和视频输入以交错方式处理,以实现统一的多模态建模,并插入显式时间戳以增强时间感知能力,尤其是在处理长视频或音视频混合上下文时。该设计使 Thinker 能够处理扩展输入,支持多达 256k 个 token、10 小时音频或 400 秒每秒一帧的 720P 视频。
-
Talker 负责根据多模态输入以及来自 Thinker 的文本输出,生成上下文相关的语音。Qwen3.5-Omni 采用了 Qwen3-Omni(Xu 等人,2025b)中引入的基于 RVQ 的语音表征,这显著提升了推理效率。
-
为支持实时交互,Qwen3.5-Omni 在 Thinker 中采用了分块流式输入处理,并设计了流式 Talker,从而实现低延迟的端到端多模态对话。
-
与 Qwen3-Omni(Xu 等人,2025b)中双轨 Talker 输入设计不同,Qwen3.5-Omni 中的 Talker 采用 ARIA 在交错文本和语音单元之前动态对齐它们。这种设计缓解了文本与语音之间 token 化速率不匹配所导致的不稳定性,从而减少了跳词、发音错误以及数字渲染模糊等问题。
在接下来的章节中,我们首先介绍 AuT 编码器,包括其训练方法。然后,描述 Thinker 如何处理各种输入。接着,我们详细说明 Talker 的多码本流式语音生成。最后,我们重点介绍在理解和生成模块上的一系列改进,旨在实现超低延迟、端到端的流式音频推理。
2.2 音频 Transformer(AuT)
我们使用从头训练的基于 Transformer 的音频编码器,该编码器采用注意力-编码器-解码器模型 AuT,如图 3 所示。Qwen3.5-Omni 编码器的训练消耗了由 Qwen3-ASR 生成的 4000 万小时音频-文本对数据。音频的滤波器组特征通过 4 个 Conv2D 模块进行 16 倍下采样,然后输入自注意力层,以获得 6.25Hz token 速率的音频 token。与 Qwen3-Omni 编码器的训练过程相比,Qwen3.5-Omni 的编码器采用了超过 20 种语言的更多多语言数据,其中中文、英文和多语言数据的比例达到 3.5 : 3.5 : 3。采用动态注意力窗口大小训练机制,以保证在实时预填充缓存推理和离线音频理解任务下的均衡性能。
2.3 感知
文本、音频、图像和视频(不含音频)。
Thinker 将文本、音频、图像和无声视频输入转换为统一的表征序列。对于文本,我们使用 Qwen3.5 分词器(Team, 2026),该分词器采用字节级字节对编码,词表大小从 15 万扩充至 25 万,在大多数语言上将编码和解码效率提升了 10% 至 60%。对于音频输入(包括从视频中提取的音频),我们将波形重采样至 16 kHz,并使用 25 毫秒窗口和 10 毫秒跳跃长度将其转换为 128 通道的梅尔频谱图。我们使用 AuT 作为音频编码器,该编码器在 4000 万小时音频数据上从头训练,每个输出帧对应约 160 毫秒的原始信号。对于视觉输入,我们采用 Qwen3.5(Team, 2026)的视觉编码器来处理图像和视频。该编码器在图像和视频混合数据上训练,在图像理解和视频理解方面均具备强大能力。为了在保持与音频流对齐的同时尽可能保留视频信息,我们以动态帧率采样视频帧。
音视频时间戳。
遵循 Qwen3-Omni(Xu 等人,2025b)的做法,我们应用 TM-RoPE 赋予模型时间感知能力,以实现音视频同步。然而,我们发现通过时间位置 ID 直接编码绝对时间,会导致来自长视频(含音频输入)的视觉补丁索引过于稀疏,从而削弱长程时间建模能力。此外,这种设计通常需要大规模且在不同帧率下均匀分布的训练样本,增加了数据构建成本。为解决这些问题,我们在每个视频或音视频时间补丁前添加一个显式时间戳,该时间戳以秒为单位的格式化文本字符串表示,使模型能够更自然地学习时间码表征。对于音频序列,我们进一步在随机间隔处插入时间戳,以改善跨模态的时间对齐。尽管这一策略会略微增加上下文长度,但它能实现更精确、更鲁棒的时间感知,尤其是在外推长上下文多模态输入时。
在多模态音视频流处理中,音频组件每 160 毫秒编码一个时间 ID。视频则被视为一系列帧,这些帧带有单调递增的时间 ID,并根据其实际时间戳动态调整,以确保每个 ID 对应 160 毫秒的一致时间分辨率。视频帧的高度和宽度 ID 的分配方式与静态图像相同。为避免处理多种模态时出现位置冲突,位置编号采用连续方式,每种后续模态从前一模态的最大位置 ID 加 1 开始。这种精细化的位置编码方法使模型能够有效整合并联合建模来自不同模态的信息。Qwen3.5-Omni 利用这些表示的时间 ID(这些 ID 明确锚定到绝对时间)来对齐它们。这种设计选择赋予了模型支持任意时长流式输入的灵活性。
2.4 语音生成
Talker 直接对 Qwen3.5-Omni-Audio-Tokenizer 生成的 RVQ token 进行操作。为了对残差码本进行建模,它采用了一个多 token 预测(MTP)模块,该模块能够实现对声学细节的精细建模和控制。结合用于波形重建的因果卷积网络,Talker 能够以低推理延迟和适中的计算开销实现高保真语音合成。
在多轮口语对话中,Talker 以 Thinker 组件提供的丰富上下文信息为条件,包括历史文本 token、多模态表示以及当前轮次的流式文本。这种条件设定使 Talker 能够根据不断变化的对话上下文,动态调节韵律、响度和情感等声学属性。
在架构层面,我们的方法与 Qwen3-Omni(Xu 等人,2025b)有两个关键区别。首先,我们为 Talker 引入了一个专用的系统提示词,用于指定目标语音特征,从而同时实现零样本语音克隆和可控语音生成。与传统的说话人嵌入相比,该提示词能够编码更丰富的多模态线索,包括文本描述和编解码序列,从而对声学实现提供更精细的控制。其次,我们提出了 ARIA(自适应速率交错对齐),它将传统的双通道生成范式统一为单通道形式。ARIA 不依赖于基于 MFA 的对齐或固定的交错速率,而是强制执行一个自适应速率约束:对于生成序列的任意前缀,累积的语音与文本 token 比率不得超过对应的项目级全局比率。尽管设计简单,但这种设计能够实现跨语言(包括编码效率相对较低的语言)的灵活文本-语音对齐,并自然地支持任意文本 token 前缀后接连贯的语音 token 延续。
2.5 流式与并发设计
在流式音视频交互场景中,首包延迟是影响用户体验的关键因素,而模型的并发能力是降低服务成本、提升响应速度的核心。本节讨论 Qwen3.5-Omni 如何通过算法和架构优化来增强并发能力并降低首包延迟。表 1 概述了 Qwen3.5-Omni 的相关架构及其对应的延迟。
| 模块 | 架构 | 流式 |
| 音频编码器 | AuT | |
| 视觉编码器 | SigLIP2 | – |
| 思考者 | 混合 MoE Transformer | |
| 说话者 | 混合 MoE Transformer | |
| MTP | 密集 Transformer | |
| Code2wav | 卷积网络 | |
| 首包延迟(音频输入) | Plus:435 毫秒 Flash:235 毫秒 | |
| 首包延迟(视频输入) | Plus:651 毫秒 Flash:426 毫秒 | |
分块预填充与混合 MoE 架构
在 Qwen3.5-Omni 中,我们保留了 Qwen3-Omni 和 Qwen2.5-Omni 中实现的分块预填充机制,其音频和视觉编码器能够沿时间维度输出数据块。这种方法显著降低了思考者和说话者的首 token 生成时间。在架构上,Qwen3.5-Omni 中的思考者和说话者均基于 Qwen3.5 引入的混合 MoE 架构构建。除了混合 MoE 的通用效率优势外,该架构还包含门控 Delta 网络模块,该模块对于加速长音频-视频序列的建模尤为有效。因此,它显著降低了长上下文推理中的 KV 缓存 I/O 开销,提升了生成吞吐量,并支持更高的服务并发能力。
基于 ARIA 的流式生成。
对于流式语音生成和高并发服务,Qwen3.5-Omni 在很大程度上继承了 Qwen3-Omni 的高效设计:说话者通过轻量级 MTP 模块预测 RVQ 编解码 token,生成的多个码本 token 由因果流式 ConvNet 编解码解码器转换为波形。这些组件在计算上保持轻量、便于批处理,并且非常适合低延迟部署。在此共同基础之上,此前引入的 ARIA 进一步将 Qwen3-Omni 中的双通道生成模式重构为文本和语音 token 上的统一交错单流形式。通过单调交错约束组织文本和语音生成,ARIA 减少了独立生成轨道之间的同步开销,在解码过程中实现了更高效的 token 调度,并更好地匹配了流式交互中自然递增的模式。
在表2中,我们报告了Qwen3.5-Omni在音频和视频输入下不同并发级别的理论首包延迟,该评估在内部vLLM上完成,并对MTP模块和编解码器启用了torch.compile和CUDA Graph加速。其中,Thinker TTFT(首Token时间)表示从接收输入流到Thinker生成第一个文本token的时间,而Talker TTFC(首Chunk时间)衡量的是Talker生成第一个音频chunk所需的时间。TPOP(每输出Token时间)代表稳态解码期间每输出一个token的延迟,其中Talker TPOP包含了Talker主干网络和MTP模块的联合延迟。TPS(每秒Token数)表示生成吞吐量。由于ARIA将文本和语音生成组织成统一的交错流,因此总体延迟不能通过简单累加几行数值获得,而是反映了到第一个可播放音频数据包的端到端关键路径。我们还注意到,由于Qwen3.5-Omni-Flash和Qwen3.5-Omni-Plus之间存在显著的规模差异,这两个变体采用了不同的部署时资源分配和并行化策略;因此,它们的延迟和吞吐量数值不适用于严格横向比较。如表所示,Qwen3.5-Omni在并发增加时保持了稳定的延迟和解码效率,同时较低的生成RTF为流畅的流式音频生成提供了充足的余量。
| Qwen3.5-Omni-Flash | Qwen3.5-Omni-Plus | |||||
| 1 并发 | 4 并发 | 8 并发 | 1 并发 | 4 并发 | 8 并发 | |
| Thinker TTFT | 80/255毫秒 | 86/446毫秒 | 103/765毫秒 | 162/377毫秒 | 183/907毫秒 | 260/1243毫秒 |
| Talker TTFC | 56/61毫秒 | 68/108毫秒 | 81/116毫秒 | 54/56毫秒 | 72/88毫秒 | 95/116毫秒 |
| Thinker TPOP | 5.6/5.9毫秒 | 8.2/9.2毫秒 | 9.6/15.8毫秒 | 17.4/18.5毫秒 | 25.6/26.9毫秒 | 33.3/40.2毫秒 |
| Talker TPOP | 14.2/14.2毫秒 | 16.9/17.0毫秒 | 20.5/20.6毫秒 | 14.9/14.9毫秒 | 21.0/21.3毫秒 | 25.8/27.1毫秒 |
| 编解码器解码 | 3~5毫秒 | |||||
| 总体延迟 | 235/426毫秒 | 298/891毫秒 | 352/1625毫秒 | 435/651毫秒 | 619/1515毫秒 | 955/1980毫秒 |
| Thinker TPS | 177/171 | 556/457 | 942/598 | 57/54 | 156/149 | 266/240 |
| Talker TPS | 70/70 | 237/235 | 389/388 | 67/67 | 191/189 | 320/296 |
| 生成RTF | 0.178 | 0.211 | 0.257 | 0.187 | 0.267 | 0.334 |
3 预训练
| 模态 | 种类数量 | 支持的语言与方言 |
| 文本 | 201 | 支持语言的完整列表请参见 Qwen3.5。 |
| 语音输入 | 113 | 74 种语言:南非荷兰语、阿拉伯语、阿斯图里亚斯语、阿塞拜疆语、巴斯克语、白俄罗斯语、孟加拉语、波斯尼亚语、保加利亚语、粤语、加泰罗尼亚语、宿务语、中文、克罗地亚语、捷克语、丹麦语、荷兰语、英语、世界语、爱沙尼亚语、菲律宾语、芬兰语、法语、加利西亚语、格鲁吉亚语、德语、希腊语、希伯来语、印地语、匈牙利语、冰岛语、印度尼西亚语、国际语、意大利语、日语、爪哇语、卡纳达语、哈萨克语、韩语、吉尔吉斯语、林加拉语、拉脱维亚语、立陶宛语、马其顿语、马来语、马拉雅拉姆语、马耳他语、毛利语、马拉地语、蒙古语、书面挪威语、新挪威语、奥里亚语、波斯语、波兰语、葡萄牙语、旁遮普语、罗马尼亚语、俄语、塞尔维亚语、斯洛伐克语、斯洛文尼亚语、西班牙语、斯瓦希里语、瑞典语、塔吉克语、泰米尔语、泰卢固语、泰语、土耳其语、乌克兰语、乌尔都语、维吾尔语和越南语。39 种汉语方言:东北官话、贵州话、广东粤语、河南话、香港粤语、上海话、陕西话、天津话、台湾国语、云南话、安徽话、福建话、甘肃话、广东官话、湖北话、湖南话、江西话、山东话、山西话、四川话、广西话、海南话、重庆话、长沙话、杭州话、合肥话、银川话、郑州话、沈阳话、温州话、武汉话、昆明话、太原话、南昌话、济南话、兰州话、南京话、客家话和闽南语。 |
| 语音输出 | 36 | 29 种语言:中文、英语、德语、意大利语、葡萄牙语、西班牙语、日语、韩语、法语、俄语、泰语、印度尼西亚语、阿拉伯语、越南语、土耳其语、芬兰语、波兰语、印地语、荷兰语、捷克语、乌尔都语、他加禄语、瑞典语、丹麦语、希伯来语、冰岛语、马来语、挪威语和波斯语。7 种汉语方言:四川话、北京话、天津话、南京话、陕西话、粤语和闽南语。 |
Qwen3.5-Omni 在包含多种语言和方言(如表 3 所示)以及多种模态的多样化数据集上进行预训练,这些模态包括图像-文本、视频-文本、音频-文本、视频-音频、视频-音频-文本以及纯文本语料。继 Qwen3-Omni(Xu 等人,2025b)之后,我们采用了更广泛范围的自然语言提示词,以增强模型的泛化能力和指令遵循能力。为了在所有模态上实现稳健的性能,我们的训练策略从预训练早期阶段就同时纳入了单模态和跨模态数据。
在 Qwen3-Omni(Xu 等人,2025b)中,我们采用 TMRoPE 来赋予模型时间感知能力。然而,我们发现了该方法的两个关键局限性:(1)通过将时间位置 ID 直接与绝对时间绑定,对于较长的音频-视频或纯视频输入,会产生过大且稀疏的时间位置 ID,这削弱了模型捕捉长程时间上下文的能力。(2)在该方案下进行有效学习通常需要大规模且在不同帧率(fps)下均匀分布的采样,这显著增加了训练数据构建的成本。为解决这些问题,我们在每个视频或音频-视频时间块前添加一个时间戳,该时间戳以格式化的文本字符串形式表示(单位为秒),使模型能够更好地学习和解读时间码表示。此外,对于音频序列,我们在随机间隔处插入时间戳,以更好地对齐不同模态的训练。虽然这种方法会略微增加上下文长度,但它使模型能够更有效、更精确地感知时间信息。
Qwen3.5-Omni 的预训练分为三个不同的阶段。在第一阶段,我们锁定大语言模型参数,专注于训练视觉和音频编码器,利用海量的音频-文本和图像-文本对来增强大语言模型内的语义理解。在第二阶段,我们解冻所有参数,并使用更广泛的多模态数据进行训练,以实现更全面的学习,序列长度为 32,768。在最后阶段,我们使用序列长度为 262,144 的数据来增强模型理解复杂长序列数据的能力:
- (1)
编码器对齐阶段(S1):在初始预训练阶段,Qwen3.5-Omni 的大语言模型部分使用 Qwen3.5 的参数进行初始化,视觉编码器采用 Qwen3.5 的,而音频编码器则使用 AuT 进行初始化。这两个编码器在固定的大语言模型上分别进行训练,两者都首先专注于训练各自的适配器,然后再训练编码器。
- (2)
通用阶段(S2):预训练的第二个阶段使用了一个包含约 4 万亿个模型 token 的大规模数据集,各模态的分布如下:文本(0.92 万亿)、音频(1.99 万亿)、图像(0.95 万亿)、视频(0.14 万亿)以及视频-音频(0.29 万亿)。在此阶段,引入更多样化的多模态数据和任务,增强了模型在听觉、视觉、文本以及视听信息方面的理解和交互能力。
- (3)
长上下文阶段(S3):在最后的预训练阶段,我们将最大 token 长度从 32,768 增加到 262,144,同时也提高了训练数据中长音频和长视频的比例。实验结果表明,这些调整显著提升了模型理解长序列数据的能力。
4 后训练
4.1 思考者
后训练阶段为 Thinker 模型设计了三阶段策略,旨在保持模型在所有模态下的能力不退化,确保在音频查询下具有高响应质量,并优化整体交互体验。训练语料采用 ChatML(OpenAI,2022)格式构建,涵盖纯文本、视觉、音频及混合模态对话数据。具体而言,该流程包含以下阶段:
-
阶段一:专家知识蒸馏 为建立全模态能力的坚实基础,我们首先通过独立的监督微调(SFT)和强化学习(RL)训练一组领域专家教师模型。所有教师模型均基于预训练的 Qwen-3.5 基础检查点进行微调。除文本相关任务(包括智能体任务、编程任务和基础推理任务)外,我们还针对视觉和音频训练了专门的教师模型。这些教师模型用于生成领域特定数据,使得各领域习得的专业能力能够被蒸馏到单一统一模型中。
-
第二阶段:基于同策略的知识蒸馏 通过上述专家蒸馏,模型已在多模态理解与推理、以及基于文本的对话、推理、编码和智能体任务等领域展现出强劲性能。然而,在音频查询条件下的响应质量与文本查询条件下的响应质量之间仍存在显著差距,尤其是在语音对话场景中。为缩小这一差距,我们引入基于同策略蒸馏(OPD)的第二阶段训练流程,旨在将模型在文本输入下更强的响应能力蒸馏至音频输入场景。具体而言,对于每个音频-文本配对查询,我们首先获取在文本条件下生成的响应——该响应通常在流畅度、推理能力和任务完成度方面表现更优。随后,我们将此响应作为对应音频条件下查询的蒸馏目标。通过基于此类同策略目标进行训练,模型逐步将其音频条件下的输出与文本条件下的行为对齐,从而提升音频输入下的响应质量,并促进模态一致的生成。
-
第三阶段:交互对齐强化学习 尽管前两个阶段显著提升了模型的领域能力与跨模态响应质量,但尚不足以使模型在实际交互场景中达到完全优化。在多轮对话中,我们观察到若干交互特有的问题,包括非预期的语言代码切换、人格不一致性,以及在长上下文场景中指令遵循能力下降。为缓解这些问题,我们引入交互对齐强化学习(Interaction-Aligned RL),这是旨在优化模型交互质量的第三阶段强化学习流程。我们构建多轮交互轨迹,并围绕这些用户体验目标设计奖励信号,使模型能够在长时间交互中学习更稳定、更一致且更对齐的行为。通过显式优化交互质量,本阶段提升了模型在实际对话场景中的整体可用性。
4.2 Talker
我们为 Talker 采用了一个四阶段训练流程,使 Qwen3.5-Omni 能够与文本协同生成自然且符合语境的语音回复。所有训练数据均以 ChatML 格式组织,以保持与 Thinker 的一致性,并便于语音克隆。
- (1)
通用阶段:在初始预训练阶段,我们使用超过 2000 万小时的多语言语音数据(配有多模态上下文)对 Qwen3.5-Omni 进行训练。特别地,引入更多样化的任务(例如指令跟随式语音生成)显著增强了上下文推理能力和副语言对齐能力,超越了从多模态表征到语音的简单单调映射。
- (2)
长上下文阶段:我们通过专门的筛选流程进行数据质量分层,并在高质量子集上执行持续预训练(CPT)。借助 Qwen3-Omni-Captioner 的增强,此阶段减轻了初始预训练阶段中由噪声数据引入的模型幻觉,并显著提升了生成语音的自然度和质量。同时,我们将最大上下文长度扩展至 64k 个模型 token,使模型能够更好地处理长而复杂的用户输入,并生成更具上下文依据的语音回复。
- (3)
强化学习阶段:我们通过直接偏好优化(DPO)(Rafailov 等人,2023)进一步使模型行为与人类偏好对齐。具体而言,我们基于人工标注构建多语言偏好对,并使用 DPO 对模型进行优化。此外,我们引入了基于规则的奖励,并采用 GSPO(Zheng 等人,2025)来进一步提升整体能力以及跨不同任务的训练稳定性。
- (4)
说话人微调阶段:最后,我们在基础模型之上执行轻量级的说话人微调,使 Qwen3.5-Omni 能够忠实地捕捉目标说话人的特征,同时进一步提升其语音输出的自然度、表现力和可控性。
5 评估
我们对 Qwen3.5-Omni-Flash 和 Qwen3.5-Omni-Plus 两个模型变体进行了全面评估。评估结果主要分为两大类:理解能力(XText)和语音生成能力(XSpeech)。
5.1 XText 评估
在本节中,我们评估了 Qwen3.5-Omni 理解多种多模态输入(文本、音频、视觉以及音视频)并生成文本回复的能力。
文本到文本
我们对 Qwen3.5-Omni 在文本到文本任务上的评估主要聚焦于通用知识任务、指令遵循、长上下文任务、STEM 任务、推理任务以及通用智能体能力。具体来说,我们使用 MMLU-Pro、MMLU-Redux、SuperGPQA 和 C-Eval 评估通用知识任务,使用 IFEval 和 IFBench 评估指令遵循能力,使用 AA-LCR 和 LongBench v2 评估长上下文任务,使用 GPQA 评估 STEM 任务,使用 LiveCodeBench v6、HMMT Nov 25 和 IMOAnswerBench 评估推理任务,使用 BFCL-V4 和 TAU2Bench 评估通用智能体能力。
音频到文本
为了评估音频到文本的能力,我们采用了涵盖四个领域的基准测试:音频理解、端到端语音对话、语音到文本翻译(S2TT)和自动语音识别(ASR)。对于音频理解,我们使用 MMAU(Sakshi 等人,2024)、MMAR(Ma 等人,2025a)、MMSU(Wang 等人,2025a)、RUL-MuchoMusic(Zang 等人,2025)和 SongFormBench(Hao 等人,2025)来评估对音效、语音和音乐的理解能力。对话性能通过 VoiceBench(Chen 等人,2024b)、URO-Bench-pro(Yan 等人,2025)、SpeechRole(Jiang 等人,2025)和 WildSpeech-Bench(Zhang 等人,2025b)进行评估。对于 S2TT,我们专注于将 Fleurs(Conneau 等人,2022)中排名前 59 的语言翻译成英语和中文。最后,ASR 性能使用 Fleurs(Conneau 等人,2022)、Common Voice(Ardila 等人,2020)、LibriSpeech(Panayotov 等人,2015)、WenetSpeech(Zhang 等人,2022)、KeSpeech(Tang 等人,2021)、Opencpop-test(Wang 等人,2022)和 MIR-1K(人声)(Hsu 和 Jang,2010)来衡量,涵盖多语言语音、中文方言和歌声转录。
视觉文本
对该模型视觉转文本能力的评估涵盖了一系列针对多样且具有挑战性任务的基准测试。为了评估其在数学和STEM推理这一专业领域的表现,我们使用了MMMU(Yue等人,2023)、MMMU-Pro(Yue等人,2024)、MathVista(Lu等人,2024)、MathVision(Wang等人,2024a)、DynaMath(Zou等人,2025)和ZEROBench(Roberts等人,2025)。对于通用视觉问答,该模型在RealWorldQA(Zhang等人,2024)、MMStar(Chen等人,2024a)、HallusionBench(Guan等人,2024)和SimpleVQA(Cheng等人,2025)上进行了评估。该模型在文档理解方面的能力通过CharXiv(Wang等人,2024e)、CC-OCR(Yang等人,2024b)、AI2D(Kembhavi等人,2016)、MMLongBench-Doc(Ma等人,2024)和OCRBench(Liu等人,2024)来衡量。此外,该模型的空间智能专门在ERQA(Team,2025b)、CountBench(Paiss等人,2023)、RefCOCO(Kazemzadeh等人,2014)、ODInW13(Li等人,2022)和EmbSpatialBench(Du等人,2024a)上进行了测试。为了评估在动态视觉数据上的表现,我们报告了六个视频理解基准测试的结果:Video-MME(Fu等人,2024)、MLVU(Zhou等人,2025a)、MVBench(Li等人,2024)、LVBench(Wang等人,2024b)、MMVU(Zhao等人,2025)和MME-VideoOCR(Shi等人,2025)。具体来说,我们还在三个成熟的基准测试上评估了该模型在医学视觉问答方面的表现:SLAKE(Liu等人,2021)、PMC-VQA(Zhang等人,2023)和MedXpertQA-MM(Zuo等人,2025)。此项评估旨在展示该模型全面的临床推理能力及其作为可靠医疗AI助手的潜在用途。
音视频视频文本
我们从多个维度评估模型的音视频理解能力。在文本查询评估方面,我们使用了 DailyOmni (Zhou et al., 2025b)、WorldSense (Hong et al., 2025)、AVUT (Yang et al., 2025b)、AV-SpeakerBench (Nguyen et al., 2025) 和 VideoMME (Fu et al., 2025)。为了评估模型在真实世界音视频交互场景中的能力,我们使用 Qualcomm IVD (Pourreza et al., 2025) 作为基于音频查询的评估基准。除了理解能力,我们还评估了模型在 OmniCloze (Ma et al., 2025b) 上的描述生成能力,以及在 OmniGAIA (Li et al., 2026) 上的工具使用能力。
5.1.1 文本能力表现
我们将 Qwen3.5-Omni-Plus 和 Qwen3.5-Omni-Flash 与 Qwen3.5-Plus-Instruct 进行了比较。如表 4 所示,Qwen3.5-Omni-Plus 在知识、指令遵循、长上下文理解、STEM、推理和通用智能体任务等多个维度上,展现出了与其纯文本版本相当的能力,凸显了其强大的语言能力。特别地,Qwen3.5-Omni 的指令遵循表现略优于基线模型。我们认为,OPD 和交互对齐强化学习对提升全模态大语言模型的指令遵循能力有积极作用。
| 数据集 | Qwen3.5-Plus-Instruct | Qwen3.5-Omni-Flash | Qwen3.5-Omni-Plus |
| 知识 | |||
| MMLU-Pro | 86.8 | 79.9 | 85.9 |
| MMLU-Redux | 94.3 | 90.0 | 94.2 |
| SuperGPQA | 67.4 | 54.9 | 66.4 |
| C-Eval | 92.3 | 86.0 | 92.0 |
| 指令遵循 | |||
| IFEval | 89.7 | 85.2 | 89.7 |
| IFBench | 51.1 | 38.4 | 52.6 |
| 长上下文 | |||
| AA-LCR | 62.0 | 46.0 | 57.0 |
| LongBench v2 | 60.2 | 46.4 | 59.6 |
| STEM | |||
| GPQA | 85.9 | 76.4 | 83.9 |
| 推理 | |||
| LiveCodeBench v6 | 67.1 | 56.6 | 65.6 |
| HMMT Nov 25 | 86.2 | 59.0 | 84.4 |
| IMOAnswerBench | 68.3 | 51.5 | 65.5 |
| 通用智能体 | |||
| BFCL-V4 | 66.1 | 55.3 | 63.3 |
| TAU2Bench | 82.7 | 78.0 | 81.0 |
5.1.2 音频能力表现
| 数据集 | Gemini-3.1 Pro | Qwen3.5-Omni-Flash | Qwen3.5-Omni-Plus |
| 音频理解 () | |||
| MMAU | 81.1 | 80.4 | 82.2 |
| MMAR | 83.7 | 74.0 | 80.0 |
| MMSU | 81.3 | 72.2 | 82.8 |
| RUL-MuchoMusic | 59.6 | 60.5 | 72.4 |
| SongFormBench-HarmonixSeta | 75.6 — 46.8 — 77.9 | 80.6 — 67.8 — 83.4 | 81.1 — 72.9 — 85.3 |
| SongFormBench-CNa | 78.1 — 43.2 — 71.9 | 86.7 — 66.4 — 84.6 | 87.1 — 65.7 — 84.2 |
| 对话 () | |||
| VoiceBench | 88.9 | 87.8 | 93.1 |
| URO-Bench-prob | 69.1 — 84.0 — 99.2 | 64.1 — 83.8 — 98.7 | 66.3 — 86.3 — 99.8 |
| SpeechRole | 124.2 | 119.8 | 123.5 |
| WildSpeech-Bench | 76.3 | 72.2 | 75.4 |
| S2TT () | |||
| Fleursc | 29.5 | 26.9 | 30.2 |
| Fleursc | 34.6 | 32.0 | 35.4 |
| Fleursc | 32.1 | 29.4 | 32.8 |
| ASR (WER) | |||
| Fleurs | 7.32 | 10.75 | 6.55 |
| CV15 | 8.59 — 13.40 — 6.78 | 4.25 — 3.45 — 2.68 | 3.46 — 1.95 — 2.27 |
| CV15 | 8.73 | 5.90 | 4.83 |
| Librispeech | 3.36 — 4.41 | 1.30 — 2.43 | 1.11 — 2.23 |
| Weneetspeech | 11.53 — 14.21 | 4.41 — 5.51 | 4.30 — 5.84 |
| Kespeech | 23.67 | 4.47 | 3.46 |
| MIR-1Kd | 8.76 | 4.94 | 4.56 |
| Opencpop | 6.83 | 1.11 | 1.49 |
- a
SongFormBench:我们使用统一的提示词,定义了类似 SRT 格式的输出时间戳格式,并采用封闭词表进行评估。该词表遵循官方代码库中指定的 SongForm-HX-8Class 分类。
- b
URO-Bench-Pro:我们使用 URO-Bench 的 pro 赛道,并将三个评估维度定义如下:U 代表理解(Understanding),R 代表推理(Reasoning),O 代表口语对话(Oral Conversation)。在口语维度上,我们使用了 GenStyle-en、GenStyle-zh 和多语言(Multilingual)任务。
- c
Fleurs:排名前 59 的语言包括英语、中文、粤语、韩语、日语、越南语、泰语、马来语、德语、俄语、意大利语、法语、西班牙语、葡萄牙语、荷兰语、印尼语、土耳其语、阿拉伯语、波兰语、印地语、乌尔都语、菲律宾语、波斯语、捷克语、希腊语、瑞典语、希伯来语、丹麦语、芬兰语、挪威语、冰岛语、孟加拉语、旁遮普语、爪哇语、马拉地语、斯瓦希里语、乌克兰语、古吉拉特语、卡纳达语、阿塞拜疆语、马拉雅拉姆语、宿务语、罗马尼亚语、匈牙利语、保加利亚语、白俄罗斯语、加泰罗尼亚语、泰米尔语、克罗地亚语、波斯尼亚语、斯洛伐克语、加利西亚语、吉尔吉斯语、马其顿语、斯洛文尼亚语、拉脱维亚语、爱沙尼亚语和阿斯图里亚斯语;与排名前 60 的语言列表相比,南非荷兰语被排除在外,因为 Fleurs S2TT 测试集不包含该语言。
- d
MIR-1K:转录文本已转换为简体中文。
在表5中,我们将Qwen3.5-Omni与Gemini-3.1 Pro在音频转文本性能方面进行了比较。与Gemini-3.1 Pro相比,Qwen3.5-Omni在MMAU、MMSU、RUL-MuchoMusic和SongFormBench上表现出更优的性能,同时在MMAR上取得了相当的结果,展示了其在多个音频领域的强大理解能力。在端到端语音对话方面,Qwen3.5-Omni在VoiceBench上显著优于Gemini-3.1 Pro,并在其他基准测试上与其表现持平,进一步验证了Qwen3.5-Omni在端到端语音交互方面的强大能力。对于S2TT和ASR,Qwen3.5-Omni始终优于Gemini-3.1 Pro,突显了其在多种语言、方言和领域中卓越的翻译和语音识别性能。
5.1.3 视觉文本性能
为了全面评估视觉转文本能力,我们将Qwen3.5-Omni-Flash和Qwen3.5-Omni-Plus与Qwen3.5-Plus-Instruct进行了比较。如表6所示,Qwen3.5-Omni-Plus取得了与Qwen3.5-Plus-Instruct相当的性能,同时在涉及短时长和长时长的视频理解任务上展现出更强的结果。这些发现突显了我们的模型在真实场景中强大的动态视觉感知能力,并表明了联合视频-音频训练范式的有效性。此外,我们认为音视频流构成了对现实世界现象最自然的表征,其中视觉和听觉模态是内在耦合的,而非独立处理的。
| 数据集 |
|
|
| |||
| STEM与谜题 | ||||||
| MMMU | 81.0 | 76.9 | 80.1 | |||
| MMMU-Pro | 73.8 | 68.2 | 73.9 | |||
| MathVision | 73.6 | 65.4 | 73.0 | |||
| Mathvista (mini) | 86.9 | 82.9 | 86.1 | |||
| DynaMath | 84.2 | 79.3 | 83.8 | |||
| ZEROBench | 6 | 1 | 5 | |||
| ZEROBench_sub | 31.1 | 26.0 | 34.4 | |||
| 通用VQA | ||||||
| RealWorldQA | 79.1 | 77.5 | 84.1 | |||
| MMStar | 80.3 | 75.7 | 79.4 | |||
| MMBenchEN-DEV-v1.1 | 93.8 | 88.8 | 92.8 | |||
| SimpleVQA | 66.1 | 54.4 | 65.3 | |||
| 文本识别与文档理解 | ||||||
| CharXiv (RQ) | 74.2 | 64.4 | 72.5 | |||
| CC-OCR | 83.0 | 80.8 | 83.4 | |||
| AI2D_TEST | 92.1 | 89.0 | 91.2 | |||
| MMLongBench-Doc | 59.7 | 53.6 | 57.5 | |||
| OCRBench | 91.4 | 89.1 | 91.3 | |||
| 空间智能 | ||||||
| ERQA | 53.8 | 50.0 | 54.8 | |||
| CountBench | 95.1 | 88.2 | 95.1 | |||
| RefCOCO(平均) | 95.2 | 92.6 | 95.0 | |||
| ODInW13 | 50.3 | 46.8 | 49.5 | |||
| EmbSpatialBench | 83.4 | 82.7 | 85.4 | |||
| 视频理解 | ||||||
| VideoMME(不含字幕) | 81.0 | 77.0 | 81.9 | |||
| MLVU(M-平均) | 85.1 | 81.9 | 86.8 | |||
| MVBench | 76.7 | 70.8 | 79.0 | |||
| LVBench | 68.6 | 65.7 | 71.2 | |||
| MMVU | 67.1 | 62.7 | 67.5 | |||
| MME-VideoOCR | 74.2 | 70.5 | 77.0 | |||
| 医学视觉问答 | ||||||
| SLAKE | 82.8 | 73.1 | 84.7 | |||
| PMC-VQA | 62.4 | 58.7 | 62.7 | |||
| MedXpertQA-MM | 55.3 | 44.8 | 54.7 | |||
5.1.4 音视频文本性能
我们在多种音视频任务上对比了 Qwen3.5-Omni 和 Gemini-3.1 Pro,结果如表 7 所示。在通用理解方面,Qwen3.5-Omni 在 DailyOmni 上达到了最先进的性能,并在 AVUT 上取得了可比较的结果。我们的模型在 Qualcomm IVD 上也大幅超越了 Gemini-3.1-Pro,证明了其在真实世界音视频交互场景中的有效性。此外,我们的模型在描述生成任务上表现强劲。它能够提供详细的音频、视觉以及音视频描述。在此版本中,我们还增强了模型的工具使用能力,在 OmniGAIA 上达到了 57.2%。
数据集 Gemini-3.1 Pro Qwen3.5-Omni-Flash Qwen3.5-Omni-Plus 文本查询问答 DailyOmni 82.7 81.8 84.6 WorldSense 65.5 57.9 62.8 AVUT 85.6 81.4 85.0 AV-SpeakerBench 75.1 65.2 71.3 VideoMME(含音频)a 89.0 79.3 83.7 音频查询问答 Qualcomm IVD 66.2 66.3 68.5 Omni-Cloze 57.2 63.0 64.8 智能体(工具使用) OmniGAIAb 68.9 33.9 57.2 a VideoMME 评估时设置 use_audio_in_video=True。b OmniGAIA 评估时不使用思考提示词,也不使用 <answer> 格式。所有结果均使用 DeepSeek-V3.2-Thinking 作为评判模型。
5.2 XSpeech 评估
在本节中,我们评估了 Qwen3.5-Omni 的语音生成能力。我们的评估主要关注基于文本和提示语音条件下的语音生成,遵循零样本文本转语音(TTS)设置。我们从四个角度研究该模型:
-
零样本语音生成:我们在 SEED(Anastassiou 等人,2024)上评估内容一致性,通过词错误率(WER)来衡量。
-
多语言语音生成:我们在 TTS 多语言测试集(Zhang 等人,2025a)以及基于 FLEURS(Conneau 等人,2022)构建的内部多语言测试集上,评估了零样本多语言语音生成中的内容一致性和说话人相似度。
-
跨语言语音生成:我们在 CV3-Eval(Du 等人,2025)上评估了零样本跨语言语音生成中的内容一致性。
-
自定义语音生成:我们在 TTS 多语言测试集(Zhang 等人,2025a)和内部多语言测试集上,评估了说话人微调模型的稳定性。
5.2.1 零样本语音生成评估
我们将 Qwen3.5-Omni 与最先进的零样本 TTS 系统进行了比较。如表 8 所示,Qwen3.5-Omni 在 SEED-TTS 基准测试上取得了极具竞争力的性能,在零样本语音生成中展现出强大的内容保真度。这些结果反映了我们的预训练和持续预训练流程在构建稳健的语音生成和上下文建模能力方面的有效性。此外,经过 RLHF 优化后,Qwen3.5-Omni 进一步提升了生成稳定性和自然度,在 test-en 子集上以 1.26 的词错误率取得了最佳性能。
| 数据集 | 模型 | 性能 |
| 内容一致性 | ||
| SEED test-zh — test-en | Seed-TTSICL(Anastassiou 等人,2024) | 1.11 — 2.24 |
| Seed-TTSRL(Anastassiou 等人,2024) | 1.00 — 1.94 | |
| MaskGCT(Wang 等人,2024c) | 2.27 — 2.62 | |
| E2 TTS(Eskimez 等人,2024) | 1.97 — 2.19 | |
| F5-TTS(Chen 等人,2024c) | 1.56 — 1.83 | |
| Spark TTS(Wang 等人,2025b) | 1.20 — 1.98 | |
| CosyVoice 2(Du 等人,2024b) | 1.45 — 2.57 | |
| CosyVoice 3(Du 等人,2025) | 0.71 — 1.45 | |
| MiniMax-Speech(Zhang 等人,2025a) | 0.83 — 1.65 | |
| MiMo-Audio-7B-Instruct(Zhang 等人,2025c) | 1.96 — 5.37 | |
| Qwen2.5-Omni-7B(Xu 等人,2025a) | 1.42 — 2.33 | |
| Qwen3-Omni-30B-A3B(Xu 等人,2025b) | 1.07 — 1.39 | |
| Qwen3.5-Omni-Plus | 0.99 — 1.26 | |
5.2.2 多语言语音生成评估
Qwen3.5-Omni 支持 29 种语言的语音生成。我们将其多语言语音生成性能与两个强大的商业系统 MiniMax-Speech 和 ElevenLabs 进行了比较。对于内部多语言测试集,我们使用 GPT-4o-transcribe-2025-03-20 进行自动语音识别。
如表 9 和表 10 所示,在多语言测试集的 29 种评估语言中,Qwen3.5-Omni 在 22 种语言上取得了最低的词错误率,在大多数情况下以明显优势超越了对比系统。在其余语言上,Qwen3.5-Omni 仍与最先进的系统保持竞争力。除了内容一致性,Qwen3.5-Omni 还展现出强大的语音克隆保真度。它在大多数评估语言中获得了最高的说话人相似度分数,并且在整体上持续优于 MiniMax-Speech 和 ElevenLabs。这些结果表明,Qwen3.5-Omni 在保持稳健的多语言语音生成质量的同时,有效地保留了说话人特征,如音色和韵律风格。
此外,在表 10 中,我们报告了在内部多语言测试集上的结果,该测试集额外涵盖了 9 种语言。Qwen3.5-Omni 在所有评估语言上持续取得强劲表现,表明其多语言语音生成能力能够很好地泛化到公开基准语言之外。
| 语言 | 内容一致性 | 说话人相似度 | |||||
| MiniMax | ElevenLabs |
| MiniMax | ElevenLabs | ||
| 中文 | 0.695 | 2.252 | 16.026 | 0.800 | 0.780 | 0.677 | |
| 英语 | 0.631 | 2.164 | 0.756 | 0.833 | 0.756 | 0.613 | |
| 德语 | 0.447 | 1.906 | 0.572 | 0.757 | 0.733 | 0.614 | |
| 意大利语 | 0.503 | 1.543 | 1.743 | 0.785 | 0.699 | 0.679 | |
| 葡萄牙语 | 1.221 | 1.877 | 1.331 | 0.792 | 0.805 | 0.711 | |
| 西班牙语 | 0.862 | 1.029 | 1.084 | 0.797 | 0.762 | 0.615 | |
| 日语 | 3.479 | 3.519 | 10.046 | 0.788 | 0.776 | 0.738 | |
| 韩语 | 1.458 | 1.747 | 1.865 | 0.747 | 0.776 | 0.700 | |
| 法语 | 2.430 | 4.099 | 5.216 | 0.730 | 0.628 | 0.535 | |
| 俄语 | 3.182 | 4.281 | 3.878 | 0.790 | 0.761 | 0.676 | |
| 泰语 | 2.170 | 2.701 | 73.936 | 0.788 | 0.800 | 0.588 | |
| 印尼语 | 0.823 | 1.237 | 1.059 | 0.780 | 0.729 | 0.660 | |
| 阿拉伯语 | 2.602 | 1.665 | 1.666 | 0.745 | 0.736 | 0.706 | |
| 越南语 | 1.143 | 0.880 | 73.415 | 0.767 | 0.743 | 0.369 | |
| 土耳其语 | 0.938 | 1.520 | 0.699 | 0.747 | 0.779 | 0.596 | |
| 芬兰语 | 2.784 | 4.666 | 2.964 | 0.859 | 0.835 | 0.759 | |
| 波兰语 | 1.427 | 1.415 | 0.766 | 0.839 | 0.802 | 0.729 | |
| 印地语 | 6.444 | 6.962 | 5.827 | 0.797 | 0.818 | 0.730 | |
| 荷兰语 | 1.238 | 1.143 | 0.803 | 0.762 | 0.738 | 0.680 | |
| 捷克语 | 2.929 | 3.875 | 2.108 | 0.802 | 0.796 | 0.685 | |
| 语言 | 内容一致性 | 说话人相似度 | |||
| 真实值 |
| 真实值 | ||
| 乌尔都语 | 14.819 | 17.822 | 0.775 | - | |
| 他加禄语 | 5.193 | 6.885 | 0.870 | - | |
| 瑞典语 | 3.760 | 4.813 | 0.822 | - | |
| 丹麦语 | 3.636 | 6.403 | 0.775 | - | |
| 希伯来语 | 7.860 | 16.178 | 0.760 | - | |
| 冰岛语 | 10.244 | 11.451 | 0.764 | - | |
| 马来语 | 3.142 | 4.628 | 0.794 | - | |
| 挪威语 | 3.613 | 4.442 | 0.825 | - | |
| 波斯语 | 11.113 | 14.469 | 0.800 | - | |
5.2.3 跨语言语音生成评估
除了多语言语音克隆,Qwen3.5-Omni 还支持跨语言语音克隆,即模型需要在生成不同目标语言的语音时保留说话人身份。我们在跨语言基准上评估了这一能力,并与 CosyVoice 系列以及 Qwen3-Omni-30B-A3B 进行了比较。
在表 11 中,我们报告了不同源语言-目标语言对上的混合错误率(英语为 WER,其他语言为 CER)。总体而言,Qwen3.5-Omni 在 12 个评估方向中的 10 个上取得了最佳性能,并在大多数以英语、日语和韩语为目标的语言对中树立了新的最先进水平。特别是,对于中文到韩语,与 CosyVoice3 相比,Qwen3.5-Omni 将错误率从 14.4 降低到了 4.03,相对降低了约 72%。Qwen3.5-Omni 在常用的语言对(如中文到英语和英语到中文)上也表现强劲,表明其在跨语言生成下具有更好的内容一致性。这些结果证明,Qwen3.5-Omni 能够有效地跨越语言边界进行泛化,同时保持目标语言的准确性。
| 语言 | Qwen3.5-Omni-Plus | Qwen3-Omni-30B-A3B | CosyVoice3 | CosyVoice2 |
| 英译中 | 4.86 | 5.37 | 5.09 | 13.5 |
| 日译中 | 3.55 | 3.32 | 3.05 | 48.1 |
| 韩译中 | 0.84 | 0.99 | 1.06 | 7.70 |
| 中译英 | 2.18 | 2.76 | 2.98 | 6.47 |
| 日译英 | 2.18 | 3.31 | 4.20 | 17.1 |
| 韩译英 | 2.51 | 3.34 | 4.19 | 11.2 |
| 中译日 | 5.92 | 8.29 | 7.08 | 13.1 |
| 英译日 | 5.12 | 7.53 | 6.80 | 14.9 |
| 韩译日 | 2.16 | 4.24 | 3.93 | 5.86 |
| 中译韩 | 4.03 | 5.13 | 14.4 | 24.8 |
| 英译韩 | 3.72 | 4.96 | 5.87 | 21.9 |
| 日译韩 | 5.12 | 6.23 | 7.92 | 21.5 |
5.2.4 自定义语音生成评估
我们在多语言场景下评估了 Qwen3.5-Omni 的自定义语音生成能力。我们将 Qwen3.5-Omni 与多个通过官方 API 在 2026 年 3 月访问的强商业系统进行了比较,包括 ElevenLabs Multilingual v2 (9YHcvj6GT2YYXdXww)、Gemini-2.5 Pro-Preview-TTS (Achernar)、GPT-Audio-2025-08-28 (Alloy) 和 MiniMax-Speech-2.8-HD (English_expressive_narrator)。
| 语言 | Qwen3.5-Omni-Plus | ElevenLabs | Gemini-2.5 Pro | GPT-Audio | MiniMax |
| 中文 | 0.785 | 3.801 | 1.890 | 0.829 | 0.786 |
| 英语 | 0.839 | 1.126 | 0.953 | 1.050 | 1.429 |
| 德语 | 0.182 | 0.500 | 0.509 | 0.558 | 1.581 |
| 意大利语 | 0.458 | 0.513 | 0.991 | 0.769 | 1.063 |
| 葡萄牙语 | 1.581 | 1.109 | 2.050 | 1.506 | 1.240 |
| 西班牙语 | 0.768 | 0.520 | 0.891 | 0.936 | 0.691 |
| 日语 | 3.306 | 11.685 | 4.420 | 4.317 | 4.254 |
| 韩语 | 1.309 | 3.981 | 4.110 | 3.999 | 3.635 |
| 法语 | 2.724 | 2.574 | 3.284 | 2.809 | 3.439 |
| 俄语 | 4.723 | 4.324 | 3.858 | 4.346 | 3.529 |
| 泰语 | 1.653 | 114.813 | 2.539 | 4.430 | 1.811 |
| 印尼语 | 1.596 | 6.094 | 1.498 | 2.362 | 1.585 |
| 阿拉伯语 | 3.183 | 5.400 | 5.525 | 5.326 | 3.309 |
| 越南语 | 1.320 | 82.849 | 1.699 | 1.854 | 1.058 |
| 土耳其语 | 1.309 | 0.551 | 2.237 | 1.389 | 0.652 |
| 芬兰语 | 4.039 | 2.522 | 5.331 | 3.270 | 2.939 |
| 波兰语 | 1.462 | 0.733 | 1.622 | 1.737 | 0.833 |
| 印地语 | 6.776 | 6.388 | 6.596 | 7.191 | 6.146 |
| 荷兰语 | 1.135 | 1.005 | 0.973 | 1.561 | 1.406 |
| 捷克语 | 3.769 | 1.916 | 3.380 | 2.859 | 1.766 |
| 乌尔都语 | 14.916 | 12.970 | 14.141 | 13.362 | 24.151 |
| 他加禄语 | 5.090 | 5.473 | 6.784 | 5.352 | 5.674 |
| 瑞典语 | 3.588 | 3.132 | 3.196 | 2.898 | 2.833 |
| 丹麦语 | 7.183 | 2.604 | 3.876 | 3.846 | 4.951 |
| 希伯来语 | 7.680 | 102.018 | 4.459 | 5.328 | 8.161 |
| 冰岛语 | 10.322 | 25.110 | 6.348 | 9.648 | 33.431 |
| 马来语 | 3.738 | 6.448 | 3.731 | 3.406 | 3.955 |
| 挪威语 | 5.576 | 7.351 | 4.304 | 3.400 | 9.492 |
| 波斯语 | 12.140 | 20.564 | 12.620 | 13.202 | 12.722 |
如表12所示,尽管Qwen3.5-Omni仅在单语数据上进行微调,但它在定制语音生成方面展现出强大的跨语言泛化能力。该模型能够将目标说话人的特征迁移至全部29种评估语言,同时保持稳定的生成质量。总体而言,Qwen3.5-Omni在10种语言上取得了最佳词错误率,并在其他多种语言上保持竞争力。特别是在日语(3.306)和韩语(1.309)等若干具有挑战性的语言上,它展现出明显优势,表明其在跨语言语音迁移下具有强大的可懂度。这些结果表明,Qwen3.5-Omni能够在广泛的语言范围内生成具有稳健语言保真度的定制语音。
6 结论
在本工作中,我们提出了Qwen3.5-Omni,一个全模态大语言模型,它统一了对文本、图像、音频及视听输入的理解、推理、生成与行动能力。基于Thinker–Talker框架构建,Qwen3.5-Omni引入了高效的混合注意力MoE架构、256k长上下文建模、通过多码本编解码预测和ARIA实现的改进型流式语音生成,以及大幅扩展的多语言语音支持。这些进步带来了三项关键能力:可控的视听字幕生成、全面的实时交互,以及通过自主工具使用和视听代码生成实现的原生全模态智能体行为。实验表明,Qwen3.5-Omni在广泛的音频和视听基准测试中取得了最先进或极具竞争力的性能,同时保持了同规模Qwen模型强大的文本和视觉能力。这些结果表明,扩展原生全模态训练可以产生统一的系统,这些系统不仅能跨模态感知和推理,还能实时交互和行动。我们希望Qwen3.5-Omni能为未来通用全模态智能体的研究提供坚实基础。
参考文献
- P. Anastassiou, J. Chen, J. Chen, Y. Chen, Z. Chen, Z. Chen, J. Cong, L. Deng, C. Ding, L. Gao, 等人 (2024) Seed-tts:一个高质量、多功能的语音生成模型系列。arXiv 预印本 arXiv:2406.02430。引用自:第1项,表8,表8。
- Anthropic (2023a) Claude 2。Anthropic 技术报告。外部链接:链接。引用自:§1。
- Anthropic (2023b) 介绍 Claude。Anthropic。外部链接:链接。引用自:§1。
- Anthropic (2024) Claude 3 模型系列:Opus、Sonnet、Haiku。Anthropic AI 技术报告。外部链接:链接。引用自:§1。
- R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. M. Tyers, 和 G. Weber (2020) Common Voice:一个大规模多语言语音语料库。载于《第12届语言资源与评估会议论文集》,LREC 2020,法国马赛,2020年5月11-16日,N. Calzolari, F. Béchet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, 和 S. Piperidis (编),第4218–4222页。外部链接:链接。引用自:§5.1。
- J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, B. Hui, L. Ji, M. Li, J. Lin, R. Lin, D. Liu, G. Liu, C. Lu, K. Lu, J. Ma, R. Men, X. Ren, X. Ren, C. Tan, S. Tan, J. Tu, P. Wang, S. Wang, W. Wang, S. Wu, B. Xu, J. Xu, A. Yang, H. Yang, J. Yang, S. Yang, Y. Yao, B. Yu, H. Yuan, Z. Yuan, J. Zhang, X. Zhang, Y. Zhang, Z. Zhang, C. Zhou, J. Zhou, X. Zhou, 和 T. Zhu (2023a) Qwen 技术报告。CoRR abs/2309.16609。引用自:§1。
- J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, 和 J. Zhou (2023b) Qwen-VL:一个具有多功能能力的先进大视觉语言模型。CoRR abs/2308.12966。引用自:§1。
- S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, 等人 (2025a) Qwen2.5-VL 技术报告。arXiv 预印本 arXiv:2502.13923。引用自:§1。
- Y. Bai, S. Tu, J. Zhang, H. Peng, X. Wang, X. Lv, S. Cao, J. Xu, L. Hou, Y. Dong, J. Tang, 和 J. Li (2025b) 《LongBench v2:迈向对现实长上下文多任务的更深层次理解与推理》。载于《第63届计算语言学协会年会论文集(第一卷:长文)》,ACL 2025,奥地利维也纳,2025年7月27日至8月1日,W. Che、J. Nabende、E. Shutova 和 M. T. Pilehvar 编,第3639–3664页。外部链接:Link 引用自:§5.1。
- M. Balunović、J. Dekoninck、I. Petrov、N. Jovanović 和 M. Vechev (2025) 《MathArena:在无污染数学竞赛中评估大语言模型》。苏黎世联邦理工学院 SRI 实验室。外部链接:Link 引用自:§5.1。
- V. Barres、H. Dong、S. Ray、X. Si 和 K. Narasimhan (2025) 《-Bench:在双控环境中评估对话智能体》。外部链接:2506.07982, Link 引用自:§5.1。
- T. Brown、B. Mann、N. Ryder、M. Subbiah、J. D. Kaplan、P. Dhariwal、A. Neelakantan、P. Shyam、G. Sastry、A. Askell 等人 (2020) 《语言模型是少样本学习者》。载于 NeurIPS。引用自:§1。
- L. Chen、J. Li、X. Dong、P. Zhang、Y. Zang、Z. Chen、H. Duan、J. Wang、Y. Qiao、D. Lin 等人 (2024a) 《我们走在评估大型视觉语言模型的正确道路上吗?》。arXiv:2403.20330。引用自:§5.1。
- Y. Chen、X. Yue、C. Zhang、X. Gao、R. T. Tan 和 H. Li (2024b) 《Voicebench:对基于大语言模型的语音助手进行基准测试》。arXiv 预印本 arXiv:2410.17196。引用自:§5.1。
- Y. Chen、Z. Niu、Z. Ma、K. Deng、C. Wang、J. Zhao、K. Yu 和 X. Chen (2024c) 《F5-tts:一个利用流匹配生成流畅且忠实语音的童话讲述者》。arXiv 预印本 arXiv:2410.06885。引用自:表8。
- X. Cheng、W. Zhang、S. Zhang、J. Yang、X. Guan、X. Wu、X. Li、G. Zhang、J. Liu、Y. Mai、Y. Zeng、Z. Wen、K. Jin、B. Wang、W. Zhou、Y. Lu、T. Li、W. Huang 和 Z. Li (2025) 《SimpleVQA:面向多模态大语言模型的多模态事实性评估》。CoRR abs/2502.13059。引用自:§5.1。
- Y. Chu、J. Xu、Q. Yang、H. Wei、X. Wei、Z. Guo、Y. Leng、Y. Lv、J. He、J. Lin 等人 (2024) 《Qwen2-Audio 技术报告》。arXiv 预印本 arXiv:2407.10759。引用自:§1。
- Y. Chu, J. Xu, X. Zhou, Q. Yang, S. Zhang, Z. Yan, C. Zhou, and J. Zhou (2023) Qwen-Audio:通过统一的大规模音频-语言模型推进通用音频理解。CoRR abs/2311.07919。引用于:§1。
- G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen, et al. (2025) Gemini 2.5:以高级推理、多模态、长上下文和下一代智能体能力突破前沿。arXiv 预印本 arXiv:2507.06261。引用于:§1。
- A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna (2022) FLEURS:语音通用表示的少样本学习评估。2022 年 IEEE 口语语言技术研讨会 (SLT),第 798–805 页。外部链接:Link。引用于:第 2 项,§5.1。
- M. Du, B. Wu, Z. Li, X. Huang, and Z. Wei (2024a) EmbSpatial-bench:基于大型视觉-语言模型对具身任务的空间理解进行基准测试。载于《第 62 届计算语言学协会年会论文集(第 2 卷:短论文)》,ACL 2024,2024 年 8 月 11-16 日,泰国曼谷,L. Ku, A. Martins, and V. Srikumar (编),第 346–355 页。引用于:§5.1。
- Z. Du, C. Gao, Y. Wang, F. Yu, T. Zhao, H. Wang, X. Lv, H. Wang, C. Ni, X. Shi, K. An, G. Yang, Y. Li, Y. Chen, Z. Gao, Q. Chen, Y. Gu, M. Chen, Y. Chen, S. Zhang, W. Wang, and J. Ye (2025) CosyVoice 3:通过规模扩展和后训练实现真实场景语音生成。CoRR abs/2505.17589。引用于:第 3 项,表 8。
- Z. Du, Y. Wang, Q. Chen, X. Shi, X. Lv, T. Zhao, Z. Gao, Y. Yang, C. Gao, H. Wang, et al. (2024b) CosyVoice 2:基于大语言模型的可扩展流式语音合成。arXiv 预印本 arXiv:2412.10117。引用于:表 8。
- A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Rozière, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. M. Kloumann, I. Misra, I. Evtimov, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, 等人 (2024) 《Llama 3 模型家族》。CoRR abs/2407.21783。引用自:§1。
- S. E. Eskimez, X. Wang, M. Thakker, C. Li, C. Tsai, Z. Xiao, H. Yang, Z. Zhu, M. Tang, X. Tan, 等人 (2024) 《E2 TTS:极其简单的全非自回归零样本语音合成》。载于《2024年IEEE口语技术研讨会(SLT)》,第682–689页。引用自:表8。
- C. Fu, Y. Dai, Y. Luo, L. Li, S. Ren, R. Zhang, Z. Wang, C. Zhou, Y. Shen, M. Zhang, 等人 (2024) 《Video-MME:首个多模态大语言模型视频分析综合评估基准》。arXiv:2405.21075。引用自:§5.1。
- C. Fu, Y. Dai, Y. Luo, L. Li, S. Ren, R. Zhang, Z. Wang, C. Zhou, Y. Shen, M. Zhang, 等人 (2025) 《Video-MME:首个多模态大语言模型视频分析综合评估基准》。载于《IEEE/CVF计算机视觉与模式识别会议论文集》,第24108–24118页。引用自:§5.1。
- A. P. Gema, J. O. J. Leang, G. Hong, A. Devoto, A. C. M. Mancino, R. Saxena, X. He, Y. Zhao, X. Du, M. R. G. Madani, 等人 (2024) 《MMLU 我们做完了吗?》。CoRR abs/2406.04127。引用自:§5.1。
- Gemini 团队 (2024) Gemini 1.5:解锁跨越数百万 token 上下文的多模态理解能力。技术报告,Google。外部链接:链接 被引用:§1。
- T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y. Yacoob, D. Manocha, 和 T. Zhou (2024) Hallusionbench:针对大型视觉语言模型中纠缠的语言幻觉与视觉错觉的高级诊断套件。收录于 IEEE/CVF 计算机视觉与模式识别会议,CVPR 2024,美国华盛顿州西雅图,2024年6月16-22日,第14375–14385页。被引用:§5.1。
- C. Hao, R. Yuan, J. Yao, Q. Deng, X. Bai, W. Xue, 和 L. Xie (2025) SongFormer:利用异构监督扩展音乐结构分析。外部链接:2510.02797,链接 被引用:§5.1。
- J. Hong, S. Yan, J. Cai, X. Jiang, Y. Hu, 和 W. Xie (2025) WorldSense:评估多模态大语言模型的真实世界全模态理解能力。CoRR abs/2502.04326。被引用:§5.1。
- C. Hsu 和 J. R. Jang (2010) 关于利用 MIR-1K 数据集改进单声道录音中歌声分离的研究。IEEE 语音与音频处理汇刊 18 (2),第310–319页。外部链接:链接,文献标识码:Document 被引用:§5.1。
- Y. Huang, Y. Bai, Z. Zhu, J. Zhang, J. Zhang, T. Su, J. Liu, C. Lv, Y. Zhang, J. Lei, Y. Fu, M. Sun, 和 J. He (2023) C-Eval:面向基础模型的多层级多学科中文评测套件。收录于 NeurIPS。被引用:§5.1。
- N. Jain, K. Han, A. Gu, W. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, 和 I. Stoica (2024) LiveCodeBench:对代码大语言模型进行整体且无污染的评估。CoRR abs/2403.07974。被引用:§5.1。
- C. Jiang, J. Sun, Y. Cao, J. Zhuang, H. Li, X. Fan, M. Zhang, J. Ye, S. Dou, Z. Xi, 等 (2025) SpeechRole:用于评估语音角色扮演智能体的大规模数据集与基准。arXiv 预印本 arXiv:2508.02013。被引用:§5.1。
- S. Kazemzadeh, V. Ordonez, M. Matten, 和 T. Berg (2014) Referitgame:指代自然场景照片中的物体。收录于 EMNLP。被引用:§5.1。
- A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, 和 A. Farhadi (2016) 一张图表胜过一打图像。收录于 ECCV。被引用:§5.1。
- J. Li, D. Li, S. Savarese 和 S. Hoi (2023) 《Blip-2:利用冻结图像编码器与大语言模型引导语言-图像预训练》。arXiv:2301.12597。引用自:§1。
- K. Li, Y. Wang, Y. He, Y. Li, Y. Wang, Y. Liu, Z. Wang, J. Xu, G. Chen, P. Luo 等 (2024) 《Mvbench:全面的多模态视频理解基准》。发表于 CVPR。引用自:§5.1。
- L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J. Hwang, K. Chang 和 J. Gao (2022) 《基于语言-图像的预训练》。发表于 CVPR,第 10955–10965 页。引用自:§5.1。
- X. Li, W. Jiao, J. Jin, S. Wang, G. Dong, J. Jin, H. Wang, Y. Wang, J. Wen, Y. Lu 等 (2026) 《OmniGAIA:迈向原生全模态 AI 智能体》。arXiv 预印本 arXiv:2602.22897。引用自:§5.1。
- B. Liu, L. Zhan, L. Xu, L. Ma, Y. Yang 和 X. Wu (2021) 《Slake:面向医学视觉问答的语义标注知识增强数据集》。发表于第 18 届 IEEE 国际生物医学成像研讨会 (ISBI 2021),法国尼斯,2021 年 4 月 13-16 日,第 1650–1654 页。引用自:§5.1。
- H. Liu, C. Li, Q. Wu 和 Y. J. Lee (2023) 《视觉指令微调》。arXiv:2304.08485。引用自:§1。
- Y. Liu, Z. Li, M. Huang, B. Yang, W. Yu, C. Li, X. Yin, C. Liu, L. Jin 和 X. Bai (2024) 《OCRBench:大型多模态模型中 OCR 的隐藏奥秘》。中国科学:信息科学 67 (12)。外部链接:ISSN 1869-1919,链接,文献。引用自:§5.1。
- P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K. Chang, M. Galley 和 J. Gao (2024) 《MathVista:评估基础模型在视觉情境中的数学推理能力》。发表于 ICLR。引用自:§5.1。
- T. Luong, D. Hwang, H. H. Nguyen, G. Ghiasi, Y. Chervonyi, I. Seo, J. Kim, G. Bingham, J. Lee, S. Mishra, A. Zhai, C. H. Hu, H. Michalewski, J. Kim, J. Ahn, J. Bae, X. Song, T. H. Trinh, Q. V. Le 和 J. Jung (2025) 《迈向稳健的数学推理》。发表于 2025 年自然语言处理实证方法会议论文集。外部链接:链接。引用自:§5.1。
- Y. Ma, Y. Zang, L. Chen, M. Chen, Y. Jiao, X. Li, X. Lu, Z. Liu, Y. Ma, X. Dong, P. Zhang, L. Pan, Y. Jiang, J. Wang, Y. Cao, 和 A. Sun (2024) MMLONGBENCH-DOC:利用可视化技术对长上下文文档理解进行基准测试。收录于《神经信息处理系统进展 38:2024 年神经信息处理系统年度大会,NeurIPS 2024,加拿大不列颠哥伦比亚省温哥华,2024 年 12 月 10 日至 15 日》,A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, 和 C. Zhang 编。引用自:§5.1。
- Z. Ma, Y. Ma, Y. Zhu, C. Yang, Y. Chao, R. Xu, W. Chen, Y. Chen, Z. Chen, J. Cong, K. Li, K. Li, S. Li, X. Li, X. Li, Z. Lian, Y. Liang, M. Liu, Z. Niu, T. Wang, Y. Wang, Y. Wang, Y. Wu, G. Yang, J. Yu, R. Yuan, Z. Zheng, Z. Zhou, H. Zhu, W. Xue, E. Benetos, K. Yu, C. E. Siong, 和 X. Chen (2025a) MMAR:针对语音、音频、音乐及其混合内容的深度推理挑战性基准。CoRR abs/2505.13032。外部链接:Link, Document, 2505.13032。引用自:§5.1。
- Z. Ma, R. Xu, Z. Xing, Y. Chu, Y. Wang, J. He, J. Xu, P. Heng, K. Yu, J. Lin, 等人 (2025b) Omni-captioner:面向全模态精细感知的数据管线、模型与基准。arXiv 预印本 arXiv:2510.12720。引用自:§5.1。
- L. T. P. Nguyen, Z. Yu, S. L. Y. Hang, S. An, J. Lee, Y. Ban, S. Chung, T. Nguyen, J. Maeng, S. Lee, 等人 (2025) 看、听、理解:在多模态大语言模型中基准测试视听人类语音理解能力。arXiv 预印本 arXiv:2512.02231。引用自:§5.1。
- OpenAI (2022) ChatML。外部链接:Link。引用自:§4.1。
- OpenAI (2023) GPT4 技术报告。CoRR abs/2303.08774。引用自:§1。
- OpenAI (2024) 你好,GPT-4o。外部链接:Link。引用自:§1。
- R. Paiss, A. Ephrat, O. Tov, S. Zada, I. Mosseri, M. Irani, 和 T. Dekel (2023) 教 CLIP 数到十。收录于《IEEE/CVF 国际计算机视觉大会,ICCV 2023,法国巴黎,2023 年 10 月 1-6 日》,第 3147–3157 页。引用自:§5.1。
- V. Panayotov, G. Chen, D. Povey 和 S. Khudanpur (2015) Librispeech:一个基于公共领域有声读物的 ASR 语料库。载于 2015 年 IEEE 国际声学、语音与信号处理会议 (ICASSP 2015),澳大利亚昆士兰州南布里斯班,2015 年 4 月 19-24 日,引用自:§5.1。
- R. Pourreza, R. Dagli, A. Bhattacharyya, S. Panchal, G. Berger 和 R. Memisevic (2025) 视觉语言模型能否回答现实世界中的面对面问题?arXiv 预印本 arXiv:2503.19356。引用自:§5.1。
- V. Pyatkin, S. Malik, V. Graf, H. Ivison, S. Huang, P. Dasigi, N. Lambert 和 H. Hajishirzi (2025) 泛化可验证的指令遵循。CoRR abs/2507.02833。外部链接:链接,文档,2507.02833。引用自:§5.1。
- R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon 和 C. Finn (2023) 直接偏好优化:你的语言模型其实是一个奖励模型。载于 NeurIPS,引用自:第 (3) 项。
- D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael 和 S. R. Bowman (2023) GPQA:一个研究生级别的、谷歌无法破解的问答基准。CoRR abs/2311.12022。引用自:§5.1。
- J. Roberts, M. R. Taesiri, A. Sharma, A. Gupta, S. Roberts, I. Croitoru, S. Bogolin, J. Tang, F. Langer, V. Raina, V. Raina, H. Xiong, V. Udandarao, J. Lu, S. Chen, S. Purkis, T. Yan, W. Lin, G. Shin, Q. Yang, A. T. Nguyen, K. Han 和 S. Albanie (2025) ZeroBench:一个针对当代大型多模态模型的、不可能完成的视觉基准。CoRR abs/2502.09696。引用自:§5.1。
- S. Sakshi, U. Tyagi, S. Kumar, A. Seth, R. Selvakumar, O. Nieto, R. Duraiswami, S. Ghosh 和 D. Manocha (2024) MMAU:一个大规模多任务音频理解与推理基准。外部链接:2410.19168,链接。引用自:§5.1。
- Y. Shi, H. Wang, W. Xie, H. Zhang, L. Zhao, Y. Zhang, X. Li, C. Fu, Z. Wen, W. Liu, Z. Zhang, X. Chen, B. Zeng, S. Yang, Y. Zhang, P. Wan, H. Wang 和 W. Yang (2025) MME-videoocr:评估多模态大语言模型在视频场景中基于 OCR 的能力。CoRR abs/2505.21333。引用自:§5.1。
- Z. Tang, D. Wang, Y. Xu, J. Sun, X. Lei, S. Zhao, C. Wen, X. Tan, C. Xie, S. Zhou, R. Yan, C. Lv, Y. Han, W. Zou, 和 X. Li (2021) KeSpeech:一个包含普通话及其八种次方言的开源语音数据集。收录于《神经信息处理系统进展——数据集与基准赛道 1》,NeurIPS 数据集与基准 2021,2021 年 12 月,线上会议,J. Vanschoren 和 S. Yeung 编,外部链接:链接,被 §5.1 引用。
- A. A. Team (2025a) 人工分析长上下文推理基准 (LCR)。注:Artificial Analysis, Inc. 数据集,被 §5.1 引用。
- G. R. Team (2025b) Gemini 机器人:将 AI 带入物理世界。CoRR abs/2503.20020。被 §5.1 引用。
- M.-A-P. Team, X. Du, Y. Yao, K. Ma, B. Wang, T. Zheng, K. Zhu, M. Liu, Y. Liang, X. Jin, Z. Wei, C. Zheng, K. Deng, S. Jia, S. Jiang, Y. Liao, R. Li, Q. Li, S. Li, Y. Li, Y. Li, D. Ma, Y. Ni, H. Que, Q. Wang, Z. Wen, S. Wu, T. Xing, M. Xu, Z. Yang, Z. M. Wang, J. Zhou, Y. Bai, X. Bu, C. Cai, L. Chen, Y. Chen, C. Cheng, T. Cheng, K. Ding, S. Huang, Y. Huang, Y. Li, Y. Li, Z. Li, T. Liang, C. Lin, H. Lin, Y. Ma, T. Pang, Z. Peng, Z. Peng, Q. Qi, S. Qiu, X. Qu, S. Quan, Y. Tan, Z. Wang, C. Wang, H. Wang, Y. Wang, Y. Wang, J. Xu, K. Yang, R. Yuan, Y. Yue, T. Zhan, C. Zhang, J. Zhang, X. Zhang, X. Zhang, Y. Zhang, Y. Zhao, X. Zheng, C. Zhong, Y. Gao, Z. Li, D. Liu, Q. Liu, T. Liu, S. Ni, J. Peng, Y. Qin, W. Su, G. Wang, S. Wang, J. Yang, M. Yang, M. Cao, X. Yue, Z. Zhang, W. Zhou, J. Liu, Q. Lin, W. Huang, 和 G. Zhang (2025) SuperGPQA:将大语言模型评估扩展至 285 个研究生学科。CoRR abs/2502.14739。外部链接:链接,文档,2502.14739,被 §5.1 引用。
- Q. Team (2026) Qwen3.5:通过原生多模态智能体加速生产力。外部链接:链接,被 §2.3 引用。
- H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, 等人 (2023) Llama 2:开放的基础模型与微调聊天模型。arXiv:2307.09288。被 §1 引用。
- D. Wang, J. Wu, J. Li, D. Yang, X. Chen, T. Zhang, 和 H. Meng (2025a) 《MMSU:大规模多任务口语语言理解与推理基准》。CoRR abs/2506.04779。外部链接:Link, Document, 2506.04779 引用自:§5.1。
- K. Wang, J. Pan, W. Shi, Z. Lu, M. Zhan, 和 H. Li (2024a) 《使用数学视觉数据集衡量多模态数学推理》。arXiv:2402.14804。引用自:§5.1。
- W. Wang, Z. He, W. Hong, Y. Cheng, X. Zhang, J. Qi, S. Huang, B. Xu, Y. Dong, M. Ding, 和 J. Tang (2024b) 《LVBench:极端长视频理解基准》。CoRR abs/2406.08035。引用自:§5.1。
- X. Wang, M. Jiang, Z. Ma, Z. Zhang, S. Liu, L. Li, Z. Liang, Q. Zheng, R. Wang, X. Feng, W. Bian, Z. Ye, S. Cheng, R. Yuan, Z. Zhao, X. Zhu, J. Pan, L. Xue, P. Zhu, Y. Chen, Z. Li, X. Chen, L. Xie, Y. Guo, 和 W. Xue (2025b) 《Spark-tts:一种基于高效大语言模型的文本转语音模型,采用单流解耦语音模型 token》。CoRR abs/2503.01710。引用自:表 8。
- Y. Wang, X. Wang, P. Zhu, J. Wu, H. Li, H. Xue, Y. Zhang, L. Xie, 和 M. Bi (2022) 《Opencpop:用于歌声合成的高质量开源中文流行歌曲语料库》。收录于:第23届国际语音通信协会年度会议,Interspeech 2022,韩国仁川,2022年9月18-22日,H. Ko 和 J. H. L. Hansen (编),第4242–4246页。外部链接:Link, Document 引用自:§5.1。
- Y. Wang, H. Zhan, L. Liu, R. Zeng, H. Guo, J. Zheng, Q. Zhang, X. Zhang, S. Zhang, 和 Z. Wu (2024c) 《Maskgct:基于掩码生成式编解码器 Transformer 的零样本文本转语音》。arXiv 预印本 arXiv:2409.00750。引用自:表 8。
- Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, T. Li, M. Ku, K. Wang, A. Zhuang, R. Fan, X. Yue, 和 W. Chen (2024d) 《MMLU-Pro:一个更鲁棒且更具挑战性的多任务语言理解基准》。CoRR abs/2406.01574。引用自:§5.1。
- Z. Wang, M. Xia, L. He, H. Chen, Y. Liu, R. Zhu, K. Liang, X. Wu, H. Liu, S. Malladi, A. Chevalier, S. Arora, 和 D. Chen (2024e) 《CharXiv:揭示多模态大语言模型在真实图表理解中的差距》。arXiv 预印本 arXiv:2406.18521。引用自:§5.1。
- J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y. Fan, K. Dang 等人 (2025a) Qwen2.5-omni 技术报告。arXiv 预印本 arXiv:2503.20215。引用自:§1, §1, §2.1, 表 8。
- J. Xu, Z. Guo, H. Hu, Y. Chu, X. Wang, J. He, Y. Wang, X. Shi, T. He, X. Zhu, Y. Lv, Y. Wang, D. Guo, H. Wang, L. Ma, P. Zhang, X. Zhang, H. Hao, Z. Guo, B. Yang, B. Zhang, Z. Ma, X. Wei, S. Bai, K. Chen, X. L. Liu, P. Wang, M. Yang, D. Liu, X. Ren, B. Zheng, R. Men, F. Zhou, B. Yu, J. Yang, L. Yu, J. Zhou, 和 J. Lin (2025b) Qwen3-omni 技术报告。ArXiv 摘要 arXiv:2509.17765。引用自:§1, §1, 第 3 项, 第 5 项, §2.1, §2.3, §2.4, §3, §3, 表 8。
- F. Yan, H. Mao, C. C. Ji, T. Zhang, S. G. Patil, I. Stoica, 和 J. E. Gonzalez (2024) Berkeley 函数调用排行榜。备注:https://gorilla.cs.berkeley.edu/blogs/8_berkeley_function_calling_leaderboard.html 引用自:§5.1。
- R. Yan, X. Li, W. Chen, Z. Niu, C. Yang, Z. Ma, K. Yu, 和 X. Chen (2025) URO-bench:面向端到端语音对话模型的综合基准测试。arXiv 预印本 arXiv:2502.17810。引用自:§5.1。
- A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv 等人 (2025a) Qwen3 技术报告。arXiv 预印本 arXiv:2505.09388。引用自:§1。
- A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang 等人 (2024a) Qwen2 技术报告。arXiv:2407.10671。引用自:§1。
- Y. Yang, J. Zhuang, G. Sun, C. Tang, Y. Li, P. Li, Y. Jiang, W. Li, Z. Ma, 和 C. Zhang (2025b) 无文本捷径的以音频为中心的视频理解基准测试。载于《2025 年自然语言处理经验方法会议论文集》,第 6580–6598 页。引用自:§5.1。
- Z. Yang, J. Tang, Z. Li, P. Wang, J. Wan, H. Zhong, X. Liu, M. Yang, P. Wang, S. Bai, L. Jin, 和 J. Lin (2024b) CC-OCR:用于评估大型多模态模型识字能力的全面且具有挑战性的 OCR 基准测试。CoRR 摘要 arXiv:2412.02210。引用自:§5.1。
- X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun 等人 (2023) MMMU:面向专家级 AGI 的大规模多学科多模态理解与推理基准测试。arXiv:2311.16502。引用自:§5.1。
- X. Yue、T. Zheng、Y. Ni、Y. Wang、K. Zhang、S. Tong、Y. Sun、M. Yin、B. Yu、G. Zhang 等人(2024)《MMMU-pro:一个更稳健的多学科多模态理解基准》。arXiv 预印本 arXiv:2409.02813。引用自:§5.1。
- Y. Zang、S. O’Brien、T. Berg-Kirkpatrick、J. McAuley 和 Z. Novack(2025)《你真的在听吗?提升音乐问答基准中的感知意识》。arXiv 预印本 arXiv:2504.00369。引用自:§5.1。
- B. Zhang、H. Lv、P. Guo、Q. Shao、C. Yang、L. Xie、X. Xu、H. Bu、X. Chen、C. Zeng、D. Wu 和 Z. Peng(2022)《WENETSPEECH:一个 10000+ 小时的多领域中文语音识别语料库》。收录于:IEEE 国际声学、语音与信号处理会议(ICASSP 2022),虚拟会议及新加坡,2022 年 5 月 23-27 日,第 6182–6186 页。外部链接:Link,DOI。引用自:§5.1。
- B. Zhang、C. Guo、G. Yang、H. Yu、H. Zhang、H. Lei、J. Mai、J. Yan、K. Yang、M. Yang、P. Huang、R. Jin、S. Jiang、W. Cheng、Y. Li、Y. Xiao、Y. Zhou、Y. Zhang、Y. Lu 和 Y. He(2025a)《MiniMax-Speech:使用可学习说话人编码器的内在零样本文本转语音》。CoRR 摘要 arXiv:2505.07916。引用自:第 2 项、第 4 项、表 8。
- L. Zhang、J. Zhang、B. Lei、C. Wu、A. Liu、W. Jia 和 X. Zhou(2025b)《WildSpeech-Bench:在真实场景中基准测试端到端语音大语言模型》。外部链接:2506.21875。引用自:§5.1。
- X. Zhang、C. Wu、Z. Zhao、W. Lin、Y. Zhang、Y. Wang 和 W. Xie(2023)《PMC-VQA:面向医学视觉问答的视觉指令微调》。CoRR 摘要 arXiv:2305.10415。引用自:§5.1。
- X. L. T. D. Zhang, G. Wang, J. Xue, K. Fang, L. Zhao, R. Ma, S. Ren, S. Liu, T. Guo, W. Zhuang, X. Zhang, X. Song, Y. Yan, Y. He, Cici, B. Shen, C. Zhu, C. Ma, C. Chen, H. Chen, J. Li, L. Li, M. Zhu, P. Li, Q. Wang, S. Deng, W. Xiong, W. Huang, W. Yang, Y. Jiang, Y. Yang, Y. Tian, Y. Ma, Y. Yu, Z. Zhang, Z. Yue, B. Xiao, B. Xia, B. Gao, B. Ye, C. Cai, C. Liu, C. He, C. Li, D. Zhu, D. Zhang, F. Shi, G. Wang, H. Zhang, H. Lv, H. Li, H. Tian, H. Qu, H. Xu, H. Zhang, H. Liu, J. Duo, J. Zuo, J. Wei, J. Xiao, J. Dong, J. Shi, J. Hu, K. Bao, K. Zhou, L. Zhang, M. Chen, N. Chen, P. Zhang, Q. Chen, Q. Wang, R. Li, S. Liu, S. Wang, S. Li, S. Yu, S. Cao, S. Chen, S. Gu, W. Wang, W. Ma, X. Deng, X. Yong, X. Zhang, X. Wang, Y. Song, Y. Zhao, Y. Zhao, Y. Gao, Y. Cheng, Y. Tu, Y. Wang, Z. Huang, Z. Tang, Z. Lin, Z. Song, Z. Xu, Z. Zheng, 和 Z. Jiang (2025c) MiMo-audio: 音频语言模型是少样本学习者。ArXiv abs/2512.23808。外部链接:Link。被表 8 引用。
- Y. Zhang, H. Zhang, H. Tian, C. Fu, S. Zhang, J. Wu, F. Li, K. Wang, Q. Wen, Z. Zhang, 等人 (2024) MME-realworld: 你的多模态大语言模型能否挑战人类也难以应对的高分辨率真实世界场景?arXiv 预印本 arXiv:2408.13257。被 §5.1 引用。
- Y. Zhao, H. Zhang, L. Xie, T. Hu, G. Gan, Y. Long, Z. Hu, W. Chen, C. Li, Z. Xu, C. Wang, Z. Shangguan, Z. Liang, Y. Liu, C. Zhao, 和 A. Cohan (2025) MMVU: 衡量专家级多学科视频理解。收录于 IEEE/CVF 计算机视觉与模式识别会议,CVPR 2025,美国田纳西州纳什维尔,2025 年 6 月 11-15 日,第 8475–8489 页。被 §5.1 引用。
- C. Zheng, S. Liu, M. Li, X. Chen, B. Yu, C. Gao, K. Dang, Y. Liu, R. Men, A. Yang, 等人 (2025) 分组序列策略优化。arXiv 预印本 arXiv:2507.18071。被第 (3) 项引用。
- J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, 和 L. Hou (2023) 大语言模型的指令遵循评估。CoRR abs/2311.07911。被 §5.1 引用。
- J. Zhou, Y. Shu, B. Zhao, B. Wu, Z. Liang, S. Xiao, M. Qin, X. Yang, Y. Xiong, B. Zhang, T. Huang, 和 Z. Liu (2025a) 《MLVU:面向多任务长视频理解的基准测试》。收录于 IEEE/CVF 计算机视觉与模式识别会议,CVPR 2025,美国田纳西州纳什维尔,2025年6月11-15日,第13691–13701页。引用自:§5.1。
- Z. Zhou, R. Wang, 和 Z. Wu (2025b) 《Daily-omni:面向跨模态时间对齐的视听推理》。CoRR abs/2505.17862。引用自:§5.1。
- D. Zhu, J. Chen, X. Shen, X. Li, 和 M. Elhoseiny (2023) 《Minigpt-4:利用先进大语言模型增强视觉语言理解》。arXiv:2304.10592。引用自:§1。
- C. Zou, X. Guo, R. Yang, J. Zhang, B. Hu, 和 H. Zhang (2025) 《DynaMath:用于评估视觉语言模型数学推理鲁棒性的动态视觉基准》。收录于 ICLR。引用自:§5.1。
- Y. Zuo, S. Qu, Y. Li, Z. Chen, X. Zhu, E. Hua, K. Zhang, N. Ding, 和 B. Zhou (2025) 《MedXpertQA:面向专家级医学推理与理解的基准测试》。收录于 ICML,A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, 和 J. Zhu 编,《机器学习研究论文集》。引用自:§5.1。
7 位作者
核心贡献者²²²按字母顺序排列。* 表示通讯作者。
韩冰 包松森 张斌 郑波 刘大一 周帆 郝鸿坤 胡航宇 徐进* 简新阳 周靖人 陈克勤 于磊 杨明坤 王鹏 张培 杨启泽 门睿 徐瑞阳 白帅 宋思博 丁鹤 程希泽 刘雪静 任兴章 石贤 王雄 张新宇 朱新发 楚云飞 吕元军 孙宇聪 王永琪 王宇轩 张洋 郭志方 郭子山 马子阳
贡献者††footnotemark:
安东·陈 安福·李 安·杨 北·陈 斌·林 炳申·穆 博涵·王 卜晓·吴 博文·徐 北晨·张 程·陈 昌·高 程根·黄 辰阳·乐 晨浩·李 成龙·刘 晨旭·吕 辰·强 陈飞·吴 辰汉·袁 程瑞东·张 楚杰·郑 大任·陈 大可·郭 飞·黄 高吉·刘 广东·周 浩·葛 慧强·江 浩然·连 鸿建·涂 浩·于 航·张 浩·周 海泉·赵 胡门·钟 嘉伟·陈 健·关 嘉怡·冷 嘉豪·李 俊荣·林 嘉伟·刘 嘉龙·唐 俊·唐 建宏·涂 建强·万 金伟·魏 建伟·张 静·周 凯·党 康祥·夏 坤·颜 可欣·杨 良浩·邓 露露·胡 林涵·马 令晨·孟 磊·谢 来文·郑 苗·洪 梅·李 明成·李 明泽·李 民生·李 明浩·吴 明峰·薛 娜·倪 鹏·刘 鹏·王 鹏飞·王 培阳·张 启东·黄 庆峰·蓝 勤通·李 阙·沈 秋月·王 勤·朱 瑞生·曹 荣耀·方 瑞·胡 瑞斌·袁 松·陈 苏·郝 申力·李 世轩·刘 树瑞·李 思齐·张 天一·唐 婷玉·夏 伟·丁 文彬·葛 伟洲·沈 伟·王 文涛·姚 曦·陈 晓彤·陈 雄辉·陈 晓东·邓 旭东·郭 欣·乐 晓丽·谢 陈 欣瑶·牛 宣成·任 雪纯·王 旭武·王 兴哲·吴 希品·魏 晓·徐 先·杨 宇轩·蔡 一中·曹 一磊·陈 宇翔·陈 一鸣·董 阳·范 彦鹏·李 宇成·李 阳·刘 彦涛·刘 宇琼·刘 宇轩·刘 雨妍·罗 宇博·马 阳·苏 月璋·王 宇浩·王 一·吴 云宝·吴 宇·席 一·张 一昌·张 英儿·张 宇翔·郑 泽宇·崔 子维·季 子越·江 兆海·李 正·李 志·李 子涵·邱 泽坤·王 志海·王 正浩·邢 志波·杨 卓瑞·叶 振如·张 志鹏·周 正阳·诸葛
8 附录
8.1 多语言评估详细结果
多语言自动语音识别
如表13所示,Qwen3.5-Omni在FLEURS测试集上展现出优于最先进竞品的语音识别能力。Qwen3.5-Omni-Plus实现了6.6%的最低平均词错误率,优于Gemini-3.1-Pro(7.3%)和GPT-4o-Transcribe(10.4%)。它在大多数语言中取得最佳表现,在粤语(2.2%对比Gemini-3.1-Pro的6.3%)、泰语和越南语等复杂声调语言及低资源语言中优势尤为显著。同时,Qwen3.5-Omni-Flash提供了高效替代方案,其10.8%的平均词错误率与Gemini-3-Flash(10.5%)相比仍具竞争力。值得注意的是,Qwen3.5-Omni-Flash在挑战性场景中展现出卓越鲁棒性,大幅降低了粤语错误率(3.1%对比Gemini-3-Flash的10.8%),并在日语和韩语中保持强劲表现,凸显了其在高价值亚洲语言对中的优势。
多语言翻译。
如表14和表15所示,Qwen3.5-Omni系列在FLEURS测试集上展现出相较于最先进竞品的显著优势,尤其在亚洲语言和特定高资源语言对中。Qwen3.5-Omni-Plus在多对多方向(en2xx/zh2xx)上全面优于Gemini-3.1-Pro,在英语到XX(33.8对比31.8)和中文到XX(21.4对比19.6)方向均取得更高平均BLEU分数。在葡萄牙语(49.4对比47.7)和印尼语(45.7对比45.1)等关键xx2en语言对中也保持领先。尽管Gemini-3.1-Pro在xx2zh整体平均值上略占优势,但Qwen3.5-Omni-Plus在粤语(+15.6 BLEU)、韩语和日语等关键亚洲语言上显著超越对手。类似地,Qwen3.5-Omni-Flash针对Gemini-3-Flash展现出定向优势。在保持整体性能竞争力的同时,它在所有方向的粤语翻译中大幅超越Gemini(例如xx2zh方向37.5对比22.4,en2xx方向37.3对比26.7),并在日语和韩语的xx2zh任务中取得更优结果。这些结果凸显了Qwen3.5-Omni针对复杂亚洲语言结构和关键区域语言的深度优化能力。
| 语言 |
|
|
|
|
| ||||||||||
| 中文 | 2.9 | 2.9 | 3.6 | 2.6 | 4.6 | ||||||||||
| 英语 | 3.2 | 3.7 | 2.7 | 3.2 | 2.9 | ||||||||||
| 粤语 | 2.2 | 3.1 | 6.3 | 5.2 | 10.8 | ||||||||||
| 阿拉伯语 | 11.7 | 13.6 | 9.2 | 13.0 | 10.1 | ||||||||||
| 德语 | 2.0 | 2.5 | 2.7 | 2.3 | 3.3 | ||||||||||
| 法语 | 2.6 | 3.3 | 3.7 | 3.7 | 3.9 | ||||||||||
| 西班牙语 | 2.2 | 2.4 | 2.5 | 2.3 | 2.7 | ||||||||||
| 葡萄牙语 | 2.1 | 2.2 | 2.6 | 2.3 | 2.9 | ||||||||||
| 印尼语 | 1.6 | 2.4 | 2.5 | 3.5 | 2.8 | ||||||||||
| 意大利语 | 0.8 | 1.0 | 1.1 | 1.4 | 1.8 | ||||||||||
| 韩语 | 1.7 | 2.1 | 2.0 | 2.1 | 2.4 | ||||||||||
| 俄语 | 3.1 | 3.6 | 3.4 | 3.7 | 3.9 | ||||||||||
| 泰语 | 2.8 | 3.2 | 4.3 | 4.9 | 4.5 | ||||||||||
| 越南语 | 1.9 | 2.5 | 2.5 | 3.5 | 3.5 | ||||||||||
| 日语 | 1.9 | 2.5 | 2.3 | 3.0 | 3.4 | ||||||||||
| 土耳其语 | 3.1 | 4.4 | 3.8 | 4.2 | 4.4 | ||||||||||
| 印地语 | 9.7 | 9.9 | 4.5 | 12.0 | 5.6 | ||||||||||
| 马来语 | 2.7 | 4.2 | 3.7 | 4.1 | 6.2 | ||||||||||
| 荷兰语 | 2.8 | 3.5 | 3.5 | 3.7 | 4.7 | ||||||||||
| 乌尔都语 | 20.8 | 31.9 | 25.2 | 19.7 | 23.0 | ||||||||||
| 挪威语 | 3.9 | 5.2 | 5.0 | 5.5 | 6.7 | ||||||||||
| 瑞典语 | 3.1 | 5.0 | 4.6 | 5.2 | 7.7 | ||||||||||
| 丹麦语 | 3.5 | 5.3 | 5.7 | 6.5 | 7.9 | ||||||||||
| 希伯来语 | 12.5 | 16.6 | 15.6 | 19.4 | 20.2 | ||||||||||
| 芬兰语 | 2.4 | 4.5 | 3.4 | 3.8 | 5.2 | ||||||||||
| 波兰语 | 1.9 | 3.1 | 2.7 | 2.8 | 4.9 | ||||||||||
| 冰岛语 | 3.6 | 8.9 | 4.7 | 10.8 | 6.8 | ||||||||||
| 捷克语 | 2.6 | 4.5 | 3.8 | 4.7 | 8.0 | ||||||||||
| 菲律宾语 | 5.1 | 7.1 | 7.6 | 7.3 | 8.5 | ||||||||||
| 波斯语 | 12.0 | 12.1 | 8.9 | 9.9 | 10.0 | ||||||||||
| 希腊语 | 4.7 | 8.1 | 5.4 | 6.5 | 7.6 | ||||||||||
| 南非荷兰语 | 10.6 | 13.7 | 12.7 | 17.9 | 18.6 | ||||||||||
| 阿斯图里亚斯语 | 15.8 | 25.9 | 23.7 | 23.8 | 48.3 | ||||||||||
| 白俄罗斯语 | 6.7 | 12.2 | 6.7 | 10.2 | 10.9 | ||||||||||
| 保加利亚语 | 6.2 | 10.7 | 5.3 | 7.0 | 7.9 | ||||||||||
| 孟加拉语 | 16.2 | 19.8 | 21.9 | 24.1 | 21.9 | ||||||||||
| 波斯尼亚语 | 5.4 | 9.5 | 6.0 | 13.9 | 11.2 | ||||||||||
| 加泰罗尼亚语 | 2.8 | 6.3 | 2.7 | 2.7 | 6.4 | ||||||||||
| 宿务语 | 10.5 | 16.6 | 13.0 | 15.1 | 12.8 | ||||||||||
| 爱沙尼亚语 | 6.7 | 16.6 | 4.9 | 7.6 | 7.3 | ||||||||||
| 加利西亚语 | 5.0 | 8.6 | 4.9 | 6.8 | 14.6 | ||||||||||
| 古吉拉特语 | 13.9 | 18.4 | 14.8 | 26.9 | 16.3 | ||||||||||
| 克罗地亚语 | 5.4 | 9.0 | 5.1 | 16.4 | 9.3 | ||||||||||
| 匈牙利语 | 4.9 | 10.6 | 5.5 | 7.4 | 11.3 | ||||||||||
| 爪哇语 | 11.8 | 18.3 | 14.1 | 24.7 | 15.7 | ||||||||||
| 哈萨克语 | 6.3 | 16.6 | 6.2 | 11.5 | 12.7 | ||||||||||
| 卡纳达语 | 16.0 | 23.8 | 16.3 | 28.1 | 16.3 | ||||||||||
| 吉尔吉斯语 | 10.0 | 19.7 | 8.3 | 20.7 | 16.3 | ||||||||||
| 拉脱维亚语 | 6.7 | 17.8 | 3.7 | 6.3 | 6.8 | ||||||||||
| 马其顿语 | 4.1 | 7.9 | 4.0 | 6.2 | 7.8 | ||||||||||
| 马拉雅拉姆语 | 18.8 | 27.0 | 18.3 | 33.5 | 20.3 | ||||||||||
| 马拉地语 | 16.3 | 23.6 | 15.3 | 26.3 | 16.0 | ||||||||||
| 旁遮普语 | 13.7 | 24.6 | 14.4 | 36.4 | 17.4 | ||||||||||
| 罗马尼亚语 | 3.2 | 6.1 | 3.4 | 4.5 | 5.9 | ||||||||||
| 斯洛伐克语 | 3.3 | 5.5 | 2.8 | 3.6 | 6.5 | ||||||||||
| 斯洛文尼亚语 | 6.1 | 14.3 | 6.3 | 8.8 | 10.0 | ||||||||||
| 斯瓦希里语 | 9.4 | 17.5 | 9.9 | 16.3 | 10.7 | ||||||||||
| 塔吉克语 | 10.0 | 41.1 | 20.4 | 20.2 | 53.1 | ||||||||||
| 阿塞拜疆语 | 7.2 | 13.0 | 5.8 | 10.6 | 13.2 | ||||||||||
| 乌克兰语 | 3.2 | 5.4 | 3.4 | 4.3 | 5.1 | ||||||||||
| 平均 | 6.6 | 10.8 | 7.3 | 10.4 | 10.5 |
| en2xx(英语 → 其他语言) | zh2xx(中文 → 其他语言) | |||||||||||||||||||||||
| 语言 |
|
|
|
|
|
|
|
| ||||||||||||||||
| 中文 | 47.8 | 46.6 | 47.4 | 46.3 | – | – | – | – | ||||||||||||||||
| 英语 | – | – | – | – | 32.2 | 31.2 | 30.1 | 29.5 | ||||||||||||||||
| 粤语 | 40.1 | 37.3 | 25.5 | 26.7 | 36.7 | 35.9 | 23.7 | 24.0 | ||||||||||||||||
| 阿拉伯语 | 31.1 | 28.2 | 27.0 | 28.5 | 16.1 | 13.9 | 14.2 | 14.4 | ||||||||||||||||
| 德语 | 43.2 | 39.6 | 41.8 | 40.9 | 23.2 | 20.8 | 22.0 | 21.4 | ||||||||||||||||
| 法语 | 50.9 | 48.8 | 48.0 | 47.8 | 30.7 | 28.8 | 29.4 | 29.2 | ||||||||||||||||
| 西班牙语 | 29.1 | 28.9 | 28.3 | 28.8 | 22.2 | 20.4 | 20.8 | 20.6 | ||||||||||||||||
| 葡萄牙语 | 51.2 | 48.6 | 47.3 | 47.2 | 28.5 | 26.8 | 25.7 | 25.6 | ||||||||||||||||
| 印度尼西亚语 | 45.3 | 43.7 | 42.1 | 41.7 | 28.8 | 26.9 | 25.4 | 25.2 | ||||||||||||||||
| 意大利语 | 32.7 | 30.7 | 31.9 | 30.9 | 23.1 | 21.1 | 21.9 | 21.3 | ||||||||||||||||
| 韩语 | 33.9 | 31.8 | 30.8 | 31.7 | 25.1 | 23.4 | 21.9 | 22.8 | ||||||||||||||||
| 俄语 | 33.8 | 31.8 | 33.1 | 33.2 | 21.5 | 18.9 | 20.0 | 19.9 | ||||||||||||||||
| 泰语 | 65.4 | 62.9 | 64.8 | 64.4 | 58.0 | 55.5 | 57.2 | 56.4 | ||||||||||||||||
| 越南语 | 43.0 | 41.8 | 41.1 | 40.2 | 31.6 | 30.5 | 28.1 | 28.4 | ||||||||||||||||
| 日语 | 53.2 | 50.6 | 51.3 | 50.8 | 45.6 | 41.6 | 43.0 | 42.0 | ||||||||||||||||
| 土耳其语 | 30.4 | 27.6 | 29.3 | 29.0 | 16.8 | 14.6 | 15.9 | 16.0 | ||||||||||||||||
| 印地语 | 33.1 | 29.1 | 28.2 | 29.2 | 19.1 | 14.3 | 17.6 | 17.5 | ||||||||||||||||
| 马来语 | 39.6 | 37.2 | 35.9 | 36.0 | 24.1 | 21.7 | 21.0 | 20.5 | ||||||||||||||||
| 荷兰语 | 30.1 | 28.2 | 28.1 | 28.8 | 21.0 | 18.8 | 19.3 | 19.1 | ||||||||||||||||
| 乌尔都语 | 25.0 | 22.1 | 23.0 | 23.0 | 15.5 | 8.6 | 14.9 | 14.7 | ||||||||||||||||
| 挪威语 | 35.3 | 32.8 | 33.1 | 33.7 | 20.3 | 17.8 | 18.4 | 18.9 | ||||||||||||||||
| 瑞典语 | 47.5 | 44.1 | 45.7 | 45.8 | 25.4 | 23.0 | 23.7 | 24.1 | ||||||||||||||||
| 丹麦语 | 48.4 | 45.2 | 45.7 | 45.2 | 25.7 | 22.8 | 23.5 | 23.2 | ||||||||||||||||
| 希伯来语 | 36.4 | 29.9 | 36.5 | 35.4 | 18.2 | 14.5 | 17.7 | 17.4 | ||||||||||||||||
| 芬兰语 | 30.1 | 26.0 | 32.1 | 32.3 | 18.1 | 15.3 | 18.6 | 17.7 | ||||||||||||||||
| 波兰语 | 25.2 | 22.5 | 24.9 | 23.4 | 17.5 | 15.0 | 15.6 | 15.5 | ||||||||||||||||
| 冰岛语 | 28.5 | 27.2 | 29.6 | 28.2 | 16.2 | 13.5 | 16.0 | 15.6 | ||||||||||||||||
| 捷克语 | 35.9 | 32.5 | 33.3 | 33.7 | 20.4 | 18.1 | 19.4 | 19.0 | ||||||||||||||||
| 菲律宾语 | 35.0 | 32.0 | 32.1 | 33.1 | 22.3 | 19.0 | 20.7 | 20.7 | ||||||||||||||||
| 波斯语 | 30.7 | 27.3 | 25.1 | 25.9 | 19.5 | 16.4 | 15.9 | 15.9 | ||||||||||||||||
| 希腊语 | 30.0 | 27.8 | 30.0 | 29.4 | 18.4 | 15.9 | 17.6 | 17.5 | ||||||||||||||||
| 阿斯图里亚斯语 | 32.4 | 27.9 | 31.5 | 30.4 | 20.4 | 16.2 | 18.6 | 18.1 | ||||||||||||||||
| 白俄罗斯语 | 16.4 | 14.7 | 16.4 | 16.5 | 12.6 | 10.8 | 12.1 | 12.2 | ||||||||||||||||
| 保加利亚语 | 45.0 | 40.7 | 41.7 | 42.5 | 25.6 | 23.0 | 24.3 | 24.0 | ||||||||||||||||
| 孟加拉语 | 18.6 | 15.7 | 14.3 | 15.0 | 10.6 | 9.2 | 9.0 | 9.3 | ||||||||||||||||
| 波斯尼亚语 | 37.5 | 34.0 | 36.3 | 35.2 | 21.4 | 18.6 | 19.9 | 19.4 | ||||||||||||||||
| 加泰罗尼亚语 | 43.9 | 41.5 | 42.7 | 42.9 | 26.6 | 17.2 | 25.0 | 25.1 | ||||||||||||||||
| 宿务语 | 28.5 | 12.7 | 28.5 | 29.2 | 19.0 | 5.6 | 17.5 | 17.7 | ||||||||||||||||
| 爱沙尼亚语 | 30.8 | 26.3 | 31.4 | 30.6 | 18.9 | 13.6 | 17.5 | 17.2 | ||||||||||||||||
| 加利西亚语 | 37.4 | 35.4 | 36.6 | 35.9 | 23.9 | 22.0 | 22.7 | 22.3 | ||||||||||||||||
| 古吉拉特语 | 23.8 | 20.9 | 21.0 | 21.5 | 14.3 | 10.8 | 12.3 | 12.3 | ||||||||||||||||
| 克罗地亚语 | 33.3 | 30.7 | 33.4 | 32.6 | 21.3 | 18.5 | 19.2 | 18.6 | ||||||||||||||||
| 匈牙利语 | 29.5 | 24.9 | 28.1 | 27.6 | 18.8 | 15.8 | 17.3 | 16.7 | ||||||||||||||||
| 爪哇语 <<<187 | 26.8 | 24.4 | 16.4 | 22.1 | 16.5 | 14.8 | 9.0 | 13.0 | ||||||||||||||||
| Kazakh | 24.9 | 21.1 | 20.6 | 22.5 | 15.0 | 12.4 | 12.9 | 13.1 | ||||||||||||||||
| 卡纳达语 | 20.0 | 17.0 | 16.1 | 17.2 | 11.7 | 6.9 | 9.5 | 10.0 | ||||||||||||||||
| 吉尔吉斯语 | 15.3 | 12.6 | 15.1 | 14.7 | 10.4 | 7.8 | 10.0 | 9.5 | ||||||||||||||||
| 拉脱维亚语 | 36.1 | 31.0 | 35.4 | 35.3 | 21.9 | 17.8 | 20.1 | 19.5 | ||||||||||||||||
| 马其顿语 | 38.1 | 34.0 | 38.5 | 38.2 | 22.3 | 20.1 | 22.0 | 21.6 | ||||||||||||||||
| 马拉雅拉姆语 | 19.3 | 11.1 | 16.1 | 16.1 | 10.4 | 5.2 | 9.9 | 9.4 | ||||||||||||||||
| 马拉地语 | 17.7 | 11.6 | 15.9 | 16.2 | 11.3 | 8.1 | 9.4 | 10.1 | ||||||||||||||||
| 旁遮普语 | 26.1 | 23.1 | 24.1 | 24.6 | 15.7 | 8.9 | 14.3 | 14.1 | ||||||||||||||||
| 罗马尼亚语 | 42.0 | 39.9 | 41.4 | 42.0 | 25.6 | 22.8 | 23.6 | 23.4 | ||||||||||||||||
| 斯洛伐克语 | 35.3 | 31.4 | 34.8 | 34.6 | 19.6 | 16.4 | 19.2 | 18.5 | ||||||||||||||||
| 斯洛文尼亚语 | 32.8 | 28.5 | 33.5 | 32.9 | 20.1 | 17.7 | 20.9 | 20.2 | ||||||||||||||||
| 斯瓦希里语 | 36.3 | 30.9 | 32.1 | 32.2 | 20.4 | 9.3 | 18.5 | 18.3 | ||||||||||||||||
| 塔吉克语 | 23.8 | 18.3 | 22.1 | 22.9 | 14.6 | 10.9 | 14.3 | 14.2 | ||||||||||||||||
| 阿塞拜疆语 | 13.7 | 9.8 | 15.2 | 14.7 | 11.5 | 9.7 | 11.3 | 10.8 | ||||||||||||||||
| 乌克兰语 | 31.7 | 29.2 | 29.8 | 30.1 | 19.5 | 15.0 | 17.0 | 17.6 | ||||||||||||||||
| 平均 | 33.8 | 30.4 | 31.8 | 31.8 | 21.4 | 18.1 | 19.6 | 19.5 | ||||||||||||||||
| xx2en(其他语言 → 英语) | xx2zh(其他语言 → 中文) | |||||||||||||||||||||||
| 语言 |
|
|
|
|
|
|
|
| ||||||||||||||||
| 中文 | 32.2 | 31.2 | 30.1 | 29.5 | – | – | – | – | ||||||||||||||||
| 英语 | – | – | – | – | 47.8 | 46.6 | 47.4 | 46.3 | ||||||||||||||||
| 粤语 | 30.3 | 29.9 | 27.7 | 26.3 | 36.8 | 37.5 | 21.2 | 22.4 | ||||||||||||||||
| 阿拉伯语 | 42.9 | 40.0 | 42.1 | 42.3 | 40.2 | 37.3 | 40.5 | 40.5 | ||||||||||||||||
| 德语 | 44.6 | 44.1 | 43.9 | 43.9 | 43.3 | 42.8 | 42.1 | 41.8 | ||||||||||||||||
| 法语 | 43.5 | 42.0 | 41.6 | 41.3 | 41.6 | 40.3 | 41.6 | 41.4 | ||||||||||||||||
| 西班牙语 | 32.3 | 31.3 | 30.3 | 30.4 | 38.8 | 38.5 | 38.4 | 38.2 | ||||||||||||||||
| 葡萄牙语 | 49.4 | 48.2 | 47.7 | 47.5 | 43.6 | 41.8 | 42.5 | 42.4 | ||||||||||||||||
| 印尼语 | 45.7 | 43.1 | 45.1 | 44.9 | 43.5 | 41.4 | 42.8 | 42.9 | ||||||||||||||||
| 意大利语 | 34.4 | 31.9 | 30.8 | 31.0 | 40.8 | 39.5 | 39.7 | 39.4 | ||||||||||||||||
| 韩语 | 34.1 | 32.4 | 32.0 | 32.1 | 39.9 | 37.5 | 37.0 | 37.2 | ||||||||||||||||
| 俄语 | 38.6 | 37.2 | 36.7 | 36.5 | 41.7 | 39.5 | 41.1 | 40.4 | ||||||||||||||||
| 泰语 | 34.1 | 32.4 | 34.2 | 33.0 | 40.2 | 37.9 | 40.0 | 39.5 | ||||||||||||||||
| 越南语 | 36.4 | 34.9 | 36.1 | 35.4 | 38.7 | 36.3 | 39.2 | 38.9 | ||||||||||||||||
| 日语 | 30.4 | 29.2 | 29.5 | 29.4 | 38.0 | 35.7 | 35.6 | 34.8 | ||||||||||||||||
| 土耳其语 | 40.3 | 39.1 | 39.5 | 38.9 | 41.8 | 40.3 | 40.6 | 41.0 | ||||||||||||||||
| 印地语 | 38.8 | 36.2 | 39.3 | 39.2 | 38.9 | 36.9 | 38.0 | 38.4 | ||||||||||||||||
| 马来语 | 42.9 | 41.1 | 44.8 | 42.4 | 41.0 | 39.5 | 42.4 | 41.5 | ||||||||||||||||
| 荷兰语 | 33.3 | 32.2 | 31.2 | 30.4 | 40.1 | 38.5 | 39.4 | 39.4 | ||||||||||||||||
| 乌尔都语 | 35.5 | 31.5 | 32.9 | 32.6 | 37.2 | 33.8 | 36.9 | 36.5 | ||||||||||||||||
| 挪威语 | 43.5 | 42.1 | 42.6 | 41.4 | 42.2 | 39.6 | 40.9 | 40.8 | ||||||||||||||||
| 瑞典语 | 47.2 | 45.2 | 46.6 | 44.7 | 42.9 | 40.7 | 41.5 | 41.6 | ||||||||||||||||
| 丹麦语 | 45.4 | 44.0 | 44.7 | 42.7 | 43.4 | 41.1 | 41.7 | 40.6 | ||||||||||||||||
| 希伯来语 | 39.7 | 36.4 | 42.5 | 39.9 | 36.7 | 34.1 | 40.3 | 37.8 | ||||||||||||||||
| 芬兰语 | 36.9 | 35.0 | 36.7 | 35.5 | 40.8 | 38.7 | 41.0 | 40.6 | ||||||||||||||||
| 波兰语 | 32.1 | 30.5 | 30.4 | 29.7 | 38.4 | 36.0 | 37.9 | 37.4 | ||||||||||||||||
| 冰岛语 | 31.5 | 27.5 | 35.1 | 35.8 | 38.2 | 31.8 | 37.9 | 37.5 | ||||||||||||||||
| 捷克语 | 42.1 | 39.3 | 40.1 | 39.5 | 40.6 | 39.4 | 40.6 | 40.0 | ||||||||||||||||
| 菲律宾语 | 42.7 | 40.9 <<<230 | 45.0 | 44.3 | 41.0 | 38.0 | 42.6 | 41.5 | ||||||||||||||||
| Persian | 40.2 | 36.8 | 38.0 | 38.1 | 40.2 | 37.0 | 41.1 | 41.1 | ||||||||||||||||
| Greek | 36.0 | 32.5 | 35.3 | 35.4 | 38.0 | 32.8 | 39.0 | 38.4 | ||||||||||||||||
| Asturian | 37.0 | 35.1 | 37.2 | 35.5 | 37.7 | 34.0 | 38.1 | 36.7 | ||||||||||||||||
| Belarusian | 23.1 | 19.9 | 20.6 | 19.9 | 33.3 | 31.2 | 33.7 | 33.7 | ||||||||||||||||
| 保加利亚语 | 39.6 | 36.0 | 41.3 | 40.0 | 40.9 | 36.6 | 41.0 | 40.4 | ||||||||||||||||
| 孟加拉语 | 32.0 | 27.6 | 34.9 | 34.3 | 35.8 | 32.4 | 38.1 | 37.5 | ||||||||||||||||
| 波斯尼亚语 | 43.1 | 40.8 | 42.7 | 42.4 | 41.5 | 39.1 | 42.3 | 41.4 | ||||||||||||||||
| 加泰罗尼亚语 | 46.6 | 42.3 | 46.2 | 45.5 | 42.2 | 38.9 | 42.6 | 41.5 | ||||||||||||||||
| 宿务语 | 37.3 | 26.3 | 38.9 | 38.2 | 34.0 | 26.5 | 36.7 | 36.6 | ||||||||||||||||
| 爱沙尼亚语 | 35.7 | 28.3 | 40.1 | 38.3 | 38.0 | 32.2 | 41.7 | 41.2 | ||||||||||||||||
| 加利西亚语 | 40.8 | 38.6 | 39.6 | 38.6 | 40.9 | 39.1 | 41.6 | 40.9 | ||||||||||||||||
| 古吉拉特语 | 33.4 | 28.3 | 40.3 | 39.7 | 35.8 | 31.4 | 39.7 | 39.0 | ||||||||||||||||
| 克罗地亚语 | 39.5 | 36.3 | 38.0 | 36.9 | 40.0 | 37.7 | 40.2 | 39.6 | ||||||||||||||||
| 匈牙利语 | 35.5 | 29.6 | 35.3 | 33.3 | 39.1 | 33.9 | 40.0 | 37.4 | ||||||||||||||||
| 爪哇语 | 35.9 | 28.4 | 38.4 | 36.5 | 34.9 | 28.3 | 36.7 | 36.3 | ||||||||||||||||
| 哈萨克语 | 34.4 | 26.9 | 35.7 | 35.5 | 37.1 | 31.4 | 39.3 | 38.8 | ||||||||||||||||
| 卡纳达语 | 26.4 | 19.8 | 33.1 | 33.6 | 32.3 | 26.1 | 38.0 | 37.8 | ||||||||||||||||
| 吉尔吉斯语 | 22.2 | 17.1 | 24.8 | 23.7 | 29.9 | 24.8 | 33.7 | 32.8 | ||||||||||||||||
| 拉脱维亚语 | 33.7 | 25.2 | 38.0 | 37.9 | 37.1 | 30.0 | 41.4 | 40.6 | ||||||||||||||||
| 马其顿语 | 43.3 | 39.6 | 43.1 | 41.9 | 41.6 | 38.0 | 42.1 | 41.8 | ||||||||||||||||
| 马拉雅拉姆语 | 31.2 | 25.7 | 34.9 | 34.2 | 36.1 | 31.5 | 38.4 | 38.0 | ||||||||||||||||
| 马拉地语 | 33.5 | 25.7 | 36.7 | 35.5 | 34.6 | 29.4 | 38.8 | 37.9 | ||||||||||||||||
| 旁遮普语 | 33.0 | 26.9 | 38.3 | 36.5 | 35.1 | 29.9 | 37.4 | 36.4 | ||||||||||||||||
| 罗马尼亚语 | 43.5 | 39.7 | 42.3 | 41.5 | 42.0 | 38.7 | 42.4 | 41.3 | ||||||||||||||||
| 斯洛伐克语 | 39.7 | 38.4 | 39.9 | 38.8 | 39.2 | 38.1 | 40.2 | 39.3 | ||||||||||||||||
| 斯洛文尼亚语 | 31.7 | 26.5 | 34.7 | 33.2 | 34.5 | 30.1 | 38.4 | 37.3 | ||||||||||||||||
| 斯瓦希里语 | 35.0 | 27.4 | 42.6 | 40.5 | 33.9 | 26.8 | 39.2 | 37.9 | ||||||||||||||||
| 塔吉克语 | 33.9 | 29.0 | 34.5 | 33.3 | 36.7 | 32.7 | 38.9 | 38.1 | ||||||||||||||||
| 阿塞拜疆语 | 25.0 | 22.0 | 23.7 | 23.3 | 33.4 | 30.5 | 33.5 | 33.3 | ||||||||||||||||
| 乌克兰语 | 42.0 | 40.1 | 41.7 | 41.4 | 41.7 | 39.7 | 41.6 | 40.7 | ||||||||||||||||
| 平均 | 37.0 | 33.5 | 37.4 | 36.6 | 38.9 | 35.7 | 39.4 | 38.9 | ||||||||||||||||
Abstract
In this work, we present Qwen3.5-Omni, the latest advancement in the Qwen-Omni model family. Representing a significant evolution over its predecessor, Qwen3.5-Omni scales to hundreds of billions of parameters and supports a 256k context length. By leveraging a massive dataset comprising heterogeneous text-vision pairs and over 100 million hours of audio-visual content, the model demonstrates robust omni-modality capabilities. Qwen3.5-Omni-Plus achieves SOTA results across 215 audio and audio-visual understanding, reasoning, and interaction subtasks and benchmarks, surpassing Gemini-3.1 Pro in key audio tasks and matching it in comprehensive audio-visual understanding. Architecturally, Qwen3.5-Omni employs a Hybrid Attention Mixture-of-Experts (MoE) framework for both Thinker and Talker, enabling efficient long-sequence inference. The model facilitates sophisticated interaction, supporting over 10 hours of audio understanding and 400 seconds of 720P video (at 1 FPS). To address the inherent instability and unnaturalness in streaming speech synthesis—often caused by encoding efficiency discrepancies between text and speech tokenizers—we introduce ARIA (Adaptive Rate Interleave Alignment). ARIA dynamically aligns text and speech units, significantly enhancing the stability and prosody of conversational speech with minimal latency impact. Furthermore, Qwen3.5-Omni expands linguistic boundaries, supporting multilingual understanding and speech generation across 10 languages with human-like emotional nuance. Beyond preset voices, the model enables zero-shot voice customization via user-provided samples. Finally, Qwen3.5-Omni exhibits superior audio-visual grounding capabilities, generating script-level structured captions with precise temporal synchronization and automated scene segmentation. Remarkably, we observed the emergence of a new capability in omnimodal models: directly performing coding based on audio-visual instructions, which we call Audio-Visual Vibe Coding. Qwen3.5-Omni is publicly accessible via API111https://www.alibabacloud.com/help/en/model-studio/qwen-omni.
1 Introduction
Human interaction with the world is inherently omnimodal and agentic, involving the integration of visual, auditory, and linguistic information, and the production of responses through text, speech, and goal-directed tool-mediated actions, facilitating information exchange with other organisms and demonstrating intelligence. Building on the rapid advances in the understanding and reasoning capabilities of large models across text (Brown et al., 2020; OpenAI, 2023; Gemini Team, 2024; Anthropic, 2023b; a; 2024; Bai et al., 2023a; Yang et al., 2024a; 2025a; Touvron et al., 2023; Dubey et al., 2024), vision (Li et al., 2023; Liu et al., 2023; Zhu et al., 2023; Bai et al., 2023b; 2025a), and audio (Chu et al., 2023; 2024), natively omnimodal systems that jointly process and generate across all modalities have drawn substantial attention (OpenAI, 2024; Comanici et al., 2025; Xu et al., 2025a; b). However, existing models predominantly operate within passive perception-response paradigms and exhibit limited capacity for scalable agentic behavior, real-time interaction, autonomous tool utilization, and cross-modal reasoning, which are essential prerequisites for practical deployment.
In this report, we present Qwen3.5-Omni, Qwen’s latest generation of fully omnimodal LLM, supporting the understanding of text, images, audio, and audio-visual content. Natively pretrained in an omnimodal manner on massive amounts of text, visual data, and more than 100 million hours of audio-visual data, Qwen3.5-Omni is designed as a native omni agent model: it not only perceives and reasons across all modalities, but also acts, autonomously invoking WebSearch, executing complex FunctionCall, generating speech outputs, and engaging in real-time streaming interaction. The model series includes Plus and Flash variants, all of which are instruct models with 256k-token long-context input.
Qwen3.5-Omni builds on the Thinker–Talker architecture introduced in Qwen2.5-Omni (Xu et al., 2025a) and introduces five key technical upgrades over Qwen3-Omni (Xu et al., 2025b): (1) both the Thinker and Talker adopt Hybrid-Attention Mixture-of-Experts (MoE) designs, enabling highly efficient inference; (2) supporting long-context modeling up to 256k tokens, supporting more than 10 hours of audio and over 400 seconds of 720P audio-visual content at 1 FPS; (3) on the speech generation side, a multi-codebook codec representation enables single-frame, immediate synthesis; (4) the Talker introduces ARIA, a technique that dynamically aligns text and speech units during streaming decoding, significantly improving naturalness and robustness; and (5) multilingual training is substantially expanded, covering 113 languages and dialects for speech recognition and 36 for speech synthesis.
Enabled by these technical advances, Qwen3.5-Omni delivers three major new capabilities over Qwen3-Omni: (1) controllable audio-visual captioning, capable of generating controllable, detailed, and structured captions as well as screenplay-level fine-grained descriptions, including automatic segmentation, timestamp annotation, and detailed descriptions of characters and their relationship to audio; (2) comprehensive real-time interaction, encompassing semantic interruption through native turn-taking intent recognition, end-to-end voice control over volume, speed, and emotion, and voice cloning from user-provided samples; and (3) native omnimodal agentic behavior, including autonomous WebSearch, complex FunctionCall invocation, and Audio-Visual Vibe Coding, an emergent capability wherein the model directly generates executable code from audio-visual instructions, enabling the model to respond to real-time queries without external orchestration.
Critically, Qwen3.5-Omni maintains state-of-the-art performance on text and visual modalities without degradation relative to same-size single-model Qwen counterparts. Across 215 audio and audio-visual understanding, reasoning, and interaction subtasks and benchmarks, covering audio-visual benchmarks, audio benchmarks, ASR benchmarks, language-specific speech-to-text translation tasks, and language-specific ASR tasks, Qwen3.5-Omni-Plus achieves SOTA results, surpassing Gemini-3.1 Pro across general audio understanding, reasoning, recognition, translation, and dialogue, while its overall audio-visual understanding reaches the level of Gemini-3.1 Pro.
2 Architecture
2.1 Overview
As shown in Figure 2, Qwen3.5-Omni continues to adopt the Thinker-Talker architecture (Xu et al., 2025a). Compared with Qwen3-Omni (Xu et al., 2025b), Qwen3.5-Omni introduces several key improvements in scalability, alignment, and real-time interaction:
-
The overall backbone adopts a Hybrid Mixture-of-Experts (MoE) design, improving scalability while better balancing capacity and efficiency across multimodal understanding and generation.
-
The Thinker receives visual and audio signals through the Vision Encoder and AuT, respectively. Audio and video inputs are interleaved for unified multimodal modeling, with explicit timestamps inserted to improve temporal perception, especially for long video or audio-video contexts. This design enables the Thinker to handle extended inputs, supporting up to 256k tokens, 10 hours of audio, or 400 seconds of 720P video at 1 FPS.
-
The Talker is responsible for contextual speech generation by conditioning on multimodal inputs together with the textual outputs from the Thinker. Qwen3.5-Omni adopts the RVQ-based speech representation introduced in Qwen3-Omni (Xu et al., 2025b), which substantially improves inference efficiency.
-
To support real-time interaction, Qwen3.5-Omni adopts both chunk-wise streaming input processing in the Thinker and a streaming Talker design, enabling low-latency end-to-end multimodal conversation.
-
Different from the dual-track Talker input design in Qwen3-Omni (Xu et al., 2025b), the Talker in Qwen3.5-Omni adopts ARIA to dynamically align text and speech units before interleaving them. This design mitigates the instability caused by mismatched tokenization rates between text and speech, thereby reducing issues such as skipped words, incorrect pronunciations, and ambiguous rendering of numbers.
In the following sections, we first introduce with the AuT encoder, including its training methodology. Then, describe how Thinker processes various inputs. We then detail Talker’s multi-codebook streaming speech generation. Finally, we highlight a series of improvements on both the understanding and generation modules aimed at achieving ultra–low-latency, end-to-end streaming audio inference.
2.2 Audio Transformer (AuT)
We use transformer based audio encoder trained from scratch in attention-encoder-decoder model AuT, as is shown in Figure 3. The training of Qwen3.5-Omni encoder consumed 40 million hours of audio-text pair data generated by Qwen3-ASR. The filter bank features of the audio are downsampled 16 times using 4 Conv2D blocks and then fed into self-attention layers to obtain audio tokens in 6.25Hz token rate. Comparing to the training process of Qwen3-Omni encoder, the encoder of Qwen3.5-Omni adapts more multilingual data of more than 20 languages, and the proportion of Chinese, English and multilingual data comes to 3.5 : 3.5 : 3. The dynamic attention window size training mechanism is adopted for guaranting balance performance of inference under real-time prefill caching and for the offline audio understanding tasks.
2.3 Perceivation
Text, Audio, Image and Video (w/o Audio).
The Thinker converts text, audio, image, and silent video inputs into a unified sequence of representations. For text, we use the Qwen3.5 tokenizer (Team, 2026), which adopts byte-level byte-pair encoding with a vocabulary size of 250k (up from 150k), improving encoding and decoding efficiency by 10–60% across most languages. For audio inputs, including audio extracted from video, we resample the waveform to 16 kHz and convert it into a 128-channel mel-spectrogram using a 25 ms window and a 10 ms hop size. We use AuT as the audio encoder, trained from scratch on 40 million hours of audio data, where each output frame corresponds to approximately 160 ms of the original signal. For visual inputs, we adopt the vision encoder from Qwen3.5 (Team, 2026) to process both images and videos. Trained on a mixture of image and video data, this encoder provides strong capabilities in both image understanding and video comprehension. To preserve video information as much as possible while maintaining alignment with the audio stream, we sample video frames at a dynamic frame rate.
Audio-visual Timestamp.
Following Qwen3-Omni (Xu et al., 2025b), we apply TM-RoPE to endow the model with temporal awareness for audio-video synchronization. However, we find that directly encoding absolute time through temporal position IDs can lead to excessively sparse indices for visual patches from long video with audio inputs, which weakens long-range temporal modeling. In addition, such a design often requires large-scale and uniformly distributed training samples across different frame rates, increasing data construction cost. To address these issues, we prepend each video or audio-video temporal patch with an explicit timestamp represented as a formatted text string in seconds, allowing the model to learn timecode representations more naturally. For audio sequences, we further insert timestamps at random intervals to improve temporal alignment across modalities. Although this strategy slightly increases the context length, it enables more precise and robust temporal perception, especially when extrapolating long-context multimodal inputs.
In the context of multimodal audio-visual streams, the audio component is encoded with a temporal ID for every 160 ms. The video is treated as a sequence of frames with monotonically increasing temporal IDs that are dynamically adjusted based on their actual timestamps to ensure a consistent temporal resolution of 160 ms per ID. The height and width IDs for video frames are assigned in the same manner as for still images. To prevent positional conflicts when processing multiple modalities, the position numbering is made contiguous, with each subsequent modality commencing from one plus the maximum position ID of the preceding modality. This refined approach to positional encoding enables the model to effectively integrate and jointly model information from diverse modalities. Qwen3.5-Omni aligns these representations using their temporal IDs, which are explicitly anchored to absolute time. This design choice affords the model the flexibility to support streaming inputs of arbitrary duration.
2.4 Speech Generation
Talker operates directly on the RVQ tokens produced by Qwen3.5-Omni-Audio-Tokenizer. To model the residual codebooks, it employs a multi-token prediction (MTP) module, which enables fine-grained modeling and control of acoustic details. Coupled with a causal ConvNet for waveform reconstruction, Talker delivers high-fidelity speech synthesis with low inference latency and modest computational overhead.
In multi-turn spoken dialogue, Talker is conditioned on the rich contextual information provided by the Thinker component, including historical text tokens, multimodal representations, and the streamed text of the current turn. Such conditioning allows Talker to dynamically modulate acoustic attributes—such as prosody, loudness, and emotion—in accordance with the evolving conversational context.
Architecturally, our approach differs from Qwen3-Omni (Xu et al., 2025b) in two key respects. First, we introduce a dedicated system prompt for Talker that specifies target voice characteristics, thereby enabling both zero-shot voice cloning and controllable speech generation. Compared with conventional speaker embeddings, this prompt can encode richer multimodal cues, including textual descriptions and codec sequences, providing substantially finer-grained control over acoustic realization. Second, we propose ARIA (Adaptive Rate Interleave Alignment), which unifies the conventional dual-channel generation paradigm into a single-channel formulation. Rather than relying on MFA-derived alignments or fixed interleaving rates, ARIA enforces an adaptive rate constraint: for any prefix of the generated sequence, the cumulative speech-to-text token ratio must not exceed the corresponding item-level global ratio. Despite its simplicity, this design affords flexible text-speech alignment across languages, including those with relatively low encoding efficiency, and naturally supports arbitrary text-token prefixes followed by coherent speech-token continuation.
2.5 Designs for Streaming and Concurrency
In streaming audio-visual interaction scenarios, the first-packet latency is a critical factor affecting user experience, and the model’s concurrency capability is key to reducing service costs and improving response speed. This section discusses how Qwen3.5-Omni enhances concurrency and reduces first-packet latency through algorithmic and architectural optimizations. Table 1 provides an overview of the relevant architecture of the Qwen3.5-Omni and its associated latency.
| Module | Architecture | Streaming |
| Audio Encoder | AuT | |
| Vision Encoder | SigLIP2 | – |
| Thinker | Hybrid MoE Transformer | |
| Talker | Hybrid MoE Transformer | |
| MTP | Dense Transformer | |
| Code2wav | ConvNet | |
| First-Packet Latency (Audio Input) | Plus: 435ms Flash: 235ms | |
| First-Packet Latency (Video Input) | Plus: 651ms Flash: 426ms | |
Chunked Prefilling and Hybrid MoE Architecture.
In Qwen3.5-Omni, we retain the chunked-prefilling mechanism as implemented in Qwen3-Omni and Qwen2.5-Omni, whose audio and vision encoders are capable of outputting chunks along the temporal dimension. This approach significantly reduces the Time-To-First-Token (TTFT) for both the Thinker and the Talker. Architecturally, both the Thinker and the Talker in Qwen3.5-Omni are built upon the Hybrid MoE architecture introduced in Qwen3.5. Beyond the general efficiency advantage of Hybrid MoE, this architecture includes the Gated Delta Net (GDN) module, which is particularly effective for accelerating the modeling of long audio-video sequences. As a result, it significantly reduces KV-cache I/O overhead in long-context inference, improving generation throughput and enabling higher serving concurrency.
Streaming Generation with ARIA.
For streaming speech generation and high-concurrency serving, Qwen3.5-Omni largely inherits the efficient design of Qwen3-Omni: Talker predicts RVQ codec tokens with a lightweight MTP module, and the generated multi-codebook tokens are converted to waveform by a causal and streaming ConvNet codec decoder. These components remain computationally lightweight, batch-friendly, and well-suited for low-latency deployment. Built on this shared foundation, the previously introduced ARIA further reformulates the dual-channel generation pattern in Qwen3-Omni into a unified interleaved single-stream formulation over text and speech tokens. By organizing text and speech generation under a monotonic interleaving constraint, ARIA reduces the synchronization overhead between separate generation tracks, enables more efficient token scheduling during decoding, and better matches the naturally incremental regime of streaming interaction.
In Table 2, we report the theoretical first-packet latency of Qwen3.5-Omni under different concurrency levels for audio and video input, evaluated on internal vLLM with torch.compile and CUDA Graph acceleration enabled for the MTP module and codec decoder. Here, Thinker TTFT (Time-To-First-Token) denotes the time from receiving the input stream to the first text token generated by Thinker, while Talker TTFC (Time-To-First-Chunk) measures the time until Talker produces the first audio chunk. TPOP (Time-Per-Output-Token) represents the per-output-token latency during steady-state decoding, where Talker TPOP includes the combined latency of the Talker backbone and the MTP module. TPS (Tokens Per Second) denotes generation throughput. Since ARIA organizes text and speech generation in a unified interleaved stream, Overall Latency cannot be obtained by simply summing several row values, but instead reflects the end-to-end critical path to the first playable audio packet. We also note that, due to the substantial scale difference between Qwen3.5-Omni-Flash and Qwen3.5-Omni-Plus, the two variants adopt different deployment-time resource allocation and parallelization strategies; therefore, their latency and throughput numbers are not intended for strict horizontal comparison. As shown in the table, Qwen3.5-Omni maintains stable latency and decoding efficiency as concurrency increases, while the low Generation RTF provides sufficient margin for smooth streaming audio generation.
| Qwen3.5-Omni-Flash | Qwen3.5-Omni-Plus | |||||
| 1 Conc. | 4 Conc. | 8 Conc. | 1 Conc. | 4 Conc. | 8 Conc. | |
| Thinker TTFT | 80/255ms | 86/446ms | 103/765ms | 162/377ms | 183/907ms | 260/1243ms |
| Talker TTFC | 56/61ms | 68/108ms | 81/116ms | 54/56ms | 72/88ms | 95/116ms |
| Thinker TPOP | 5.6/5.9ms | 8.2/9.2ms | 9.6/15.8ms | 17.4/18.5ms | 25.6/26.9ms | 33.3/40.2ms |
| Talker TPOP | 14.2/14.2ms | 16.9/17.0ms | 20.5/20.6ms | 14.9/14.9ms | 21.0/21.3ms | 25.8/27.1ms |
| Codec Decode | 3~5ms | |||||
| Overall Latency | 235/426ms | 298/891ms | 352/1625ms | 435/651ms | 619/1515ms | 955/1980ms |
| Thinker TPS | 177/171 | 556/457 | 942/598 | 57/54 | 156/149 | 266/240 |
| Talker TPS | 70/70 | 237/235 | 389/388 | 67/67 | 191/189 | 320/296 |
| Generation RTF | 0.178 | 0.211 | 0.257 | 0.187 | 0.267 | 0.334 |
3 Pretraining
| Modality | # Varieties | Supported languages and dialects |
| Text | 201 | See Qwen3.5 for the complete list of supported languages. |
| Speech Input | 113 | 74 languages: Afrikaans, Arabic, Asturian, Azerbaijani, Basque, Belarusian, Bengali, Bosnian, Bulgarian, Cantonese, Catalan, Cebuano, Chinese, Croatian, Czech, Danish, Dutch, English, Esperanto, Estonian, Filipino, Finnish, French, Galician, Georgian, German, Greek, Hebrew, Hindi, Hungarian, Icelandic, Indonesian, Interlingua, Italian, Japanese, Javanese, Kannada, Kazakh, Korean, Kyrgyz, Lingala, Latvian, Lithuanian, Macedonian, Malay, Malayalam, Maltese, Maori, Marathi, Mongolian, Norwegian Bokmål, Norwegian Nynorsk, Oriya, Persian, Polish, Portuguese, Punjabi, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swahili, Swedish, Tajiki, Tamil, Telugu, Thai, Turkish, Ukrainian, Urdu, Uyghur, and Vietnamese. 39 Chinese dialects: Northeastern Mandarin, Guizhou dialect, Guangdong Cantonese, Henan dialect, Hong Kong Cantonese, Shanghainese, Shaanxi dialect, Tianjin dialect, Taiwanese Mandarin, Yunnan dialect, Anhui dialect, Fujian dialect, Gansu dialect, Guangdong Mandarin, Hubei dialect, Hunan dialect, Jiangxi dialect, Shandong dialect, Shanxi dialect, Sichuanese, Guangxi dialect, Hainan dialect, Chongqing dialect, Changsha dialect, Hangzhou dialect, Hefei dialect, Yinchuan dialect, Zhengzhou dialect, Shenyang dialect, Wenzhou dialect, Wuhan dialect, Kunming dialect, Taiyuan dialect, Nanchang dialect, Jinan dialect, Lanzhou dialect, Nanjing dialect, Hakka, and Southern Min. |
| Speech Output | 36 | 29 languages: Chinese, English, German, Italian, Portuguese, Spanish, Japanese, Korean, French, Russian, Thai, Indonesian, Arabic, Vietnamese, Turkish, Finnish, Polish, Hindi, Dutch, Czech, Urdu, Tagalog, Swedish, Danish, Hebrew, Icelandic, Malay, Norwegian, and Persian. 7 Chinese dialects: Sichuanese, Beijing dialect, Tianjin dialect, Nanjing dialect, Shaanxi dialect, Cantonese, and Southern Min. |
Qwen3.5-Omni is pre-trained on a diverse dataset that encompasses multiple languages and dialects as shown in Table 3 and modalities, including image-text, video-text, audio-text, video-audio, video-audio-text, and pure text corpora. Following Qwen3-Omni (Xu et al., 2025b), we employ a wider range of natural language prompts to enhance both the generalization ability and instruction-following capabilities. To achieve robust performance across all modalities, our training strategy incorporates both unimodal and cross-modal data from the early pretraining stage.
In Qwen3-Omni (Xu et al., 2025b), we employ TMRoPE to endow the model with temporal awareness. However, we identify two key limitations of this approach: (1) By directly tying temporal position IDs to absolute time, it produces excessively large and sparse temporal position IDs for long audio-video or video inputs, which undermines the model’s ability to capture long-range temporal contexts. (2) Effective learning under this scheme typically requires large-scale and uniformly distributed sampling across different frame rates (fps), significantly increasing the cost of training data construction. To address these issues, we prepend each video or audio-video temporal patch with a timestamp represented as a formatted text string in seconds, enabling the model to better learn and interpret timecode representations. In addition, for audio sequences, we insert timestamps at random intervals to better align training across different modalities. Although this approach introduces a modest increase in context length, it allows the model to perceive temporal information more effectively and precisely.
The pre-training of Qwen3.5-Omni is structured into three distinct stages. In the first stage, we lock the LLM parameters and focus on training the vision and audio encoders, utilizing a vast corpus of audio-text and image-text pairs to enhance semantic understanding within the LLM. In the second stage, we unfreeze all parameters and train with a wider range of multimodal data for more comprehensive learning with a sequence length of 32,768. In the final stage, we use data with a sequence length of 262,144 to enhance the model’s ability to understand complex long-sequence data:
- (1)
Encoder Alignment Stage (S1): During the initial pretraining phase, the LLM component of Qwen3.5-Omni is initialized with parameters from Qwen3.5, while the vision encoder is adopted from Qwen3.5, and the audio encoder is initialized with AuT. The two encoders are trained separately on the fixed LLM, with both initially focusing on training their respective adapters before training the encoders.
- (2)
General Stage (S2): The second phase of pretraining utilizes a large-scale dataset containing approximately 4 trillion tokens, with the following distribution across modalities: text (0.92 trillion), audio (1.99 trillion), image (0.95 trillion), video (0.14 trillion), and video-audio 0.29 trillion). During this stage, the introduction of more diverse multimodal data and tasks enhances the model’s understanding and interaction capabilities in auditory, visual, textual, and audio-visual information.
- (3)
Long Context Stage (S3): In the final pre-training phase, we increased the maximum token length from 32,768 to 262,144 and also raised the proportion of long audio and long video in the training data. Experimental results indicate that these adjustments lead to significant improvements in the model’s ability to understand long sequence data.
4 Post-training
4.1 Thinker
The post-training phase employs a three-stage strategy for the Thinker, designed to preserve the model’s capabilities across all modalities without degradation, ensure high response quality under audio queries, and optimize the overall interaction experience. The training corpus, structured in the ChatML (OpenAI, 2022) format, encompasses pure text, visual, audio, and mixed-modality conversational data. Specifically, the process consists of the following stages:
-
Stage 1: Specialist Distillation To establish a strong foundation for omnimodal capabilities, we first train a suite of domain-specialized teacher models via independent Supervised Fine-Tuning (SFT) and reinforcement learning (RL). All teacher models are fine-tuned from the pre-trained Qwen-3.5 base checkpoint. Beyond text-related tasks, including agentic, coding, and foundational reasoning tasks, we also train specialized teacher models for vision and audio. These teacher models are used to generate domain-specific data, enabling the specialized capabilities learned in each domain to be distilled into a single unified model.
-
Stage 2: On-Policy Distillation Through the specialist distillation described above, the model already achieves strong performance in domains such as multimodal understanding and reasoning, as well as text-based dialogue, reasoning, coding, and agentic tasks. Nevertheless, a substantial gap remains between the quality of responses conditioned on audio queries and that of responses conditioned on text queries, particularly in speech dialogue. To reduce this gap, we introduce a second-stage training procedure based on on-policy distillation (OPD), with the goal of distilling the model’s stronger response capabilities under text inputs into the audio-input setting. Concretely, for each audio-text paired query, we first obtain a response generated under the text condition, which typically exhibits higher quality in terms of fluency, reasoning, and task completion. We then use this response as the distillation target for the corresponding audio-conditioned query. By training on such on-policy targets, the model gradually aligns its audio-conditioned outputs with its text-conditioned behavior, thereby improving response quality under audio inputs and promoting modality-consistent generation.
-
Stage 3: Interaction-Aligned Reinforcement Learning Although the previous two stages substantially improve the model’s domain capabilities and cross-modal response quality, they are not sufficient to fully optimize the model for real-world interactive use. In multi-turn conversations, we observe several interaction-specific issues, including unintended language code-switching, persona inconsistency, and degraded instruction-following over extended contexts. To mitigate these issues, we introduce Interaction-Aligned RL, a third-stage reinforcement learning procedure aimed at optimizing the model for interaction quality. We construct multi-turn interaction trajectories and design reward signals around these user experience objectives, enabling the model to learn behaviors that are more stable, consistent, and aligned in prolonged interactions. By explicitly optimizing for interaction quality, this stage improves the model’s overall usability in practical conversational scenarios.
4.2 Talker
We employ a four-stage training pipeline for Talker, enabling Qwen3.5-Omni to generate natural and contextually appropriate spoken responses jointly with text. All training data is organized in the ChatML format to maintain consistency with Thinker and to facilitate voice cloning.
- (1)
General Stage: In the initial pre-training stage, we train Qwen3.5-Omni on more than 20 million hours of multilingual speech data paired with multimodal context. In particular, the introduction of more diverse tasks, such as instruction-following speech generation, substantially enhance contextual reasoning and paralinguistic alignment, going beyond a simple monotonic mapping from multimodal representations to speech.
- (2)
Long-Context Stage: We perform data quality stratification through a dedicated curation pipeline and conduct continual pre-training (CPT) on high-quality subsets. Augmented by Qwen3-Omni-Captioner, this stage mitigates hallucinations introduced by noisy data in the initial pre-training phase and substantially improves the naturalness and quality of generated speech. Meanwhile, we extend the maximum context length to 64k tokens, allowing the model to better handle long and complex user inputs and to produce more contextually grounded speech responses.
- (3)
Reinforcement Learning Stage: We further align model behavior with human preferences through Direct Preference Optimization (DPO) (Rafailov et al., 2023). Concretely, we construct multilingual preference pairs based on human annotations and optimize the model with DPO. In addition, we incorporate rule-based rewards and adopt GSPO (Zheng et al., 2025) to further improve overall capability and training stability across diverse tasks.
- (4)
Speaker Fine-tuning Stage: Finally, we perform lightweight speaker fine-tuning on top of the base model, enabling Qwen3.5-Omni to faithfully capture target speaker characteristics while further improving the naturalness, expressiveness, and controllability of its speech outputs.
5 Evaluation
A comprehensive evaluation was performed on two variants of models, including Qwen3.5-Omni-Flash and Qwen3.5-Omni-Plus. The evaluation results are divided into two main categories: understanding (XText) and speech generation (XSpeech).
5.1 Evaluation of XText
In this section, we evaluate Qwen3.5-Omni’s ability to comprehend various multimodal inputs (text, audio, vision, and audio-visual video) and generate textual responses.
TextText
Our evaluation of Qwen3.5-Omni on text text primarily focuses on general knowledge tasks, instruction following, long context tasks, STEM tasks, reasoning tasks and general agent ability. Specifically, we utilize MMLU-Pro (Wang et al., 2024d), MMLU-Redux (Gema et al., 2024), SuperGPQA (Team et al., 2025) and C-Eval (Huang et al., 2023) for general knowledge tasks, IFEval (Zhou et al., 2023) and IFBench (Pyatkin et al., 2025) for instruction following, AA-LCR (Team, 2025a) and LongBench v2 (Bai et al., 2025b) for long context tasks, GPQA (Rein et al., 2023) for STEM tasks, LiveCodeBench v6 (Jain et al., 2024), HMMT Nov 25 (Balunović et al., 2025) and IMOAnswerBench (Luong et al., 2025) for reasoning tasks, BFCL-V4 (Yan et al., 2024) and TAU2Bench (Barres et al., 2025) for general agent ability.
AudioText
To evaluate audio-to-text capabilities, we employ benchmarks across four domains: audio understanding, end-to-end speech dialogue, speech-to-text translation (S2TT), and automatic speech recognition (ASR). For audio understanding, we utilize MMAU (Sakshi et al., 2024), MMAR (Ma et al., 2025a), MMSU (Wang et al., 2025a), RUL-MuchoMusic (Zang et al., 2025), and SongFormBench (Hao et al., 2025) to assess comprehension of sound effects, speech, and music. Dialogue performance is evaluated via VoiceBench (Chen et al., 2024b), URO-Bench-pro (Yan et al., 2025), SpeechRole (Jiang et al., 2025), and WildSpeech-Bench (Zhang et al., 2025b). For S2TT, we focus on the translation of the top 59 languages in Fleurs (Conneau et al., 2022) into English and Chinese. Finally, ASR performance is measured using Fleurs (Conneau et al., 2022), Common Voice (Ardila et al., 2020), LibriSpeech (Panayotov et al., 2015), WenetSpeech (Zhang et al., 2022), KeSpeech (Tang et al., 2021), Opencpop-test (Wang et al., 2022), and MIR-1K (vocal) (Hsu and Jang, 2010), covering multilingual speech, Chinese dialects, and singing voice transcription.
VisionText
The evaluation of the model’s vision-to-text capabilities encompasses a suite of benchmarks targeting diverse and challenging tasks. To assess performance in specialized domain of mathematical and STEM reasoning, we utilize MMMU (Yue et al., 2023), MMMU-Pro (Yue et al., 2024), MathVista (Lu et al., 2024), MathVision (Wang et al., 2024a), DynaMath (Zou et al., 2025), ZEROBench (Roberts et al., 2025). For the general visual question answering, the model is evaluated on RealWorldQA (Zhang et al., 2024), MMStar (Chen et al., 2024a), HallusionBench (Guan et al., 2024), and SimpleVQA (Cheng et al., 2025). The model’s proficiency in document understanding is measured using the CharXiv (Wang et al., 2024e), CC-OCR (Yang et al., 2024b), AI2D (Kembhavi et al., 2016), MMLongBench-Doc (Ma et al., 2024), and OCRBench (Liu et al., 2024). Furthermore, the model’s spatial intelligence is specifically tested on ERQA (Team, 2025b), CountBench (Paiss et al., 2023), RefCOCO (Kazemzadeh et al., 2014), ODInW13 (Li et al., 2022), and EmbSpatialBench (Du et al., 2024a). To evaluate performance on dynamic visual data, we report results on six video understanding benchmarks: Video-MME (Fu et al., 2024), MLVU (Zhou et al., 2025a), MVBench (Li et al., 2024), LVBench (Wang et al., 2024b), MMVU (Zhao et al., 2025) and MME-VideoOCR (Shi et al., 2025). Specifically, we evaluate the model’s performance on medical VQA across three established benchmarks: SLAKE (Liu et al., 2021), PMC-VQA (Zhang et al., 2023), and MedXpertQA-MM (Zuo et al., 2025). This assessment is designed to demonstrate the model’s comprehensive clinical reasoning capabilities and its potential utility as a reliable healthcare AI assistant.
Audio-Visual VideoText
We evaluate our model’s audio-visual understanding capabilities from multiple perspectives. For text-query evaluation, we use DailyOmni (Zhou et al., 2025b), WorldSense (Hong et al., 2025), AVUT (Yang et al., 2025b), AV-SpeakerBench (Nguyen et al., 2025), and VideoMME (Fu et al., 2025). To assess the model’s ability in real-world audio-visual interactive scenarios, we use Qualcomm IVD (Pourreza et al., 2025) as the benchmark for audio-query-based evaluation. Beyond understanding, we also evaluate the model’s captioning capability on OmniCloze (Ma et al., 2025b) and its tool-use ability on OmniGAIA (Li et al., 2026).
5.1.1 Performance of TextText
We compare Qwen3.5-Omni-Plus and Qwen3.5-Omni-Flash with Qwen3.5-Plus-Instruct. As shown in Table 4, Qwen3.5-Omni-Plus demonstrates text capabilities that are on par with its text-only counterpart across multiple dimensions, including knowledge, instruction following, long-context understanding, STEM, reasoning, and general agent tasks, highlighting its strong language ability. In particular, Qwen3.5-Omni ’s instruction-following performance is slightly better than the baseline. We believe that OPD and interaction-aligned RL have a positive effect on improving the instruction-following capabilities of an omni-model LLM.
| Datasets | Qwen3.5-Plus-Instruct | Qwen3.5-Omni-Flash | Qwen3.5-Omni-Plus |
| Knowledge | |||
| MMLU-Pro | 86.8 | 79.9 | 85.9 |
| MMLU-Redux | 94.3 | 90.0 | 94.2 |
| SuperGPQA | 67.4 | 54.9 | 66.4 |
| C-Eval | 92.3 | 86.0 | 92.0 |
| Instruction Following | |||
| IFEval | 89.7 | 85.2 | 89.7 |
| IFBench | 51.1 | 38.4 | 52.6 |
| Long Context | |||
| AA-LCR | 62.0 | 46.0 | 57.0 |
| LongBench v2 | 60.2 | 46.4 | 59.6 |
| STEM | |||
| GPQA | 85.9 | 76.4 | 83.9 |
| Reasoning | |||
| LiveCodeBench v6 | 67.1 | 56.6 | 65.6 |
| HMMT Nov 25 | 86.2 | 59.0 | 84.4 |
| IMOAnswerBench | 68.3 | 51.5 | 65.5 |
| General Agent | |||
| BFCL-V4 | 66.1 | 55.3 | 63.3 |
| TAU2Bench | 82.7 | 78.0 | 81.0 |
5.1.2 Performance of AudioText
| Datasets | Gemini-3.1 Pro | Qwen3.5-Omni-Flash | Qwen3.5-Omni-Plus |
| Audio Understanding () | |||
| MMAU | 81.1 | 80.4 | 82.2 |
| MMAR | 83.7 | 74.0 | 80.0 |
| MMSU | 81.3 | 72.2 | 82.8 |
| RUL-MuchoMusic | 59.6 | 60.5 | 72.4 |
| SongFormBench-HarmonixSeta | 75.6 — 46.8 — 77.9 | 80.6 — 67.8 — 83.4 | 81.1 — 72.9 — 85.3 |
| SongFormBench-CNa | 78.1 — 43.2 — 71.9 | 86.7 — 66.4 — 84.6 | 87.1 — 65.7 — 84.2 |
| Dialogue () | |||
| VoiceBench | 88.9 | 87.8 | 93.1 |
| URO-Bench-prob | 69.1 — 84.0 — 99.2 | 64.1 — 83.8 — 98.7 | 66.3 — 86.3 — 99.8 |
| SpeechRole | 124.2 | 119.8 | 123.5 |
| WildSpeech-Bench | 76.3 | 72.2 | 75.4 |
| S2TT () | |||
| Fleursc | 29.5 | 26.9 | 30.2 |
| Fleursc | 34.6 | 32.0 | 35.4 |
| Fleursc | 32.1 | 29.4 | 32.8 |
| ASR (WER) | |||
| Fleurs | 7.32 | 10.75 | 6.55 |
| CV15 | 8.59 — 13.40 — 6.78 | 4.25 — 3.45 — 2.68 | 3.46 — 1.95 — 2.27 |
| CV15 | 8.73 | 5.90 | 4.83 |
| Librispeech | 3.36 — 4.41 | 1.30 — 2.43 | 1.11 — 2.23 |
| Weneetspeech | 11.53 — 14.21 | 4.41 — 5.51 | 4.30 — 5.84 |
| Kespeech | 23.67 | 4.47 | 3.46 |
| MIR-1Kd | 8.76 | 4.94 | 4.56 |
| Opencpop | 6.83 | 1.11 | 1.49 |
- a
SongFormBench: We use a unified prompt defining an SRT-like output timestamp format and a closed vocabulary for evaluation. The vocabulary follows the SongForm-HX-8Class specified in the official codebase.
- b
URO-Bench-Pro: We use the pro track of URO-Bench and denote the three evaluation dimensions as follows: U for Understanding, R for Reasoning, and O for Oral Conversation. We use GenStyle-en, GenStyle-zh, Multilingual tasks for oral dimension.
- c
Fleurs: The top59 languages are English, Chinese, Cantonese, Korean, Japanese, Vietnamese, Thai, Malay, German, Russian, Italian, French, Spanish, Portuguese, Dutch, Indonesian, Turkish, Arabic, Polish, Hindi, Urdu, Filipino, Persian, Czech, Greek, Swedish, Hebrew, Danish, Finnish, Norwegian, Icelandic, Bengali, Punjabi, Javanese, Marathi, Swahili, Ukrainian, Gujarati, Kannada, Azerbaijani, Malayalam, Cebuano, Romanian, Hungarian, Bulgarian, Belarusian, Catalan, Tamil, Croatian, Bosnian, Slovak, Galician, Kyrgyz, Macedonian, Slovenian, Latvian, Estonian, and Asturian; compared with the top60 list, Afrikaans is excluded because the Fleurs S2TT test set does not cover this language.
- d
MIR-1K: Transcription is converted into Simplified Chinese.
In Table 5, we compare Qwen3.5-Omni with Gemini-3.1 Pro in terms of audio-to-text performance. Compared to Gemini-3.1 Pro, Qwen3.5-Omni exhibits superior performance on MMAU, MMSU, RUL-MuchoMusic, and SongFormBench, while achieving comparable results on MMAR, demonstrating its strong comprehension capabilities across multiple audio domains. Regarding end-to-end speech dialogue, Qwen3.5-Omni significantly outperforms Gemini-3.1 Pro on VoiceBench and matches its performance on other benchmarks, further validating Qwen3.5-Omni ’s robust capabilities in end-to-end voice interaction. For S2TT and ASR, Qwen3.5-Omni consistently outperforms Gemini-3.1 Pro, underscoring its superior translation and speech recognition performance across diverse languages, dialects, and domains.
5.1.3 Performance of Vision Text
To comprehensively evaluate vision-to-text capabilities, we compare Qwen3.5-Omni-Flash and Qwen3.5-Omni-Plus with Qwen3.5-Plus-Instruct. As shown in Table 6, Qwen3.5-Omni-Plus achieves performance comparable to that of Qwen3.5-Plus-Instruct, while demonstrating stronger results on video understanding tasks involving both short and long videos. These findings highlight the strong dynamic visual perception ability of our model in real-world scenarios and suggest the effectiveness of joint video-audio training paradigms. Furthermore, we posit that audio-visual streams constitute the most naturalistic representation of real-world phenomena, wherein visual and auditory modalities are intrinsically coupled rather than independently processed.
| Datasets |
|
|
| |||
| STEM and Puzzle | ||||||
| MMMU | 81.0 | 76.9 | 80.1 | |||
| MMMU-Pro | 73.8 | 68.2 | 73.9 | |||
| MathVision | 73.6 | 65.4 | 73.0 | |||
| Mathvista (mini) | 86.9 | 82.9 | 86.1 | |||
| DynaMath | 84.2 | 79.3 | 83.8 | |||
| ZEROBench | 6 | 1 | 5 | |||
| ZEROBench_sub | 31.1 | 26.0 | 34.4 | |||
| General VQA | ||||||
| RealWorldQA | 79.1 | 77.5 | 84.1 | |||
| MMStar | 80.3 | 75.7 | 79.4 | |||
| MMBenchEN-DEV-v1.1 | 93.8 | 88.8 | 92.8 | |||
| SimpleVQA | 66.1 | 54.4 | 65.3 | |||
| Text Recognition and Document Understanding | ||||||
| CharXiv (RQ) | 74.2 | 64.4 | 72.5 | |||
| CC-OCR | 83.0 | 80.8 | 83.4 | |||
| AI2D_TEST | 92.1 | 89.0 | 91.2 | |||
| MMLongBench-Doc | 59.7 | 53.6 | 57.5 | |||
| OCRBench | 91.4 | 89.1 | 91.3 | |||
| Spatial Intelligence | ||||||
| ERQA | 53.8 | 50.0 | 54.8 | |||
| CountBench | 95.1 | 88.2 | 95.1 | |||
| RefCOCO(avg) | 95.2 | 92.6 | 95.0 | |||
| ODInW13 | 50.3 | 46.8 | 49.5 | |||
| EmbSpatialBench | 83.4 | 82.7 | 85.4 | |||
| Video Understanding | ||||||
| VideoMME(w/o sub.) | 81.0 | 77.0 | 81.9 | |||
| MLVU(M-Avg) | 85.1 | 81.9 | 86.8 | |||
| MVBench | 76.7 | 70.8 | 79.0 | |||
| LVBench | 68.6 | 65.7 | 71.2 | |||
| MMVU | 67.1 | 62.7 | 67.5 | |||
| MME-VideoOCR | 74.2 | 70.5 | 77.0 | |||
| Medical VQA | ||||||
| SLAKE | 82.8 | 73.1 | 84.7 | |||
| PMC-VQA | 62.4 | 58.7 | 62.7 | |||
| MedXpertQA-MM | 55.3 | 44.8 | 54.7 | |||
5.1.4 Performance of Audio-Visual VideoText
We compare Qwen3.5-Omni and Gemini-3.1 Pro across a diverse range of audio-visual tasks, as shown in Table 7. For general understanding, Qwen3.5-Omni achieves state-of-the-art performance on DailyOmni and obtains comparable results on AVUT. Our model also surpasses Gemini-3.1-Pro by a substantial margin on Qualcomm IVD, demonstrating its effectiveness in real-world audio-visual interactive scenarios. Moreover, our model shows strong performance on captioning tasks. It can provide detailed audio, visual, and audio-visual captions. In this version, we also enhance the model’s tool-use capability, achieving 57.2% on OmniGAIA.
Datasets Gemini-3.1 Pro Qwen3.5-Omni-Flash Qwen3.5-Omni-Plus Text Query QA DailyOmni 82.7 81.8 84.6 WorldSense 65.5 57.9 62.8 AVUT 85.6 81.4 85.0 AV-SpeakerBench 75.1 65.2 71.3 VideoMMEw/ audioa 89.0 79.3 83.7 Audio Query QA Qualcomm IVD 66.2 66.3 68.5 Omni-Cloze 57.2 63.0 64.8 Agent (Tool Use) OmniGAIAb 68.9 33.9 57.2 a VideoMME is evaluated with use_audio_in_video=True. b OmniGAIA is evaluated without a thinking prompt and without <answer> formatting. All results are evaluated using DeepSeek-V3.2-Thinking as the judge.
5.2 Evaluation of XSpeech
In this section, we evaluate the speech generation capability of Qwen3.5-Omni. Our evaluation mainly focuses on speech generation conditioned on text and prompt speech, following a zero-shot text-to-speech (TTS) setting. We study the model from four perspectives:
-
Zero-Shot Speech Generation: We evaluate content consistency, measured by WER, on SEED (Anastassiou et al., 2024).
-
Cross-Lingual Speech Generation: We evaluate content consistency in zero-shot cross-lingual speech generation on CV3-Eval (Du et al., 2025).
-
Custom-Voice Speech Generation: We evaluate the stability of our speaker fine-tuned model on the TTS multilingual test set (Zhang et al., 2025a) and our internal multilingual test set.
5.2.1 Evaluation of Zero-Shot Speech Generation
We compare Qwen3.5-Omni with state-of-the-art zero-shot TTS systems. As shown in Table 8, Qwen3.5-Omni achieves highly competitive performance on the SEED-TTS benchmark, demonstrating strong content fidelity in zero-shot speech generation. These results reflect the effectiveness of our pretraining and continual pretraining pipeline in building robust speech generation and context modeling capabilities. Moreover, after RLHF optimization, Qwen3.5-Omni further improves generation stability and naturalness, achieving the best performance on the test-en split with a WER of 1.26.
| Datasets | Model | Performance |
| Content Consistency | ||
| SEED test-zh — test-en | Seed-TTSICL (Anastassiou et al., 2024) | 1.11 — 2.24 |
| Seed-TTSRL (Anastassiou et al., 2024) | 1.00 — 1.94 | |
| MaskGCT (Wang et al., 2024c) | 2.27 — 2.62 | |
| E2 TTS (Eskimez et al., 2024) | 1.97 — 2.19 | |
| F5-TTS (Chen et al., 2024c) | 1.56 — 1.83 | |
| Spark TTS (Wang et al., 2025b) | 1.20 — 1.98 | |
| CosyVoice 2 (Du et al., 2024b) | 1.45 — 2.57 | |
| CosyVoice 3 (Du et al., 2025) | 0.71 — 1.45 | |
| MiniMax-Speech (Zhang et al., 2025a) | 0.83 — 1.65 | |
| MiMo-Audio-7B-Instruct (Zhang et al., 2025c) | 1.96 — 5.37 | |
| Qwen2.5-Omni-7B (Xu et al., 2025a) | 1.42 — 2.33 | |
| Qwen3-Omni-30B-A3B (Xu et al., 2025b) | 1.07 — 1.39 | |
| Qwen3.5-Omni-Plus | 0.99 — 1.26 | |
5.2.2 Evaluation of Multilingual Speech Generation
Qwen3.5-Omni supports speech generation in 29 languages. We compare its multilingual speech generation performance with two strong commercial systems, MiniMax-Speech and ElevenLabs. For the internal multilingual test set, we use GPT-4o-transcribe-2025-03-20 for automatic speech recognition.
As shown in Table 9 and Table 10, Qwen3.5-Omni achieves the lowest WER in 22 out of 29 evaluated languages on the multilingual test sets, outperforming the comparison systems by a clear margin in most cases. On the remaining languages, Qwen3.5-Omni remains competitive with state-of-the-art systems. In addition to content consistency, Qwen3.5-Omni also shows strong voice cloning fidelity. It obtains the highest speaker similarity scores in the majority of evaluated languages and consistently outperforms both MiniMax-Speech and ElevenLabs overall. These results suggest that Qwen3.5-Omni effectively preserves speaker characteristics, such as timbre and prosodic style, while maintaining robust multilingual speech generation quality.
Furthermore, in Table 10, we report results on our internal multilingual test set, covering an additional 9 languages. Qwen3.5-Omni continues to achieve strong performance across all evaluated languages, indicating that its multilingual speech generation ability generalizes well beyond the public benchmark languages.
| Language | Content Consistency | Speaker Similarity | |||||
| MiniMax | ElevenLabs |
| MiniMax | ElevenLabs | ||
| Chinese | 0.695 | 2.252 | 16.026 | 0.800 | 0.780 | 0.677 | |
| English | 0.631 | 2.164 | 0.756 | 0.833 | 0.756 | 0.613 | |
| German | 0.447 | 1.906 | 0.572 | 0.757 | 0.733 | 0.614 | |
| Italian | 0.503 | 1.543 | 1.743 | 0.785 | 0.699 | 0.679 | |
| Portuguese | 1.221 | 1.877 | 1.331 | 0.792 | 0.805 | 0.711 | |
| Spanish | 0.862 | 1.029 | 1.084 | 0.797 | 0.762 | 0.615 | |
| Japanese | 3.479 | 3.519 | 10.046 | 0.788 | 0.776 | 0.738 | |
| Korean | 1.458 | 1.747 | 1.865 | 0.747 | 0.776 | 0.700 | |
| French | 2.430 | 4.099 | 5.216 | 0.730 | 0.628 | 0.535 | |
| Russian | 3.182 | 4.281 | 3.878 | 0.790 | 0.761 | 0.676 | |
| Thai | 2.170 | 2.701 | 73.936 | 0.788 | 0.800 | 0.588 | |
| Indonesian | 0.823 | 1.237 | 1.059 | 0.780 | 0.729 | 0.660 | |
| Arabic | 2.602 | 1.665 | 1.666 | 0.745 | 0.736 | 0.706 | |
| Vietnamese | 1.143 | 0.880 | 73.415 | 0.767 | 0.743 | 0.369 | |
| Turkish | 0.938 | 1.520 | 0.699 | 0.747 | 0.779 | 0.596 | |
| Finnish | 2.784 | 4.666 | 2.964 | 0.859 | 0.835 | 0.759 | |
| Polish | 1.427 | 1.415 | 0.766 | 0.839 | 0.802 | 0.729 | |
| Hindi | 6.444 | 6.962 | 5.827 | 0.797 | 0.818 | 0.730 | |
| Dutch | 1.238 | 1.143 | 0.803 | 0.762 | 0.738 | 0.680 | |
| Czech | 2.929 | 3.875 | 2.108 | 0.802 | 0.796 | 0.685 | |
| Language | Content Consistency | Speaker Similarity | |||
| Ground Truth |
| Ground Truth | ||
| Urdu | 14.819 | 17.822 | 0.775 | - | |
| Tagalog | 5.193 | 6.885 | 0.870 | - | |
| Swedish | 3.760 | 4.813 | 0.822 | - | |
| Danish | 3.636 | 6.403 | 0.775 | - | |
| Hebrew | 7.860 | 16.178 | 0.760 | - | |
| Icelandic | 10.244 | 11.451 | 0.764 | - | |
| Malay | 3.142 | 4.628 | 0.794 | - | |
| Norwegian | 3.613 | 4.442 | 0.825 | - | |
| Persian | 11.113 | 14.469 | 0.800 | - | |
5.2.3 Evaluation of Cross-Lingual Speech Generation
Beyond multilingual voice cloning, Qwen3.5-Omni also supports cross-lingual voice cloning, where the model is required to preserve speaker identity while generating speech in a different target language. We evaluate this capability on the Cross-Lingual benchmark and compare against the CosyVoice series as well as Qwen3-Omni-30B-A3B.
In Table 11, we report the mixed error rate (WER for English and CER for the other languages) across different source–target language pairs. Overall, Qwen3.5-Omni achieves the best performance in 10 out of 12 evaluated directions and sets a new state of the art on most English-, Japanese-, and Korean-targeted pairs. In particular, for zh-to-ko, Qwen3.5-Omni reduces the error rate from 14.4 to 4.03 compared with CosyVoice3, corresponding to an approximately 72% relative reduction. Qwen3.5-Omni also performs strongly on commonly used language pairs such as zh-to-en and en-to-zh, indicating better content consistency under cross-lingual generation. These results demonstrate that Qwen3.5-Omni generalizes effectively across language boundaries while preserving target linguistic accuracy.
| Language | Qwen3.5-Omni-Plus | Qwen3-Omni-30B-A3B | CosyVoice3 | CosyVoice2 |
| English-to-Chinese | 4.86 | 5.37 | 5.09 | 13.5 |
| Japanese-to-Chinese | 3.55 | 3.32 | 3.05 | 48.1 |
| Korean-to-Chinese | 0.84 | 0.99 | 1.06 | 7.70 |
| Chinese-to-English | 2.18 | 2.76 | 2.98 | 6.47 |
| Japanese-to-English | 2.18 | 3.31 | 4.20 | 17.1 |
| Korean-to-English | 2.51 | 3.34 | 4.19 | 11.2 |
| Chinese-to-Japanese | 5.92 | 8.29 | 7.08 | 13.1 |
| English-to-Japanese | 5.12 | 7.53 | 6.80 | 14.9 |
| Korean-to-Japanese | 2.16 | 4.24 | 3.93 | 5.86 |
| Chinese-to-Korean | 4.03 | 5.13 | 14.4 | 24.8 |
| English-to-Korean | 3.72 | 4.96 | 5.87 | 21.9 |
| Japanese-to-Korean | 5.12 | 6.23 | 7.92 | 21.5 |
5.2.4 Evaluation of Custom-Voice Speech Generation
We evaluate the custom-voice speech generation capability of Qwen3.5-Omni in multilingual settings. We compare Qwen3.5-Omni with several strong commercial systems accessed through their official APIs in March 2026, including ElevenLabs Multilingual v2 (9YHcvj6GT2YYXdXww), Gemini-2.5 Pro-Preview-TTS (Achernar), GPT-Audio-2025-08-28 (Alloy), and MiniMax-Speech-2.8-HD (English_expressive_narrator).
| Language | Qwen3.5-Omni-Plus | ElevenLabs | Gemini-2.5 Pro | GPT-Audio | MiniMax |
| Chinese | 0.785 | 3.801 | 1.890 | 0.829 | 0.786 |
| English | 0.839 | 1.126 | 0.953 | 1.050 | 1.429 |
| German | 0.182 | 0.500 | 0.509 | 0.558 | 1.581 |
| Italian | 0.458 | 0.513 | 0.991 | 0.769 | 1.063 |
| Portuguese | 1.581 | 1.109 | 2.050 | 1.506 | 1.240 |
| Spanish | 0.768 | 0.520 | 0.891 | 0.936 | 0.691 |
| Japanese | 3.306 | 11.685 | 4.420 | 4.317 | 4.254 |
| Korean | 1.309 | 3.981 | 4.110 | 3.999 | 3.635 |
| French | 2.724 | 2.574 | 3.284 | 2.809 | 3.439 |
| Russian | 4.723 | 4.324 | 3.858 | 4.346 | 3.529 |
| Thai | 1.653 | 114.813 | 2.539 | 4.430 | 1.811 |
| Indonesian | 1.596 | 6.094 | 1.498 | 2.362 | 1.585 |
| Arabic | 3.183 | 5.400 | 5.525 | 5.326 | 3.309 |
| Vietnamese | 1.320 | 82.849 | 1.699 | 1.854 | 1.058 |
| Turkish | 1.309 | 0.551 | 2.237 | 1.389 | 0.652 |
| Finnish | 4.039 | 2.522 | 5.331 | 3.270 | 2.939 |
| Polish | 1.462 | 0.733 | 1.622 | 1.737 | 0.833 |
| Hindi | 6.776 | 6.388 | 6.596 | 7.191 | 6.146 |
| Dutch | 1.135 | 1.005 | 0.973 | 1.561 | 1.406 |
| Czech | 3.769 | 1.916 | 3.380 | 2.859 | 1.766 |
| Urdu | 14.916 | 12.970 | 14.141 | 13.362 | 24.151 |
| Tagalog | 5.090 | 5.473 | 6.784 | 5.352 | 5.674 |
| Swedish | 3.588 | 3.132 | 3.196 | 2.898 | 2.833 |
| Danish | 7.183 | 2.604 | 3.876 | 3.846 | 4.951 |
| Hebrew | 7.680 | 102.018 | 4.459 | 5.328 | 8.161 |
| Icelandic | 10.322 | 25.110 | 6.348 | 9.648 | 33.431 |
| Malay | 3.738 | 6.448 | 3.731 | 3.406 | 3.955 |
| Norwegian | 5.576 | 7.351 | 4.304 | 3.400 | 9.492 |
| Persian | 12.140 | 20.564 | 12.620 | 13.202 | 12.722 |
As shown in Table 12, although Qwen3.5-Omni is fine-tuned only on monolingual data, it demonstrates strong cross-lingual generalization in custom-voice speech generation. The model is able to transfer the target speaker characteristics to all 29 evaluated languages while maintaining stable generation quality. Overall, Qwen3.5-Omni achieves the best WER in 10 languages and remains competitive in many others. In particular, it shows clear advantages in several challenging languages, including Japanese (3.306) and Korean (1.309), indicating strong intelligibility under cross-lingual voice transfer. These results suggest that Qwen3.5-Omni can generate custom-voice speech with robust linguistic fidelity across a wide range of languages.
6 Conclusion
In this work, we present Qwen3.5-Omni, a fully omnimodal large language model that unifies understanding, reasoning, generation, and action across text, images, audio, and audio-visual inputs. Built on the Thinker–Talker framework, Qwen3.5-Omni introduces efficient Hybrid-Attention MoE architectures, 256k long-context modeling, improved streaming speech generation with multi-codebook codec prediction and ARIA, and substantially expanded multilingual speech support. These advances enable three key capabilities: controllable audio-visual captioning, comprehensive real-time interaction, and native omnimodal agentic behavior through autonomous tool use and audio-visual code generation. Empirically, Qwen3.5-Omni achieves state-of-the-art or highly competitive performance across a broad range of audio and audio-visual benchmarks, while maintaining the strong text and vision capabilities of same-scale Qwen models. These results suggest that scaling native omnimodal training can produce unified systems that not only perceive and reason across modalities, but also interact and act in real time. We hope Qwen3.5-Omni provides a strong foundation for future research on general-purpose omnimodal agents.
References
- P. Anastassiou, J. Chen, J. Chen, Y. Chen, Z. Chen, Z. Chen, J. Cong, L. Deng, C. Ding, L. Gao, et al. (2024) Seed-tts: a family of high-quality versatile speech generation models. arXiv preprint arXiv:2406.02430. Cited by: 1st item, Table 8, Table 8.
- Anthropic (2023a) Claude 2. Technical report Anthropic. External Links: Link Cited by: §1.
- Anthropic (2023b) Introducing Claude. Anthropic. External Links: Link Cited by: §1.
- Anthropic (2024) The Claude 3 model family: Opus, Sonnet, Haiku. Technical report Anthropic, AI. External Links: Link Cited by: §1.
- R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. M. Tyers, and G. Weber (2020) Common voice: A massively-multilingual speech corpus. In Proceedings of The 12th Language Resources and Evaluation Conference, LREC 2020, Marseille, France, May 11-16, 2020, N. Calzolari, F. Béchet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, and S. Piperidis (Eds.), pp. 4218–4222. External Links: Link Cited by: §5.1.
- J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, B. Hui, L. Ji, M. Li, J. Lin, R. Lin, D. Liu, G. Liu, C. Lu, K. Lu, J. Ma, R. Men, X. Ren, X. Ren, C. Tan, S. Tan, J. Tu, P. Wang, S. Wang, W. Wang, S. Wu, B. Xu, J. Xu, A. Yang, H. Yang, J. Yang, S. Yang, Y. Yao, B. Yu, H. Yuan, Z. Yuan, J. Zhang, X. Zhang, Y. Zhang, Z. Zhang, C. Zhou, J. Zhou, X. Zhou, and T. Zhu (2023a) Qwen technical report. CoRR abs/2309.16609. Cited by: §1.
- J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou (2023b) Qwen-VL: a frontier large vision-language model with versatile abilities. CoRR abs/2308.12966. Cited by: §1.
- S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al. (2025a) Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923. Cited by: §1.
- Y. Bai, S. Tu, J. Zhang, H. Peng, X. Wang, X. Lv, S. Cao, J. Xu, L. Hou, Y. Dong, J. Tang, and J. Li (2025b) LongBench v2: towards deeper understanding and reasoning on realistic long-context multitasks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 3639–3664. External Links: Link Cited by: §5.1.
- M. Balunović, J. Dekoninck, I. Petrov, N. Jovanović, and M. Vechev (2025) MathArena: evaluating llms on uncontaminated math competitions. SRI Lab, ETH Zurich. External Links: Link Cited by: §5.1.
- V. Barres, H. Dong, S. Ray, X. Si, and K. Narasimhan (2025) -Bench: evaluating conversational agents in a dual-control environment. External Links: 2506.07982, Link Cited by: §5.1.
- T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. (2020) Language models are few-shot learners. In NeurIPS, Cited by: §1.
- L. Chen, J. Li, X. Dong, P. Zhang, Y. Zang, Z. Chen, H. Duan, J. Wang, Y. Qiao, D. Lin, et al. (2024a) Are we on the right way for evaluating large vision-language models?. arXiv:2403.20330. Cited by: §5.1.
- Y. Chen, X. Yue, C. Zhang, X. Gao, R. T. Tan, and H. Li (2024b) Voicebench: benchmarking llm-based voice assistants. arXiv preprint arXiv:2410.17196. Cited by: §5.1.
- Y. Chen, Z. Niu, Z. Ma, K. Deng, C. Wang, J. Zhao, K. Yu, and X. Chen (2024c) F5-tts: a fairytaler that fakes fluent and faithful speech with flow matching. arXiv preprint arXiv:2410.06885. Cited by: Table 8.
- X. Cheng, W. Zhang, S. Zhang, J. Yang, X. Guan, X. Wu, X. Li, G. Zhang, J. Liu, Y. Mai, Y. Zeng, Z. Wen, K. Jin, B. Wang, W. Zhou, Y. Lu, T. Li, W. Huang, and Z. Li (2025) SimpleVQA: multimodal factuality evaluation for multimodal large language models. CoRR abs/2502.13059. Cited by: §5.1.
- Y. Chu, J. Xu, Q. Yang, H. Wei, X. Wei, Z. Guo, Y. Leng, Y. Lv, J. He, J. Lin, et al. (2024) Qwen2-audio technical report. arXiv preprint arXiv:2407.10759. Cited by: §1.
- Y. Chu, J. Xu, X. Zhou, Q. Yang, S. Zhang, Z. Yan, C. Zhou, and J. Zhou (2023) Qwen-Audio: advancing universal audio understanding via unified large-scale audio-language models. CoRR abs/2311.07919. Cited by: §1.
- G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen, et al. (2025) Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261. Cited by: §1.
- A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna (2022) FLEURS: few-shot learning evaluation of universal representations of speech. 2022 IEEE Spoken Language Technology Workshop (SLT), pp. 798–805. External Links: Link Cited by: 2nd item, §5.1.
- M. Du, B. Wu, Z. Li, X. Huang, and Z. Wei (2024a) EmbSpatial-bench: benchmarking spatial understanding for embodied tasks with large vision-language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), pp. 346–355. Cited by: §5.1.
- Z. Du, C. Gao, Y. Wang, F. Yu, T. Zhao, H. Wang, X. Lv, H. Wang, C. Ni, X. Shi, K. An, G. Yang, Y. Li, Y. Chen, Z. Gao, Q. Chen, Y. Gu, M. Chen, Y. Chen, S. Zhang, W. Wang, and J. Ye (2025) CosyVoice 3: towards in-the-wild speech generation via scaling-up and post-training. CoRR abs/2505.17589. Cited by: 3rd item, Table 8.
- Z. Du, Y. Wang, Q. Chen, X. Shi, X. Lv, T. Zhao, Z. Gao, Y. Yang, C. Gao, H. Wang, et al. (2024b) CosyVoice 2: scalable streaming speech synthesis with large language models. arXiv preprint arXiv:2412.10117. Cited by: Table 8.
- A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Rozière, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. M. Kloumann, I. Misra, I. Evtimov, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, and et al. (2024) The Llama 3 herd of models. CoRR abs/2407.21783. Cited by: §1.
- S. E. Eskimez, X. Wang, M. Thakker, C. Li, C. Tsai, Z. Xiao, H. Yang, Z. Zhu, M. Tang, X. Tan, et al. (2024) E2 tts: embarrassingly easy fully non-autoregressive zero-shot tts. In 2024 IEEE Spoken Language Technology Workshop (SLT), pp. 682–689. Cited by: Table 8.
- C. Fu, Y. Dai, Y. Luo, L. Li, S. Ren, R. Zhang, Z. Wang, C. Zhou, Y. Shen, M. Zhang, et al. (2024) Video-mme: the first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. arXiv:2405.21075. Cited by: §5.1.
- C. Fu, Y. Dai, Y. Luo, L. Li, S. Ren, R. Zhang, Z. Wang, C. Zhou, Y. Shen, M. Zhang, et al. (2025) Video-mme: the first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 24108–24118. Cited by: §5.1.
- A. P. Gema, J. O. J. Leang, G. Hong, A. Devoto, A. C. M. Mancino, R. Saxena, X. He, Y. Zhao, X. Du, M. R. G. Madani, et al. (2024) Are we done with mmlu?. CoRR abs/2406.04127. Cited by: §5.1.
- Gemini Team (2024) Gemini 1.5: unlocking multimodal understanding across millions of tokens of context. Technical report Google. External Links: Link Cited by: §1.
- T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y. Yacoob, D. Manocha, and T. Zhou (2024) Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, pp. 14375–14385. Cited by: §5.1.
- C. Hao, R. Yuan, J. Yao, Q. Deng, X. Bai, W. Xue, and L. Xie (2025) SongFormer: scaling music structure analysis with heterogeneous supervision. External Links: 2510.02797, Link Cited by: §5.1.
- J. Hong, S. Yan, J. Cai, X. Jiang, Y. Hu, and W. Xie (2025) WorldSense: evaluating real-world omnimodal understanding for multimodal llms. CoRR abs/2502.04326. Cited by: §5.1.
- C. Hsu and J. R. Jang (2010) On the improvement of singing voice separation for monaural recordings using the MIR-1K dataset. IEEE Trans. Speech Audio Process. 18 (2), pp. 310–319. External Links: Link, Document Cited by: §5.1.
- Y. Huang, Y. Bai, Z. Zhu, J. Zhang, J. Zhang, T. Su, J. Liu, C. Lv, Y. Zhang, J. Lei, Y. Fu, M. Sun, and J. He (2023) C-Eval: a multi-level multi-discipline chinese evaluation suite for foundation models. In NeurIPS, Cited by: §5.1.
- N. Jain, K. Han, A. Gu, W. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica (2024) LiveCodeBench: holistic and contamination free evaluation of large language models for code. CoRR abs/2403.07974. Cited by: §5.1.
- C. Jiang, J. Sun, Y. Cao, J. Zhuang, H. Li, X. Fan, M. Zhang, J. Ye, S. Dou, Z. Xi, et al. (2025) SpeechRole: a large-scale dataset and benchmark for evaluating speech role-playing agents. arXiv preprint arXiv:2508.02013. Cited by: §5.1.
- S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg (2014) Referitgame: referring to objects in photographs of natural scenes. In EMNLP, Cited by: §5.1.
- A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi (2016) A diagram is worth a dozen images. In ECCV, Cited by: §5.1.
- J. Li, D. Li, S. Savarese, and S. Hoi (2023) Blip-2: bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv:2301.12597. Cited by: §1.
- K. Li, Y. Wang, Y. He, Y. Li, Y. Wang, Y. Liu, Z. Wang, J. Xu, G. Chen, P. Luo, et al. (2024) Mvbench: a comprehensive multi-modal video understanding benchmark. In CVPR, Cited by: §5.1.
- L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J. Hwang, K. Chang, and J. Gao (2022) Grounded language-image pre-training. In CVPR, pp. 10955–10965. Cited by: §5.1.
- X. Li, W. Jiao, J. Jin, S. Wang, G. Dong, J. Jin, H. Wang, Y. Wang, J. Wen, Y. Lu, et al. (2026) OmniGAIA: towards native omni-modal ai agents. arXiv preprint arXiv:2602.22897. Cited by: §5.1.
- B. Liu, L. Zhan, L. Xu, L. Ma, Y. Yang, and X. Wu (2021) Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. In 18th IEEE International Symposium on Biomedical Imaging, ISBI 2021, Nice, France, April 13-16, 2021, pp. 1650–1654. Cited by: §5.1.
- H. Liu, C. Li, Q. Wu, and Y. J. Lee (2023) Visual instruction tuning. arXiv:2304.08485. Cited by: §1.
- Y. Liu, Z. Li, M. Huang, B. Yang, W. Yu, C. Li, X. Yin, C. Liu, L. Jin, and X. Bai (2024) OCRBench: on the hidden mystery of ocr in large multimodal models. Science China Information Sciences 67 (12). External Links: ISSN 1869-1919, Link, Document Cited by: §5.1.
- P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K. Chang, M. Galley, and J. Gao (2024) MathVista: evaluating mathematical reasoning of foundation models in visual contexts. In ICLR, Cited by: §5.1.
- T. Luong, D. Hwang, H. H. Nguyen, G. Ghiasi, Y. Chervonyi, I. Seo, J. Kim, G. Bingham, J. Lee, S. Mishra, A. Zhai, C. H. Hu, H. Michalewski, J. Kim, J. Ahn, J. Bae, X. Song, T. H. Trinh, Q. V. Le, and J. Jung (2025) Towards robust mathematical reasoning. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, External Links: Link Cited by: §5.1.
- Y. Ma, Y. Zang, L. Chen, M. Chen, Y. Jiao, X. Li, X. Lu, Z. Liu, Y. Ma, X. Dong, P. Zhang, L. Pan, Y. Jiang, J. Wang, Y. Cao, and A. Sun (2024) MMLONGBENCH-DOC: benchmarking long-context document understanding with visualizations. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang (Eds.), Cited by: §5.1.
- Z. Ma, Y. Ma, Y. Zhu, C. Yang, Y. Chao, R. Xu, W. Chen, Y. Chen, Z. Chen, J. Cong, K. Li, K. Li, S. Li, X. Li, X. Li, Z. Lian, Y. Liang, M. Liu, Z. Niu, T. Wang, Y. Wang, Y. Wang, Y. Wu, G. Yang, J. Yu, R. Yuan, Z. Zheng, Z. Zhou, H. Zhu, W. Xue, E. Benetos, K. Yu, C. E. Siong, and X. Chen (2025a) MMAR: A challenging benchmark for deep reasoning in speech, audio, music, and their mix. CoRR abs/2505.13032. External Links: Link, Document, 2505.13032 Cited by: §5.1.
- Z. Ma, R. Xu, Z. Xing, Y. Chu, Y. Wang, J. He, J. Xu, P. Heng, K. Yu, J. Lin, et al. (2025b) Omni-captioner: data pipeline, models, and benchmark for omni detailed perception. arXiv preprint arXiv:2510.12720. Cited by: §5.1.
- L. T. P. Nguyen, Z. Yu, S. L. Y. Hang, S. An, J. Lee, Y. Ban, S. Chung, T. Nguyen, J. Maeng, S. Lee, et al. (2025) See, hear, and understand: benchmarking audiovisual human speech understanding in multimodal large language models. arXiv preprint arXiv:2512.02231. Cited by: §5.1.
- OpenAI (2022) ChatML. External Links: Link Cited by: §4.1.
- OpenAI (2023) GPT4 technical report. CoRR abs/2303.08774. Cited by: §1.
- OpenAI (2024) Hello GPT-4o. External Links: Link Cited by: §1.
- R. Paiss, A. Ephrat, O. Tov, S. Zada, I. Mosseri, M. Irani, and T. Dekel (2023) Teaching CLIP to count to ten. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023, pp. 3147–3157. Cited by: §5.1.
- V. Panayotov, G. Chen, D. Povey, and S. Khudanpur (2015) Librispeech: an ASR corpus based on public domain audio books. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2015, South Brisbane, Queensland, Australia, April 19-24, 2015, Cited by: §5.1.
- R. Pourreza, R. Dagli, A. Bhattacharyya, S. Panchal, G. Berger, and R. Memisevic (2025) Can vision-language models answer face to face questions in the real-world?. arXiv preprint arXiv:2503.19356. Cited by: §5.1.
- V. Pyatkin, S. Malik, V. Graf, H. Ivison, S. Huang, P. Dasigi, N. Lambert, and H. Hajishirzi (2025) Generalizing verifiable instruction following. CoRR abs/2507.02833. External Links: Link, Document, 2507.02833 Cited by: §5.1.
- R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn (2023) Direct preference optimization: your language model is secretly a reward model. In NeurIPS, Cited by: item (3).
- D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman (2023) GPQA: a graduate-level Google-proof Q&A benchmark. CoRR abs/2311.12022. Cited by: §5.1.
- J. Roberts, M. R. Taesiri, A. Sharma, A. Gupta, S. Roberts, I. Croitoru, S. Bogolin, J. Tang, F. Langer, V. Raina, V. Raina, H. Xiong, V. Udandarao, J. Lu, S. Chen, S. Purkis, T. Yan, W. Lin, G. Shin, Q. Yang, A. T. Nguyen, K. Han, and S. Albanie (2025) ZeroBench: an impossible visual benchmark for contemporary large multimodal models. CoRR abs/2502.09696. Cited by: §5.1.
- S. Sakshi, U. Tyagi, S. Kumar, A. Seth, R. Selvakumar, O. Nieto, R. Duraiswami, S. Ghosh, and D. Manocha (2024) MMAU: a massive multi-task audio understanding and reasoning benchmark. External Links: 2410.19168, Link Cited by: §5.1.
- Y. Shi, H. Wang, W. Xie, H. Zhang, L. Zhao, Y. Zhang, X. Li, C. Fu, Z. Wen, W. Liu, Z. Zhang, X. Chen, B. Zeng, S. Yang, Y. Zhang, P. Wan, H. Wang, and W. Yang (2025) MME-videoocr: evaluating ocr-based capabilities of multimodal llms in video scenarios. CoRR abs/2505.21333. Cited by: §5.1.
- Z. Tang, D. Wang, Y. Xu, J. Sun, X. Lei, S. Zhao, C. Wen, X. Tan, C. Xie, S. Zhou, R. Yan, C. Lv, Y. Han, W. Zou, and X. Li (2021) KeSpeech: an open source speech dataset of mandarin and its eight subdialects. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual, J. Vanschoren and S. Yeung (Eds.), External Links: Link Cited by: §5.1.
- A. A. Team (2025a) Artificial analysis long context reasoning benchmark (lcr). Note: Artificial Analysis, Inc.Dataset Cited by: §5.1.
- G. R. Team (2025b) Gemini robotics: bringing AI into the physical world. CoRR abs/2503.20020. Cited by: §5.1.
- M.-A-P. Team, X. Du, Y. Yao, K. Ma, B. Wang, T. Zheng, K. Zhu, M. Liu, Y. Liang, X. Jin, Z. Wei, C. Zheng, K. Deng, S. Jia, S. Jiang, Y. Liao, R. Li, Q. Li, S. Li, Y. Li, Y. Li, D. Ma, Y. Ni, H. Que, Q. Wang, Z. Wen, S. Wu, T. Xing, M. Xu, Z. Yang, Z. M. Wang, J. Zhou, Y. Bai, X. Bu, C. Cai, L. Chen, Y. Chen, C. Cheng, T. Cheng, K. Ding, S. Huang, Y. Huang, Y. Li, Y. Li, Z. Li, T. Liang, C. Lin, H. Lin, Y. Ma, T. Pang, Z. Peng, Z. Peng, Q. Qi, S. Qiu, X. Qu, S. Quan, Y. Tan, Z. Wang, C. Wang, H. Wang, Y. Wang, Y. Wang, J. Xu, K. Yang, R. Yuan, Y. Yue, T. Zhan, C. Zhang, J. Zhang, X. Zhang, X. Zhang, Y. Zhang, Y. Zhao, X. Zheng, C. Zhong, Y. Gao, Z. Li, D. Liu, Q. Liu, T. Liu, S. Ni, J. Peng, Y. Qin, W. Su, G. Wang, S. Wang, J. Yang, M. Yang, M. Cao, X. Yue, Z. Zhang, W. Zhou, J. Liu, Q. Lin, W. Huang, and G. Zhang (2025) SuperGPQA: scaling LLM evaluation across 285 graduate disciplines. CoRR abs/2502.14739. External Links: Link, Document, 2502.14739 Cited by: §5.1.
- Q. Team (2026) Qwen3.5: accelerating productivity with native multimodal agents. External Links: Link Cited by: §2.3.
- H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. (2023) Llama 2: open foundation and fine-tuned chat models. arXiv:2307.09288. Cited by: §1.
- D. Wang, J. Wu, J. Li, D. Yang, X. Chen, T. Zhang, and H. Meng (2025a) MMSU: A massive multi-task spoken language understanding and reasoning benchmark. CoRR abs/2506.04779. External Links: Link, Document, 2506.04779 Cited by: §5.1.
- K. Wang, J. Pan, W. Shi, Z. Lu, M. Zhan, and H. Li (2024a) Measuring multimodal mathematical reasoning with math-vision dataset. arXiv:2402.14804. Cited by: §5.1.
- W. Wang, Z. He, W. Hong, Y. Cheng, X. Zhang, J. Qi, S. Huang, B. Xu, Y. Dong, M. Ding, and J. Tang (2024b) LVBench: an extreme long video understanding benchmark. CoRR abs/2406.08035. Cited by: §5.1.
- X. Wang, M. Jiang, Z. Ma, Z. Zhang, S. Liu, L. Li, Z. Liang, Q. Zheng, R. Wang, X. Feng, W. Bian, Z. Ye, S. Cheng, R. Yuan, Z. Zhao, X. Zhu, J. Pan, L. Xue, P. Zhu, Y. Chen, Z. Li, X. Chen, L. Xie, Y. Guo, and W. Xue (2025b) Spark-tts: an efficient llm-based text-to-speech model with single-stream decoupled speech tokens. CoRR abs/2503.01710. Cited by: Table 8.
- Y. Wang, X. Wang, P. Zhu, J. Wu, H. Li, H. Xue, Y. Zhang, L. Xie, and M. Bi (2022) Opencpop: A high-quality open source chinese popular song corpus for singing voice synthesis. In 23rd Annual Conference of the International Speech Communication Association, Interspeech 2022, Incheon, Korea, September 18-22, 2022, H. Ko and J. H. L. Hansen (Eds.), pp. 4242–4246. External Links: Link, Document Cited by: §5.1.
- Y. Wang, H. Zhan, L. Liu, R. Zeng, H. Guo, J. Zheng, Q. Zhang, X. Zhang, S. Zhang, and Z. Wu (2024c) Maskgct: zero-shot text-to-speech with masked generative codec transformer. arXiv preprint arXiv:2409.00750. Cited by: Table 8.
- Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, T. Li, M. Ku, K. Wang, A. Zhuang, R. Fan, X. Yue, and W. Chen (2024d) MMLU-Pro: A more robust and challenging multi-task language understanding benchmark. CoRR abs/2406.01574. Cited by: §5.1.
- Z. Wang, M. Xia, L. He, H. Chen, Y. Liu, R. Zhu, K. Liang, X. Wu, H. Liu, S. Malladi, A. Chevalier, S. Arora, and D. Chen (2024e) CharXiv: charting gaps in realistic chart understanding in multimodal llms. arXiv preprint arXiv:2406.18521. Cited by: §5.1.
- J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y. Fan, K. Dang, et al. (2025a) Qwen2. 5-omni technical report. arXiv preprint arXiv:2503.20215. Cited by: §1, §1, §2.1, Table 8.
- J. Xu, Z. Guo, H. Hu, Y. Chu, X. Wang, J. He, Y. Wang, X. Shi, T. He, X. Zhu, Y. Lv, Y. Wang, D. Guo, H. Wang, L. Ma, P. Zhang, X. Zhang, H. Hao, Z. Guo, B. Yang, B. Zhang, Z. Ma, X. Wei, S. Bai, K. Chen, X. L. Liu, P. Wang, M. Yang, D. Liu, X. Ren, B. Zheng, R. Men, F. Zhou, B. Yu, J. Yang, L. Yu, J. Zhou, and J. Lin (2025b) Qwen3-omni technical report. ArXiv abs/2509.17765. Cited by: §1, §1, 3rd item, 5th item, §2.1, §2.3, §2.4, §3, §3, Table 8.
- F. Yan, H. Mao, C. C. Ji, T. Zhang, S. G. Patil, I. Stoica, and J. E. Gonzalez (2024) Berkeley function calling leaderboard. Note: https://gorilla.cs.berkeley.edu/blogs/8_berkeley_function_calling_leaderboard.html Cited by: §5.1.
- R. Yan, X. Li, W. Chen, Z. Niu, C. Yang, Z. Ma, K. Yu, and X. Chen (2025) URO-bench: a comprehensive benchmark for end-to-end spoken dialogue models. arXiv preprint arXiv:2502.17810. Cited by: §5.1.
- A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al. (2025a) Qwen3 technical report. arXiv preprint arXiv:2505.09388. Cited by: §1.
- A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, et al. (2024a) Qwen2 technical report. arXiv:2407.10671. Cited by: §1.
- Y. Yang, J. Zhuang, G. Sun, C. Tang, Y. Li, P. Li, Y. Jiang, W. Li, Z. Ma, and C. Zhang (2025b) Audio-centric video understanding benchmark without text shortcut. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 6580–6598. Cited by: §5.1.
- Z. Yang, J. Tang, Z. Li, P. Wang, J. Wan, H. Zhong, X. Liu, M. Yang, P. Wang, S. Bai, L. Jin, and J. Lin (2024b) CC-OCR: A comprehensive and challenging OCR benchmark for evaluating large multimodal models in literacy. CoRR abs/2412.02210. Cited by: §5.1.
- X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al. (2023) Mmmu: a massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. arXiv:2311.16502. Cited by: §5.1.
- X. Yue, T. Zheng, Y. Ni, Y. Wang, K. Zhang, S. Tong, Y. Sun, M. Yin, B. Yu, G. Zhang, et al. (2024) MMMU-pro: a more robust multi-discipline multimodal understanding benchmark. arXiv preprint arXiv:2409.02813. Cited by: §5.1.
- Y. Zang, S. O’Brien, T. Berg-Kirkpatrick, J. McAuley, and Z. Novack (2025) Are you really listening? boosting perceptual awareness in music-qa benchmarks. arXiv preprint arXiv:2504.00369. Cited by: §5.1.
- B. Zhang, H. Lv, P. Guo, Q. Shao, C. Yang, L. Xie, X. Xu, H. Bu, X. Chen, C. Zeng, D. Wu, and Z. Peng (2022) WENETSPEECH: A 10000+ hours multi-domain mandarin corpus for speech recognition. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2022, Virtual and Singapore, 23-27 May 2022, pp. 6182–6186. External Links: Link, Document Cited by: §5.1.
- B. Zhang, C. Guo, G. Yang, H. Yu, H. Zhang, H. Lei, J. Mai, J. Yan, K. Yang, M. Yang, P. Huang, R. Jin, S. Jiang, W. Cheng, Y. Li, Y. Xiao, Y. Zhou, Y. Zhang, Y. Lu, and Y. He (2025a) MiniMax-speech: intrinsic zero-shot text-to-speech with a learnable speaker encoder. CoRR abs/2505.07916. Cited by: 2nd item, 4th item, Table 8.
- L. Zhang, J. Zhang, B. Lei, C. Wu, A. Liu, W. Jia, and X. Zhou (2025b) WildSpeech-bench: benchmarking end-to-end speechllms in the wild. External Links: 2506.21875 Cited by: §5.1.
- X. Zhang, C. Wu, Z. Zhao, W. Lin, Y. Zhang, Y. Wang, and W. Xie (2023) PMC-VQA: visual instruction tuning for medical visual question answering. CoRR abs/2305.10415. Cited by: §5.1.
- X. L. T. D. Zhang, G. Wang, J. Xue, K. Fang, L. Zhao, R. Ma, S. Ren, S. Liu, T. Guo, W. Zhuang, X. Zhang, X. Song, Y. Yan, Y. He, Cici, B. Shen, C. Zhu, C. Ma, C. Chen, H. Chen, J. Li, L. Li, M. Zhu, P. Li, Q. Wang, S. Deng, W. Xiong, W. Huang, W. Yang, Y. Jiang, Y. Yang, Y. Tian, Y. Ma, Y. Yu, Z. Zhang, Z. Yue, B. Xiao, B. Xia, B. Gao, B. Ye, C. Cai, C. Liu, C. He, C. Li, D. Zhu, D. Zhang, F. Shi, G. Wang, H. Zhang, H. Lv, H. Li, H. Tian, H. Qu, H. Xu, H. Zhang, H. Liu, J. Duo, J. Zuo, J. Wei, J. Xiao, J. Dong, J. Shi, J. Hu, K. Bao, K. Zhou, L. Zhang, M. Chen, N. Chen, P. Zhang, Q. Chen, Q. Wang, R. Li, S. Liu, S. Wang, S. Li, S. Yu, S. Cao, S. Chen, S. Gu, W. Wang, W. Ma, X. Deng, X. Yong, X. Zhang, X. Wang, Y. Song, Y. Zhao, Y. Zhao, Y. Gao, Y. Cheng, Y. Tu, Y. Wang, Z. Huang, Z. Tang, Z. Lin, Z. Song, Z. Xu, Z. Zheng, and Z. Jiang (2025c) MiMo-audio: audio language models are few-shot learners. ArXiv abs/2512.23808. External Links: Link Cited by: Table 8.
- Y. Zhang, H. Zhang, H. Tian, C. Fu, S. Zhang, J. Wu, F. Li, K. Wang, Q. Wen, Z. Zhang, et al. (2024) MME-realworld: could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans?. arXiv preprint arXiv:2408.13257. Cited by: §5.1.
- Y. Zhao, H. Zhang, L. Xie, T. Hu, G. Gan, Y. Long, Z. Hu, W. Chen, C. Li, Z. Xu, C. Wang, Z. Shangguan, Z. Liang, Y. Liu, C. Zhao, and A. Cohan (2025) MMVU: measuring expert-level multi-discipline video understanding. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025, Nashville, TN, USA, June 11-15, 2025, pp. 8475–8489. Cited by: §5.1.
- C. Zheng, S. Liu, M. Li, X. Chen, B. Yu, C. Gao, K. Dang, Y. Liu, R. Men, A. Yang, et al. (2025) Group sequence policy optimization. arXiv preprint arXiv:2507.18071. Cited by: item (3).
- J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, and L. Hou (2023) Instruction-following evaluation for large language models. CoRR abs/2311.07911. Cited by: §5.1.
- J. Zhou, Y. Shu, B. Zhao, B. Wu, Z. Liang, S. Xiao, M. Qin, X. Yang, Y. Xiong, B. Zhang, T. Huang, and Z. Liu (2025a) MLVU: benchmarking multi-task long video understanding. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025, Nashville, TN, USA, June 11-15, 2025, pp. 13691–13701. Cited by: §5.1.
- Z. Zhou, R. Wang, and Z. Wu (2025b) Daily-omni: towards audio-visual reasoning with temporal alignment across modalities. CoRR abs/2505.17862. Cited by: §5.1.
- D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny (2023) Minigpt-4: enhancing vision-language understanding with advanced large language models. arXiv:2304.10592. Cited by: §1.
- C. Zou, X. Guo, R. Yang, J. Zhang, B. Hu, and H. Zhang (2025) DynaMath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models. In ICLR, Cited by: §5.1.
- Y. Zuo, S. Qu, Y. Li, Z. Chen, X. Zhu, E. Hua, K. Zhang, N. Ding, and B. Zhou (2025) MedXpertQA: benchmarking expert-level medical reasoning and understanding. In ICML, A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu (Eds.), Proceedings of Machine Learning Research. Cited by: §5.1.
7 Authors
Core Contributors222Alphabetical order. * denotes the corresponding author.
Bing Han
Baosong Yang
Bin Zhang
Bo Zheng
Dayiheng Liu
Fan Zhou
Hongkun Hao
Hangrui Hu
Jin Xu∗
Jianxin Yang
Jingren Zhou
Keqin Chen
Le Yu
Mingkun Yang
Peng Wang
Pei Zhang
Qize Yang
Rui Men
Ruiyang Xu
Shuai Bai
Sibo Song
Ting He
Xize Cheng
Xuejing Liu
Xingzhang Ren
Xian Shi
Xiong Wang
Xinyu Zhang
Xinfa Zhu
Yunfei Chu
Yuanjun Lv
Yuchong Sun
Yongqi Wang
Yuxuan Wang
Yang Zhang
Zhifang Guo
Zishan Guo
Ziyang Ma
Contributors††footnotemark:
Andong Chen
Anfeng Li
An Yang
Bei Chen
Bin Lin
Bingshen Mu
Bohan Wang
Buxiao Wu
Bowen Xu
Beichen Zhang
Cheng Chen
Chang Gao
Chengen Huang
Chenyang Le
Chenhao Li
Chenglong Liu
Chenxu Lv
Chen Qiang
Chenfei Wu
Chenhan Yuan
Chengruidong Zhang
Chujie Zheng
Daren Chen
Dake Guo
Fei Huang
Gaoji Liu
Guangdong Zhou
Hao Ge
Huiqiang Jiang
Haoran Lian
Hongjian Tu
Hao Yu
Hang Zhang
Hao Zhou
Haiquan Zhao
Humen Zhong
Jiawei Chen
Jian Guan
Jiayi Leng
Jiahao Li
Junrong Lin
Jiawei Liu
Jialong Tang
Jun Tang
Jianhong Tu
Jianqiang Wan
Jinxi Wei
Jianwei Zhang
Jing Zhou
Kai Dang
Kangxiang Xia
Kun Yan
Kexin Yang
Lianghao Deng
Lulu Hu
Linhan Ma
Lingchen Meng
Lei Xie
Laiwen Zheng
Miao Hong
Mei Li
Mingcheng Li
Mingze Li
Minsheng Li
Minghao Wu
Mingfeng Xue
Na Ni
Peng Liu
Peng Wang
Pengfei Wang
Peiyang Zhang
Qidong Huang
Qingfeng Lan
Qintong Li
Que Shen
Qiuyue Wang
Qin Zhu
Ruisheng Cao
Rongyao Fang
Rui Hu
Ruibin Yuan
Song Chen
Su Hao
Shen Li
Shixuan Liu
Shurui Li
Siqi Zhang
Tianyi Tang
Tingyu Xia
Wei Ding
Wenbin Ge
Weizhou Shen
Wei Wang
Wentao Yao
Xi Chen
Xiaotong Chen
Xionghui Chen
Xiaodong Deng
Xudong Guo
Xin Le
Xiao Li
Xie Chen
Xinyao Niu
Xuancheng Ren
Xuechun Wang
Xuwu Wang
Xingzhe Wu
Xipin Wei
Xiao Xu
Xian Yang
Yuxuan Cai
Yizhong Cao
Yilei Chen
Yuxiang Chen
Yiming Dong
Yang Fan
Yanpeng Li
Yucheng Li
Yang Liu
Yantao Liu
Yuqiong Liu
Yuxuan Liu
Yuyan Luo
Yubo Ma
Yang Su
Yuezhang Wang
Yuhao Wang
Yi Wu
Yunbao Wu
Yu Xi
Yi Zhang
Yichang Zhang
Yinger Zhang
Yuxiang Zheng
Zeyu Cui
Ziwei Ji
Ziyue Jiang
Zhaohai Li
Zheng Li
Zhi Li
Zihan Qiu
Zekun Wang
Zhihai Wang
Zhenghao Xing
Zhibo Yang
Zhuorui Ye
Zhenru Zhang
Zhipeng Zhou
Zhengyang Zhuge
8 Appendix
8.1 Detailed Multilingual Evaluation Results
Multilingual ASR.
As presented in Table 13, Qwen3.5-Omni demonstrates superior speech recognition capabilities compared to state-of-the-art competitors on the FLEURS test set. Qwen3.5-Omni-Plus achieves the lowest average WER of 6.6%, outperforming both Gemini-3.1-Pro (7.3%) and GPT-4o-Transcribe (10.4%). It secures the best performance in the majority of languages, with particularly significant margins in complex tonal and low-resource languages such as Cantonese (2.2% vs. 6.3% for Gemini-3.1-Pro), Thai, and Vietnamese. Meanwhile, Qwen3.5-Omni-Flash offers a highly efficient alternative, achieving an average WER of 10.8% that remains competitive against Gemini-3-Flash (10.5%). Notably, Qwen3.5-Omni-Flash exhibits exceptional robustness in challenging scenarios, drastically reducing errors in Cantonese (3.1% vs. 10.8% for Gemini-3-Flash) and maintaining strong performance in Japanese and Korean, thereby highlighting its advantage for high-value Asian language pairs.
Multilingual Translation.
As shown in Tables 14 and 15, the Qwen3.5-Omni series demonstrates distinct advantages over state-of-the-art competitors on the FLEURS test set, particularly in Asian languages and specific high-resource pairs. Qwen3.5-Omni-Plus exhibits comprehensive superiority over Gemini-3.1-Pro in the many-to-many directions (en2xx/zh2xx), achieving higher average BLEU scores in both English-to-XX (33.8 vs. 31.8) and Chinese-to-XX (21.4 vs. 19.6). It also leads in key xx2en pairs such as Portuguese (49.4 vs. 47.7) and Indonesian (45.7 vs. 45.1). Although Gemini-3.1-Pro holds a slight edge in overall xx2zh averages, Qwen3.5-Omni-Plus significantly outperforms it in critical Asian languages, including Cantonese (+15.6 BLEU), Korean, and Japanese. Similarly, Qwen3.5-Omni-Flash shows targeted strengths against Gemini-3-Flash. While maintaining competitive general performance, it vastly surpasses Gemini in Cantonese translation across all directions (e.g., 37.5 vs. 22.4 in xx2zh and 37.3 vs. 26.7 in en2xx) and delivers better results in Japanese and Korean xx2zh tasks. These results underscore Qwen3.5-Omni’s robust optimization for complex Asian linguistic structures and key regional languages.
| Language |
|
|
|
|
| ||||||||||
| Chinese | 2.9 | 2.9 | 3.6 | 2.6 | 4.6 | ||||||||||
| English | 3.2 | 3.7 | 2.7 | 3.2 | 2.9 | ||||||||||
| Cantonese | 2.2 | 3.1 | 6.3 | 5.2 | 10.8 | ||||||||||
| Arabic | 11.7 | 13.6 | 9.2 | 13.0 | 10.1 | ||||||||||
| German | 2.0 | 2.5 | 2.7 | 2.3 | 3.3 | ||||||||||
| French | 2.6 | 3.3 | 3.7 | 3.7 | 3.9 | ||||||||||
| Spanish | 2.2 | 2.4 | 2.5 | 2.3 | 2.7 | ||||||||||
| Portuguese | 2.1 | 2.2 | 2.6 | 2.3 | 2.9 | ||||||||||
| Indonesian | 1.6 | 2.4 | 2.5 | 3.5 | 2.8 | ||||||||||
| Italian | 0.8 | 1.0 | 1.1 | 1.4 | 1.8 | ||||||||||
| Korean | 1.7 | 2.1 | 2.0 | 2.1 | 2.4 | ||||||||||
| Russian | 3.1 | 3.6 | 3.4 | 3.7 | 3.9 | ||||||||||
| Thai | 2.8 | 3.2 | 4.3 | 4.9 | 4.5 | ||||||||||
| Vietnamese | 1.9 | 2.5 | 2.5 | 3.5 | 3.5 | ||||||||||
| Japanese | 1.9 | 2.5 | 2.3 | 3.0 | 3.4 | ||||||||||
| Turkish | 3.1 | 4.4 | 3.8 | 4.2 | 4.4 | ||||||||||
| Hindi | 9.7 | 9.9 | 4.5 | 12.0 | 5.6 | ||||||||||
| Malay | 2.7 | 4.2 | 3.7 | 4.1 | 6.2 | ||||||||||
| Dutch | 2.8 | 3.5 | 3.5 | 3.7 | 4.7 | ||||||||||
| Urdu | 20.8 | 31.9 | 25.2 | 19.7 | 23.0 | ||||||||||
| Norwegian | 3.9 | 5.2 | 5.0 | 5.5 | 6.7 | ||||||||||
| Swedish | 3.1 | 5.0 | 4.6 | 5.2 | 7.7 | ||||||||||
| Danish | 3.5 | 5.3 | 5.7 | 6.5 | 7.9 | ||||||||||
| Hebrew | 12.5 | 16.6 | 15.6 | 19.4 | 20.2 | ||||||||||
| Finnish | 2.4 | 4.5 | 3.4 | 3.8 | 5.2 | ||||||||||
| Polish | 1.9 | 3.1 | 2.7 | 2.8 | 4.9 | ||||||||||
| Icelandic | 3.6 | 8.9 | 4.7 | 10.8 | 6.8 | ||||||||||
| Czech | 2.6 | 4.5 | 3.8 | 4.7 | 8.0 | ||||||||||
| Filipino | 5.1 | 7.1 | 7.6 | 7.3 | 8.5 | ||||||||||
| Persian | 12.0 | 12.1 | 8.9 | 9.9 | 10.0 | ||||||||||
| Greek | 4.7 | 8.1 | 5.4 | 6.5 | 7.6 | ||||||||||
| Afrikaans | 10.6 | 13.7 | 12.7 | 17.9 | 18.6 | ||||||||||
| Asturian | 15.8 | 25.9 | 23.7 | 23.8 | 48.3 | ||||||||||
| Belarusian | 6.7 | 12.2 | 6.7 | 10.2 | 10.9 | ||||||||||
| Bulgarian | 6.2 | 10.7 | 5.3 | 7.0 | 7.9 | ||||||||||
| Bengali | 16.2 | 19.8 | 21.9 | 24.1 | 21.9 | ||||||||||
| Bosnian | 5.4 | 9.5 | 6.0 | 13.9 | 11.2 | ||||||||||
| Catalan | 2.8 | 6.3 | 2.7 | 2.7 | 6.4 | ||||||||||
| Cebuano | 10.5 | 16.6 | 13.0 | 15.1 | 12.8 | ||||||||||
| Estonian | 6.7 | 16.6 | 4.9 | 7.6 | 7.3 | ||||||||||
| Galician | 5.0 | 8.6 | 4.9 | 6.8 | 14.6 | ||||||||||
| Gujarati | 13.9 | 18.4 | 14.8 | 26.9 | 16.3 | ||||||||||
| Croatian | 5.4 | 9.0 | 5.1 | 16.4 | 9.3 | ||||||||||
| Hungarian | 4.9 | 10.6 | 5.5 | 7.4 | 11.3 | ||||||||||
| Javanese | 11.8 | 18.3 | 14.1 | 24.7 | 15.7 | ||||||||||
| Kazakh | 6.3 | 16.6 | 6.2 | 11.5 | 12.7 | ||||||||||
| Kannada | 16.0 | 23.8 | 16.3 | 28.1 | 16.3 | ||||||||||
| Kyrgyz | 10.0 | 19.7 | 8.3 | 20.7 | 16.3 | ||||||||||
| Latvian | 6.7 | 17.8 | 3.7 | 6.3 | 6.8 | ||||||||||
| Macedonian | 4.1 | 7.9 | 4.0 | 6.2 | 7.8 | ||||||||||
| Malayalam | 18.8 | 27.0 | 18.3 | 33.5 | 20.3 | ||||||||||
| Marathi | 16.3 | 23.6 | 15.3 | 26.3 | 16.0 | ||||||||||
| Punjabi | 13.7 | 24.6 | 14.4 | 36.4 | 17.4 | ||||||||||
| Romanian | 3.2 | 6.1 | 3.4 | 4.5 | 5.9 | ||||||||||
| Slovak | 3.3 | 5.5 | 2.8 | 3.6 | 6.5 | ||||||||||
| Slovenian | 6.1 | 14.3 | 6.3 | 8.8 | 10.0 | ||||||||||
| Swahili | 9.4 | 17.5 | 9.9 | 16.3 | 10.7 | ||||||||||
| Tajik | 10.0 | 41.1 | 20.4 | 20.2 | 53.1 | ||||||||||
| Azerbaijani | 7.2 | 13.0 | 5.8 | 10.6 | 13.2 | ||||||||||
| Ukrainian | 3.2 | 5.4 | 3.4 | 4.3 | 5.1 | ||||||||||
| Average | 6.6 | 10.8 | 7.3 | 10.4 | 10.5 |
| en2xx (English → Other Languages) | zh2xx (Chinese → Other Languages) | |||||||||||||||||||||||
| Language |
|
|
|
|
|
|
|
| ||||||||||||||||
| Chinese | 47.8 | 46.6 | 47.4 | 46.3 | – | – | – | – | ||||||||||||||||
| English | – | – | – | – | 32.2 | 31.2 | 30.1 | 29.5 | ||||||||||||||||
| Cantonese | 40.1 | 37.3 | 25.5 | 26.7 | 36.7 | 35.9 | 23.7 | 24.0 | ||||||||||||||||
| Arabic | 31.1 | 28.2 | 27.0 | 28.5 | 16.1 | 13.9 | 14.2 | 14.4 | ||||||||||||||||
| German | 43.2 | 39.6 | 41.8 | 40.9 | 23.2 | 20.8 | 22.0 | 21.4 | ||||||||||||||||
| French | 50.9 | 48.8 | 48.0 | 47.8 | 30.7 | 28.8 | 29.4 | 29.2 | ||||||||||||||||
| Spanish | 29.1 | 28.9 | 28.3 | 28.8 | 22.2 | 20.4 | 20.8 | 20.6 | ||||||||||||||||
| Portuguese | 51.2 | 48.6 | 47.3 | 47.2 | 28.5 | 26.8 | 25.7 | 25.6 | ||||||||||||||||
| Indonesian | 45.3 | 43.7 | 42.1 | 41.7 | 28.8 | 26.9 | 25.4 | 25.2 | ||||||||||||||||
| Italian | 32.7 | 30.7 | 31.9 | 30.9 | 23.1 | 21.1 | 21.9 | 21.3 | ||||||||||||||||
| Korean | 33.9 | 31.8 | 30.8 | 31.7 | 25.1 | 23.4 | 21.9 | 22.8 | ||||||||||||||||
| Russian | 33.8 | 31.8 | 33.1 | 33.2 | 21.5 | 18.9 | 20.0 | 19.9 | ||||||||||||||||
| Thai | 65.4 | 62.9 | 64.8 | 64.4 | 58.0 | 55.5 | 57.2 | 56.4 | ||||||||||||||||
| Vietnamese | 43.0 | 41.8 | 41.1 | 40.2 | 31.6 | 30.5 | 28.1 | 28.4 | ||||||||||||||||
| Japanese | 53.2 | 50.6 | 51.3 | 50.8 | 45.6 | 41.6 | 43.0 | 42.0 | ||||||||||||||||
| Turkish | 30.4 | 27.6 | 29.3 | 29.0 | 16.8 | 14.6 | 15.9 | 16.0 | ||||||||||||||||
| Hindi | 33.1 | 29.1 | 28.2 | 29.2 | 19.1 | 14.3 | 17.6 | 17.5 | ||||||||||||||||
| Malay | 39.6 | 37.2 | 35.9 | 36.0 | 24.1 | 21.7 | 21.0 | 20.5 | ||||||||||||||||
| Dutch | 30.1 | 28.2 | 28.1 | 28.8 | 21.0 | 18.8 | 19.3 | 19.1 | ||||||||||||||||
| Urdu | 25.0 | 22.1 | 23.0 | 23.0 | 15.5 | 8.6 | 14.9 | 14.7 | ||||||||||||||||
| Norwegian | 35.3 | 32.8 | 33.1 | 33.7 | 20.3 | 17.8 | 18.4 | 18.9 | ||||||||||||||||
| Swedish | 47.5 | 44.1 | 45.7 | 45.8 | 25.4 | 23.0 | 23.7 | 24.1 | ||||||||||||||||
| Danish | 48.4 | 45.2 | 45.7 | 45.2 | 25.7 | 22.8 | 23.5 | 23.2 | ||||||||||||||||
| Hebrew | 36.4 | 29.9 | 36.5 | 35.4 | 18.2 | 14.5 | 17.7 | 17.4 | ||||||||||||||||
| Finnish | 30.1 | 26.0 | 32.1 | 32.3 | 18.1 | 15.3 | 18.6 | 17.7 | ||||||||||||||||
| Polish | 25.2 | 22.5 | 24.9 | 23.4 | 17.5 | 15.0 | 15.6 | 15.5 | ||||||||||||||||
| Icelandic | 28.5 | 27.2 | 29.6 | 28.2 | 16.2 | 13.5 | 16.0 | 15.6 | ||||||||||||||||
| Czech | 35.9 | 32.5 | 33.3 | 33.7 | 20.4 | 18.1 | 19.4 | 19.0 | ||||||||||||||||
| Filipino | 35.0 | 32.0 | 32.1 | 33.1 | 22.3 | 19.0 | 20.7 | 20.7 | ||||||||||||||||
| Persian | 30.7 | 27.3 | 25.1 | 25.9 | 19.5 | 16.4 | 15.9 | 15.9 | ||||||||||||||||
| Greek | 30.0 | 27.8 | 30.0 | 29.4 | 18.4 | 15.9 | 17.6 | 17.5 | ||||||||||||||||
| Asturian | 32.4 | 27.9 | 31.5 | 30.4 | 20.4 | 16.2 | 18.6 | 18.1 | ||||||||||||||||
| Belarusian | 16.4 | 14.7 | 16.4 | 16.5 | 12.6 | 10.8 | 12.1 | 12.2 | ||||||||||||||||
| Bulgarian | 45.0 | 40.7 | 41.7 | 42.5 | 25.6 | 23.0 | 24.3 | 24.0 | ||||||||||||||||
| Bengali | 18.6 | 15.7 | 14.3 | 15.0 | 10.6 | 9.2 | 9.0 | 9.3 | ||||||||||||||||
| Bosnian | 37.5 | 34.0 | 36.3 | 35.2 | 21.4 | 18.6 | 19.9 | 19.4 | ||||||||||||||||
| Catalan | 43.9 | 41.5 | 42.7 | 42.9 | 26.6 | 17.2 | 25.0 | 25.1 | ||||||||||||||||
| Cebuano | 28.5 | 12.7 | 28.5 | 29.2 | 19.0 | 5.6 | 17.5 | 17.7 | ||||||||||||||||
| Estonian | 30.8 | 26.3 | 31.4 | 30.6 | 18.9 | 13.6 | 17.5 | 17.2 | ||||||||||||||||
| Galician | 37.4 | 35.4 | 36.6 | 35.9 | 23.9 | 22.0 | 22.7 | 22.3 | ||||||||||||||||
| Gujarati | 23.8 | 20.9 | 21.0 | 21.5 | 14.3 | 10.8 | 12.3 | 12.3 | ||||||||||||||||
| Croatian | 33.3 | 30.7 | 33.4 | 32.6 | 21.3 | 18.5 | 19.2 | 18.6 | ||||||||||||||||
| Hungarian | 29.5 | 24.9 | 28.1 | 27.6 | 18.8 | 15.8 | 17.3 | 16.7 | ||||||||||||||||
| Javanese | 26.8 | 24.4 | 16.4 | 22.1 | 16.5 | 14.8 | 9.0 | 13.0 | ||||||||||||||||
| Kazakh | 24.9 | 21.1 | 20.6 | 22.5 | 15.0 | 12.4 | 12.9 | 13.1 | ||||||||||||||||
| Kannada | 20.0 | 17.0 | 16.1 | 17.2 | 11.7 | 6.9 | 9.5 | 10.0 | ||||||||||||||||
| Kyrgyz | 15.3 | 12.6 | 15.1 | 14.7 | 10.4 | 7.8 | 10.0 | 9.5 | ||||||||||||||||
| Latvian | 36.1 | 31.0 | 35.4 | 35.3 | 21.9 | 17.8 | 20.1 | 19.5 | ||||||||||||||||
| Macedonian | 38.1 | 34.0 | 38.5 | 38.2 | 22.3 | 20.1 | 22.0 | 21.6 | ||||||||||||||||
| Malayalam | 19.3 | 11.1 | 16.1 | 16.1 | 10.4 | 5.2 | 9.9 | 9.4 | ||||||||||||||||
| Marathi | 17.7 | 11.6 | 15.9 | 16.2 | 11.3 | 8.1 | 9.4 | 10.1 | ||||||||||||||||
| Punjabi | 26.1 | 23.1 | 24.1 | 24.6 | 15.7 | 8.9 | 14.3 | 14.1 | ||||||||||||||||
| Romanian | 42.0 | 39.9 | 41.4 | 42.0 | 25.6 | 22.8 | 23.6 | 23.4 | ||||||||||||||||
| Slovak | 35.3 | 31.4 | 34.8 | 34.6 | 19.6 | 16.4 | 19.2 | 18.5 | ||||||||||||||||
| Slovenian | 32.8 | 28.5 | 33.5 | 32.9 | 20.1 | 17.7 | 20.9 | 20.2 | ||||||||||||||||
| Swahili | 36.3 | 30.9 | 32.1 | 32.2 | 20.4 | 9.3 | 18.5 | 18.3 | ||||||||||||||||
| Tajik | 23.8 | 18.3 | 22.1 | 22.9 | 14.6 | 10.9 | 14.3 | 14.2 | ||||||||||||||||
| Azerbaijani | 13.7 | 9.8 | 15.2 | 14.7 | 11.5 | 9.7 | 11.3 | 10.8 | ||||||||||||||||
| Ukrainian | 31.7 | 29.2 | 29.8 | 30.1 | 19.5 | 15.0 | 17.0 | 17.6 | ||||||||||||||||
| Average | 33.8 | 30.4 | 31.8 | 31.8 | 21.4 | 18.1 | 19.6 | 19.5 | ||||||||||||||||
| xx2en (Other Languages → English) | xx2zh (Other Languages → Chinese) | |||||||||||||||||||||||
| Language |
|
|
|
|
|
|
|
| ||||||||||||||||
| Chinese | 32.2 | 31.2 | 30.1 | 29.5 | – | – | – | – | ||||||||||||||||
| English | – | – | – | – | 47.8 | 46.6 | 47.4 | 46.3 | ||||||||||||||||
| Cantonese | 30.3 | 29.9 | 27.7 | 26.3 | 36.8 | 37.5 | 21.2 | 22.4 | ||||||||||||||||
| Arabic | 42.9 | 40.0 | 42.1 | 42.3 | 40.2 | 37.3 | 40.5 | 40.5 | ||||||||||||||||
| German | 44.6 | 44.1 | 43.9 | 43.9 | 43.3 | 42.8 | 42.1 | 41.8 | ||||||||||||||||
| French | 43.5 | 42.0 | 41.6 | 41.3 | 41.6 | 40.3 | 41.6 | 41.4 | ||||||||||||||||
| Spanish | 32.3 | 31.3 | 30.3 | 30.4 | 38.8 | 38.5 | 38.4 | 38.2 | ||||||||||||||||
| Portuguese | 49.4 | 48.2 | 47.7 | 47.5 | 43.6 | 41.8 | 42.5 | 42.4 | ||||||||||||||||
| Indonesian | 45.7 | 43.1 | 45.1 | 44.9 | 43.5 | 41.4 | 42.8 | 42.9 | ||||||||||||||||
| Italian | 34.4 | 31.9 | 30.8 | 31.0 | 40.8 | 39.5 | 39.7 | 39.4 | ||||||||||||||||
| Korean | 34.1 | 32.4 | 32.0 | 32.1 | 39.9 | 37.5 | 37.0 | 37.2 | ||||||||||||||||
| Russian | 38.6 | 37.2 | 36.7 | 36.5 | 41.7 | 39.5 | 41.1 | 40.4 | ||||||||||||||||
| Thai | 34.1 | 32.4 | 34.2 | 33.0 | 40.2 | 37.9 | 40.0 | 39.5 | ||||||||||||||||
| Vietnamese | 36.4 | 34.9 | 36.1 | 35.4 | 38.7 | 36.3 | 39.2 | 38.9 | ||||||||||||||||
| Japanese | 30.4 | 29.2 | 29.5 | 29.4 | 38.0 | 35.7 | 35.6 | 34.8 | ||||||||||||||||
| Turkish | 40.3 | 39.1 | 39.5 | 38.9 | 41.8 | 40.3 | 40.6 | 41.0 | ||||||||||||||||
| Hindi | 38.8 | 36.2 | 39.3 | 39.2 | 38.9 | 36.9 | 38.0 | 38.4 | ||||||||||||||||
| Malay | 42.9 | 41.1 | 44.8 | 42.4 | 41.0 | 39.5 | 42.4 | 41.5 | ||||||||||||||||
| Dutch | 33.3 | 32.2 | 31.2 | 30.4 | 40.1 | 38.5 | 39.4 | 39.4 | ||||||||||||||||
| Urdu | 35.5 | 31.5 | 32.9 | 32.6 | 37.2 | 33.8 | 36.9 | 36.5 | ||||||||||||||||
| Norwegian | 43.5 | 42.1 | 42.6 | 41.4 | 42.2 | 39.6 | 40.9 | 40.8 | ||||||||||||||||
| Swedish | 47.2 | 45.2 | 46.6 | 44.7 | 42.9 | 40.7 | 41.5 | 41.6 | ||||||||||||||||
| Danish | 45.4 | 44.0 | 44.7 | 42.7 | 43.4 | 41.1 | 41.7 | 40.6 | ||||||||||||||||
| Hebrew | 39.7 | 36.4 | 42.5 | 39.9 | 36.7 | 34.1 | 40.3 | 37.8 | ||||||||||||||||
| Finnish | 36.9 | 35.0 | 36.7 | 35.5 | 40.8 | 38.7 | 41.0 | 40.6 | ||||||||||||||||
| Polish | 32.1 | 30.5 | 30.4 | 29.7 | 38.4 | 36.0 | 37.9 | 37.4 | ||||||||||||||||
| Icelandic | 31.5 | 27.5 | 35.1 | 35.8 | 38.2 | 31.8 | 37.9 | 37.5 | ||||||||||||||||
| Czech | 42.1 | 39.3 | 40.1 | 39.5 | 40.6 | 39.4 | 40.6 | 40.0 | ||||||||||||||||
| Filipino | 42.7 | 40.9 | 45.0 | 44.3 | 41.0 | 38.0 | 42.6 | 41.5 | ||||||||||||||||
| Persian | 40.2 | 36.8 | 38.0 | 38.1 | 40.2 | 37.0 | 41.1 | 41.1 | ||||||||||||||||
| Greek | 36.0 | 32.5 | 35.3 | 35.4 | 38.0 | 32.8 | 39.0 | 38.4 | ||||||||||||||||
| Asturian | 37.0 | 35.1 | 37.2 | 35.5 | 37.7 | 34.0 | 38.1 | 36.7 | ||||||||||||||||
| Belarusian | 23.1 | 19.9 | 20.6 | 19.9 | 33.3 | 31.2 | 33.7 | 33.7 | ||||||||||||||||
| Bulgarian | 39.6 | 36.0 | 41.3 | 40.0 | 40.9 | 36.6 | 41.0 | 40.4 | ||||||||||||||||
| Bengali | 32.0 | 27.6 | 34.9 | 34.3 | 35.8 | 32.4 | 38.1 | 37.5 | ||||||||||||||||
| Bosnian | 43.1 | 40.8 | 42.7 | 42.4 | 41.5 | 39.1 | 42.3 | 41.4 | ||||||||||||||||
| Catalan | 46.6 | 42.3 | 46.2 | 45.5 | 42.2 | 38.9 | 42.6 | 41.5 | ||||||||||||||||
| Cebuano | 37.3 | 26.3 | 38.9 | 38.2 | 34.0 | 26.5 | 36.7 | 36.6 | ||||||||||||||||
| Estonian | 35.7 | 28.3 | 40.1 | 38.3 | 38.0 | 32.2 | 41.7 | 41.2 | ||||||||||||||||
| Galician | 40.8 | 38.6 | 39.6 | 38.6 | 40.9 | 39.1 | 41.6 | 40.9 | ||||||||||||||||
| Gujarati | 33.4 | 28.3 | 40.3 | 39.7 | 35.8 | 31.4 | 39.7 | 39.0 | ||||||||||||||||
| Croatian | 39.5 | 36.3 | 38.0 | 36.9 | 40.0 | 37.7 | 40.2 | 39.6 | ||||||||||||||||
| Hungarian | 35.5 | 29.6 | 35.3 | 33.3 | 39.1 | 33.9 | 40.0 | 37.4 | ||||||||||||||||
| Javanese | 35.9 | 28.4 | 38.4 | 36.5 | 34.9 | 28.3 | 36.7 | 36.3 | ||||||||||||||||
| Kazakh | 34.4 | 26.9 | 35.7 | 35.5 | 37.1 | 31.4 | 39.3 | 38.8 | ||||||||||||||||
| Kannada | 26.4 | 19.8 | 33.1 | 33.6 | 32.3 | 26.1 | 38.0 | 37.8 | ||||||||||||||||
| Kyrgyz | 22.2 | 17.1 | 24.8 | 23.7 | 29.9 | 24.8 | 33.7 | 32.8 | ||||||||||||||||
| Latvian | 33.7 | 25.2 | 38.0 | 37.9 | 37.1 | 30.0 | 41.4 | 40.6 | ||||||||||||||||
| Macedonian | 43.3 | 39.6 | 43.1 | 41.9 | 41.6 | 38.0 | 42.1 | 41.8 | ||||||||||||||||
| Malayalam | 31.2 | 25.7 | 34.9 | 34.2 | 36.1 | 31.5 | 38.4 | 38.0 | ||||||||||||||||
| Marathi | 33.5 | 25.7 | 36.7 | 35.5 | 34.6 | 29.4 | 38.8 | 37.9 | ||||||||||||||||
| Punjabi | 33.0 | 26.9 | 38.3 | 36.5 | 35.1 | 29.9 | 37.4 | 36.4 | ||||||||||||||||
| Romanian | 43.5 | 39.7 | 42.3 | 41.5 | 42.0 | 38.7 | 42.4 | 41.3 | ||||||||||||||||
| Slovak | 39.7 | 38.4 | 39.9 | 38.8 | 39.2 | 38.1 | 40.2 | 39.3 | ||||||||||||||||
| Slovenian | 31.7 | 26.5 | 34.7 | 33.2 | 34.5 | 30.1 | 38.4 | 37.3 | ||||||||||||||||
| Swahili | 35.0 | 27.4 | 42.6 | 40.5 | 33.9 | 26.8 | 39.2 | 37.9 | ||||||||||||||||
| Tajik | 33.9 | 29.0 | 34.5 | 33.3 | 36.7 | 32.7 | 38.9 | 38.1 | ||||||||||||||||
| Azerbaijani | 25.0 | 22.0 | 23.7 | 23.3 | 33.4 | 30.5 | 33.5 | 33.3 | ||||||||||||||||
| Ukrainian | 42.0 | 40.1 | 41.7 | 41.4 | 41.7 | 39.7 | 41.6 | 40.7 | ||||||||||||||||
| Average | 37.0 | 33.5 | 37.4 | 36.6 | 38.9 | 35.7 | 39.4 | 38.9 | ||||||||||||||||