Shobhit Banga Shobhitbanga 关注
VoiceArena
Manas Dhir manasdhir04 关注
VoiceArena
Bhaskar Singh bhaskarJT 关注
VoiceArena
Manmeet Kaur manmeet-voicearena 关注
VoiceArena
Aaditya Pareek pareek-voicearena 关注
VoiceArena
Walecha Amritansh8675 关注
VoiceArena
Sagar Jain sagarjain268380 关注
VoiceArena
Hanuman Sidh hanuman44420 关注
VoiceArena
Vanshika Chhabra vanshikachhabra-voicearena 关注
VoiceArena
Voice Arena 与 Hugging Face 合作,为印地语和印度英语推出开放 ASR 评测。基准决定什么会被构建。一个在 Open ASR 排行榜上得分高的模型会被采用并持续迭代,而排行榜未衡量的能力往往得不到改进。近期排行榜上的大量工作都集中在让评测指标更加可信:
- 留出的私有数据划分。
- 基准拟合分析,用于量化模型在多大程度上是在复现参考转写文本,而非仅基于音频进行转写。
- 缩小归一化器中的差距,确保正确的预测/变体不会被误判扣分。
所有这一切都让一个数字(WER)更难被钻空子。但它仍然只是一个数字。一系列长期研究表明,ASR 错误率在使用者之间并非均匀分布。关于自动语音识别中的种族差异的研究发现,商业系统对黑人说话者的错误率大约是白人的两倍,而《量化自动语音识别中的偏差》进一步发现了性别、年龄和口音方面的差异。这些在排行榜上都不可见,而这并非因为排行榜在隐瞒什么——它运行的测试集记录的是说了什么,而几乎不记录是谁说的。
为了解决这一缺口,我们向 Open ASR 排行榜引入了两个评测集:Monsoon en-IN 和 Monsoon hi-IN。印地语有超过五亿使用者,是当前仅覆盖欧洲语言的多语言标签页上第一个印度语言。每个评测集都发布一个公开划分,可供自行评分,以及一个私有划分,用于限制针对特定基准的优化。这四个划分在说话者上互不重叠,共包含 4,888 位说话者,并为每位说话者记录了 12 项说话者属性。
数据收集的设计
一个测试集只能暴露它所变化的维度上的失败模式。大多数基准测试都是用任何现成的音频构建的。Monsoon 则是围绕九个维度构建的:地理、年龄、性别、词汇、设备、声学环境、语音类型、语速,以及同一音频存在多个有效转写文本的情况。每一种维度都是一种方式,使得聚合 WER 在平均值上正确,但对特定人群却是错误的。
采集方法正是由此而来。
地理多样性来自于在数百个地区招募贡献者,而不是在少数几个地方进行更长时间的录音。设备和声学条件来自于贡献者使用自己的手机和网络,在室内和室外录制,而不是在安静的房间里使用提供的硬件。词汇、语音类型和语速来自于提示词:日常话题促使贡献者表达观点、分歧、叙述和回忆,而这些正是命名实体、数字和未经排练的措辞出现的地方。年龄和性别按说话人记录并经过验证。多个有效转写文本是参考标注的属性,而非音频本身的属性,这将在后面的章节中讨论。
数据集构成
四个数据划分,两种语言,通过同一条流水线采集。
| 数据集 | 语言 | 时长 | 说话人数 | 片段长度(平均 / 中位数) | 男/女 | 地区数 | 邦/联邦属地 | 设备数 | 风格 | 转写 |
|---|---|---|---|---|---|---|---|---|---|---|
| Monsoon en-IN 公开 | 印度英语 | 5.62 小时 | 1,444 | 9.6 秒 / 10.4 秒 | 50/50 | 428 | 24/6 | 556 | 对话式、自发性 | 规范化,含不流畅标记 |
| Monsoon en-IN 私有 | 印度英语 | 5.58 小时 | 1,405 | 9.6 秒 / 10.4 秒 | 45/55 | 420 | 24/6 | 560 | 对话式、自发性 | 规范化,含不流畅标记 |
| Monsoon hi-IN 公开 | 印地语 | 1.33 小时 | 468 | 6.4 秒 / 5.0 秒 | 54/46 | 202 | 11/3 | 315 | 对话式、自发性 | 格网(接受的拼写变体) |
| Monsoon hi-IN 私有 | 印地语 | 4.47 小时 | 1,571 | 6.6 秒 / 5.3 秒 | 55/45 | 295 | 12/3 | 582 | 对话式、自发性 | 格网(接受的拼写变体) |
数据来源于无脚本的双通道自发性对话,片段从单个通道切分,因此每个片段只包含一位说话人。除了表中报告的字段外,每个片段还记录了职业、教育程度、婚姻状况、收入区间、手机品牌、当前所在城市以及在当前地区的居住年数。
以下是从公开的印度英语数据集中抽取的五段音频片段,以及每段所附带的元数据:
29岁女性,来自特里普拉邦西特里普拉县。学生,使用三星SM-G781B手机。
32岁女性,来自中央邦萨特纳。无业,使用三星SM-E146B手机。
22岁男性,来自比哈尔邦罗塔斯。学生,使用摩托罗拉moto g54 5G手机。
27岁女性,来自特伦甘纳邦瓦朗加尔。无业,使用vivo V2247手机。
57岁男性,来自本地治里。私企职员,使用小米M2006C3LI手机。
英语数据集使用标准的字符串引用,排行榜的归一化器可以消除大部分拼写差异。而印地语的拼写差异要大得多,且没有任何归一化器能解决这个问题,因为这些变体并非两种书写规范之间的固定映射关系。因此,印地语数据集采用了一种格状结构:对于转录文本的每一段,都附上一个被认定为正确拼写的列表。
说话人覆盖情况
Monsoon数据集按小时衡量规模很小,但按说话人数量衡量规模却很大。这是设计使然,也是其大部分价值所在。
除上述字段外,说话人的集中度与多样性情况。
| Monsoon 印地语(印度)公开集 | Monsoon 印地语(印度)私有集 | Monsoon 英语(印度)公开集 | Monsoon 英语(印度)私有集 | |
|---|---|---|---|---|
| 每位说话人的平均片段数 | 1.61 | 1.56 | 1.46 | 1.48 |
| 仅有一个片段的说话人数 | 261 | 994 | 956 | 924 |
| 每位说话人的音频时长(中位数) | 8.34秒 | 8.28秒 | 12.36秒 | 12.39秒 |
| 前10位说话人所占时长份额 | 6.8% | 3.1% | 2.8% | 2.9% |
| 覆盖城市数量 | 289 | 814 | 641 | 584 |
| 设备制造商数量 | 18 | 25 | 23 | 20 |
接下来是三个特性,每一个都是关于多样性而非体量的论断。
没有单一声音主导成绩:贡献最大的十位说话人合计仅占总时长的2.8%至6.8%,且超过一半的说话人只出现一次。Monsoon上的评测结果是数百种不同声音的平均表现,而非少数几位说话人的长时间录音。同等时长的测试集通常采用另一种构建方式。
也没有单一地区或手机型号主导成绩:印度英语公开集覆盖了30个邦和中央直辖区的428个本土地区;印地语数据集作为印地语地带语言,分布更为集中,但仍覆盖了202至295个地区。录音来自315至582种不同的设备型号,且在任何子集中,单一型号所占片段比例均未超过2.1%。在标准化硬件上采集的语料库往往会过拟合于某一种麦克风响应;而本数据集则不会。
这里的印度英语并非单一口音:这是英语在整个印度各地的实际使用形态,而非某一地区的英语。印度六大区域均有代表:在公开数据集中,35% 的语音片段来自南部 speakers,18% 来自东部,18% 来自中部,16% 来自北部,11% 来自西部。由此产生的口音差异被记录在元数据中,而非仅作笼统宣称。
元数据字段
Monsoon 为每个语音片段提供 18 列数据,其中 12 列为元数据,而大多数公开 ASR 测试集仅提供标识符、转写文本和时长。人口统计字段完整或近乎完整;贡献者已同意将这些数据用于此用途。
| 分组 | 字段 |
|---|---|
| 语音片段 | id、audio、audiolengths、language |
| 参考文本 | lattice(印地语)或 text(印度英语) |
| 说话人 | speakerid、gender、dateofbirth |
| 背景 | occupation、educationalbackground、maritalstatus、income |
| 地理 | nativedistrict、nativestate、currentcity、yearsspentincurrentdistrict |
| 录音设备 | devicemanufacturer、devicemodel |
这两种语言具有不同的地理分布形态,而这种形态本身就蕴含信息。印地语数据集集中在印地语地带,北方邦约占说话人总数的 40%,这正是一个按人口比例采样的印地语语料库应有的样子。印度英语数据集则平坦得多:没有任何一个邦占比超过 13%,且有三分之一的说话人来自人口最多的八个邦之外。公开和私有两半数据在这两方面都高度吻合。
印度各邦的边界是按语言划分的,因此地区和邦别携带了真实的口音信号,这正是这些字段被公开而非被概括掉的原因。此类分析已在印度 ASR 领域以大规模方式得到报告:地区级错误率跨度约为 4% 至 44%,代表性不足的地区远落后于印地语地带和大都市,同时还按音频质量、语速、话语时长、性别、年龄和设备进行了细分。这些实验是在封闭基准上进行的。Monsoon 使得同类分析可以在公开排行榜测试集上开展。
数据采集与质量控制
要实现广泛的地理覆盖,需要在数百个地区进行招募,而非依赖少数发言人的更长录音;而如此规模的分布式招募会引入小规模采集所不会面临的故障模式:贡献者刷任务、将回放音频冒充实时语音提交,以及标注不认真。每一项都通过明确的检查机制加以应对。
招募与录音:贡献者通过 Voice Arena 社区招募,这是一个全球性数字平台,其触达范围延伸至语音语料库很少覆盖的农村和半城市地区。随后,成对贡献者通过点对点界面录制双人对话,采用双声道,围绕指定的日常话题展开。贡献者使用自己的手机和自己的网络连接。其中许多是低端设备、带宽不稳定,这正是已发布音频中保留了这种条件而非将其过滤掉的原因。有意参与的贡献者在获得录音权限前需完成语言能力筛查,获得报酬,并在知情同意的前提下同意其语音用于训练和分发。每位发言人的时长上限按语言分别校准,依据该语言的使用人口规模和地理分布而定,防止少数高产贡献者主导某一语言或地区;这些数据集中超过一半的发言人恰好贡献一个片段。
引导采集:大规模引导自发语音本身就有难度,因为贡献者在没有结构化引导的情况下往往只会给出简短、稀疏的回应。因此,每段对话都以一个开放式叙事提示作为起点,并逐步揭示后续追问,话题涵盖旅行、医疗、农业、教育和数字服务等领域,引导交流走向更长的描述,而非照本宣科。候选话题由大语言模型生成,随后由母语语言学家审阅并进行本地化。
质量控制:每条录音在转写前都需通过一组门控检查。口语语种通过基于人工标注数据、覆盖 30 多种语言训练的语言识别模型,与指定语种进行核验。说话人性别通过专用分类器与自我报告标签进行确认,该分类器仅作为自我报告的佐证,而非替代。另有模型用于区分真实自发对话与预录或回放音频。信噪比估算可剔除清晰度受损的录音,而自然环境背景噪声则有意保留,以维持真实场景语音的声学真实感。通过上述检查的录音经语音活动检测进行切分,在连续静音两秒处断开,或达到十五秒软上限后于下一个检测到的静音处闭合。切分按声道独立进行,因此每个片段均为单说话人、单声道。切分后的片段随后通过 DNSMOS P.808 检查。
参考转写文本由人工完成。初稿由基于领域内数据训练的内部ASR模型生成,这些模型均未出现在任何公开排行榜上,因此在这些评测集上接受评估的系统,其得分所参照的参考文本并非由它们自身参与生成。后续每个阶段均由母语语言学家在基于严格分工的五级协议下完成,每一轮修正之后都紧跟一轮由不同标注员执行的独立复核,确保没有任何语言学家审核自己的产出。语言学家首先对照声学信号逐段修正初稿;第二位复核人再次核验并标记遗留分歧。后续级别由新的标注员重复这一循环,逐步解决模糊的音位实现、语码转换边界、命名实体以及跨拼写变体的正字法一致性问题。数字以单词形式写出,使转写文本与口语内容直接对应。在最终级别仍被标记的片段,在收录前会被退回重新转写。标注员行为全程受到自动监控,凡是提交内容包含目标文字系统之外的字符、异常的字符或词语重复,以及编辑次数异常偏低或偏高的,都会被标记出来。
区域差异
以下是一个示例,运行在公开的印度英语子集上,用以展示这些元数据所能支撑的评估类型。这并不是这些数据集旨在得出的结论,而是说明一旦每个片段都带有说话人信息,哪些问题就变得可以回答了。
排行榜上有八个模型在该子集上的词错误率(WER)介于4.81到4.99之间。从最好到最差仅相差0.18个百分点,这在五小时数据所能分辨的范围之内。按该语料库排名,它们实际上是同一个模型。
按地区对说话人进行分组,则呈现出另一番景象。每位说话人的母语地区被汇总到其所属的区议会(印度内政部对各邦的划分),由此得到五个采样充分的区域。openai/whisper-large-v3-turbo 在这些区域间的得分差异为 0.46 个百分点。mistralai/Voxtral-Mini-3B-2507 在语料库上的整体得分落后前者 0.14 个百分点,但其区域间差异达到 1.68——在中央区得分为 4.38,而在东部地区则为 6.06。两个在排行榜上难分伯仲的系统,其准确率对说话人地域的依赖程度却相差近四倍。
哪个区域最难识别也并非固定不变。ibm-granite/granite-speech-3.3-2b 在北部表现最差,microsoft/VibeVoice-ASR-HF 在南部最差,而 mistralai/Voxtral-Mini-3B-2507 则在东部最差。如果仅仅是因为某个区域的音频本身更难转写,那么所有模型对各区域的排名顺序应当一致。事实并非如此,这说明差异源于模型本身,而非音频。
地区只是十二项记录属性之一,上述区域划分则是 428 个地区的粗略汇总。同样的细分分析也适用于年龄、教育程度、职业和手机型号,发布的数据文件中包含复现这一切所需的全部信息。而对于仅记录说话内容的测试集,这些信息一概不可得。
印地语的正字法差异
英语的正字法差异是有限的。英式拼写与美式拼写、标点、大小写、数字与单词之间的差异:规范化器大多可以将其映射为单一形式,排行榜所采用的规范化器正是如此。印地语则不具备这种有限性。日常口语大量夹杂语码混用,源自英语的词汇没有固定的天城文拼写,复合词形式根据个人偏好选择连写或分写。同一个短语可以有十种甚至更多种有效的书写形式,且不存在固定的映射关系能将它们归并,因为并没有一个所谓的规范侧可供映射。
使用单一参考答案进行评分时,WER 会奖励系统输出恰好与标注者所选拼写一致的结果。两个对音频识别效果同样出色的系统,仅因正字法差异就可能相差数个百分点。
因此,印地语数据集采用了一种格状结构:对于转录文本的每个片段,都提供一组被接受为正确的书写形式。构建这一结构需要人工操作。候选变体来自同一音频的多个 ASR 转录结果,并通过语言模型进行扩展,然后由母语语言学家判断哪些变体对该话语有效,并剔除其余部分,从而只保留与所说内容一致的书写形式。
因此,对于印地语,我们报告的是由 AI4Bharat 提出的“正字法感知词错误率”(OIWER),而非 WER。假设结果在每个片段处与接受集合进行对齐,因此任何被接受的书写形式都计为正确,只有真正的识别错误才会被计入。
为了量化这一影响,我们对相同的假设结果进行了两次评分。将每个格状结构展平为每个片段的第一个变体,就得到了传统基准测试所提供的那种单一字符串参考;任何被接受的变体都能同样有效地作为参考,而不同的选择会产生不同的参考。与展平后的参考进行评分时,每个系统的错误率都会上升,而且上升幅度并不一致。因此,排名也随之发生变化。下图显示了两对系统在两种参考下顺序发生反转的情况:在单一参考下,系统在一定程度上会因为复现了标注者的正字法而获得加分,而格状结构评分则只衡量识别本身。
我们还开源了我们的实现 voi-oiwer,以便这些数据集上的所有结果都可以直接复现。
接受评估
对于私有测试集,请将您的模型提交到 Open ASR Leaderboard,Hugging Face 团队将负责运行评估。与之前一样,将模型添加到排行榜的流程在 Open ASR Leaderboard 的 GitHub 上进行:
- 提交一个拉取请求,会出现一个模型检查清单。与之前一样,您需要在公共数据集上报告您的结果。
- 我们将在公共数据集上验证结果,并在私有数据集上计算指标。
- 确认我们获得的结果。
印地英语(Indian-English)现已加入主排行榜,成为 Voice Arena Monsoon 的默认列之一,而非可选的开关项,因此它会为每个模型的整体平均词错误率(Average WER)贡献数据。私有数据划分与 Appen 和 DataoceanAI 的数据一起,汇入聚合的“私有(对话式)”列。公开和私有的印地语数据则出现在“多语言”标签页中,在该标签页下,模型只有在支持所有选定语言时才会被排名,这使得该列成为同类对比。或者,也可以从“语言数据集细分”下拉菜单中选择“印地语”。
接下来是什么
印地语是一个普遍问题的一个典型例子。任何以多种方式书写、且其使用者未被基准测试所采样的语言,都会同时带有这里所描述的两种缺陷。这些数据集并不能修复这些问题。它们所增加的是提供一种看见问题的方式:一个位于该领域已经关注的排行榜上的测试集,它包含了关于每位说话者和每个参考文本的足够信息,使得两个系统之间的差异可以追溯到是谁在说话以及他们如何书写,而不是消失在单一数字中。
这四个数据集是 Monsoon 的一部分,Monsoon 是 Voice Arena 为全球南方(Global South)发起的更广泛的数据集计划。
本文提到的模型 4
本文提到的 Spaces 1
本文提到的论文 2
音频语音基准 ## 衡量语音识别中的基准优化 +3 57 2026年8月21日 tlebryk02 等人
本文提到的模型 4
本文提到的 Spaces 1
本文提到的论文 2
Shobhit Banga Shobhitbanga Follow
VoiceArena
Manas Dhir manasdhir04 Follow
VoiceArena
Bhaskar Singh bhaskarJT Follow
VoiceArena
Manmeet Kaur manmeet-voicearena Follow
VoiceArena
Aaditya Pareek pareek-voicearena Follow
VoiceArena
Walecha Amritansh8675 Follow
VoiceArena
Sagar Jain sagarjain268380 Follow
VoiceArena
Hanuman Sidh hanuman44420 Follow
VoiceArena
Vanshika Chhabra vanshikachhabra-voicearena Follow
VoiceArena
Voice Arena and Hugging Face partner to launch open ASR evaluation for Hindi and Indian English Benchmarks decide what gets built. A model that scores well on the Open ASR Leaderboard gets adopted and iterated on, while capabilities the leaderboard does not measure tend not to improve. Much of the recent work on the leaderboard has gone into making the evaluation metrics more trustworthy:
- Held-out private splits.
- Benchmark-fitting analysis to quantify how much models are reproducing reference transcripts rather than transcribing solely on the audio.
- Closing the gaps in normalisers to ensure correct predictions/variants are not penalized.
All of that makes one number (WER) harder to game. It is still one number. A long line of work has shown that ASR error rates are not evenly distributed across the people using them. Racial disparities in automated speech recognition found commercial systems roughly twice as bad for Black speakers as for white speakers, and Quantifying Bias in Automatic Speech Recognition found further differences by gender, age and accent. None of that is visible on a leaderboard, and not because the leaderboard is hiding it. The test sets it runs on record what was said and almost nothing about who said it.
To address this gap, we introduce two evaluation sets to the Open ASR Leaderboard: Monsoon en-IN and Monsoon hi-IN. Hindi, spoken by more than half a billion people, is the first Indic language on a multilingual tab that currently covers only European languages. Each set is released as a public split, available for self-scoring, and a private split withheld to limit benchmark-specific optimisation. The four splits are speaker-disjoint, comprising 4,888 speakers, with 12 speaker attributes recorded for each.
Design of the collection
A test set can only expose a failure mode it varies along. Most benchmarks are built from whatever audio was readily available. Monsoon was built to vary along nine axes: geography, age, gender, vocabulary, devices, acoustic environments, speech type, speech rate, and the existence of multiple valid transcripts for the same audio. Each is a way an aggregate WER can be right on average and wrong for a particular population.
The collection method follows from that.
Geography comes from recruiting across hundreds of districts rather than recording longer sessions in fewer places. Devices and acoustic conditions come from contributors using their own handsets and connections, indoors and out, rather than supplied hardware in a quiet room. Vocabulary, speech type and speech rate come from the prompts: everyday topics that push contributors toward opinion, disagreement, narration and recall, which is where named entities, numbers and unrehearsed phrasing appear. Age and gender are recorded per speaker and verified. Multiple valid transcripts is a property of the reference rather than the audio, and it is the subject of a later section.
Dataset composition
Four splits, two languages, collected through one pipeline.
| Set | Language | Duration | Speakers | Clip length (mean / median) | M/F | Districts | States/UTs | Devices | Style | Transcription |
|---|---|---|---|---|---|---|---|---|---|---|
| Monsoon en-IN public | Indian English | 5.62 h | 1,444 | 9.6s / 10.4s | 50/50 | 428 | 24/6 | 556 | Conversational, spontaneous | Normalised, disfluencies |
| Monsoon en-IN private | Indian English | 5.58 h | 1,405 | 9.6s / 10.4s | 45/55 | 420 | 24/6 | 560 | Conversational, spontaneous | Normalised, disfluencies |
| Monsoon hi-IN public | Hindi | 1.33 h | 468 | 6.4s / 5.0s | 54/46 | 202 | 11/3 | 315 | Conversational, spontaneous | Lattice (accepted orthographic variants) |
| Monsoon hi-IN private | Hindi | 4.47 h | 1,571 | 6.6s / 5.3s | 55/45 | 295 | 12/3 | 582 | Conversational, spontaneous | Lattice (accepted orthographic variants) |
The data is sourced from unscripted dual-channel spontaneous conversations, with clips segmented from a single channel so that each clip carries one speaker. Along with the fields reported in the table, each clip also records occupation, education, marital status, income band, handset brand, current city and years in the current district.
Five clips from the public Indian English split, with the metadata each one carries:
29-year-old woman, West Tripura, Tripura. Student, samsung SM-G781B.
32-year-old woman, Satna, Madhya Pradesh. Unemployed, samsung SM-E146B.
22-year-old man, Rohtas, Bihar. Student, motorola moto g54 5G.
27-year-old woman, Warangal, Telangana. Unemployed, vivo V2247.
57-year-old man, Puducherry. Private job, Xiaomi M2006C3LI.
The English sets use standard string references, where the leaderboard's normaliser collapses most spelling variation. Hindi has far more of it, and no normaliser can resolve it, because the variants are not a fixed mapping between two conventions. The Hindi sets therefore ship a lattice: for each span of the transcript, a list of the spellings that are accepted as correct.
Speaker coverage
Monsoon is small measured in hours and large measured in speakers. That is the design, and it is where most of the value sits.
Speaker concentration and diversity beyond the fields above.
| Monsoon hi-IN public | Monsoon hi-IN private | Monsoon en-IN public | Monsoon en-IN private | |
|---|---|---|---|---|
| Segments per speaker (mean) | 1.61 | 1.56 | 1.46 | 1.48 |
| Speakers with a single segment | 261 | 994 | 956 | 924 |
| Audio per speaker (median) | 8.34 s | 8.28 s | 12.36 s | 12.39 s |
| Share held by top 10 speakers | 6.8% | 3.1% | 2.8% | 2.9% |
| Current cities | 289 | 814 | 641 | 584 |
| Device manufacturers | 18 | 25 | 23 | 20 |
Three properties follow, and each is a claim about variance rather than volume.
No voice carries the score: The ten largest contributors account for between 2.8% and 6.8% of total duration, and more than half of all speakers appear exactly once. A result on Monsoon is an average over hundreds of distinct voices, not a small number of talkers recorded at length. Test sets of comparable duration are usually constructed the other way.
No region or handset carries it either: The Indian English public set draws on 428 native districts across 30 states and union territories; the Hindi sets, being a Hindi-belt language, concentrate more tightly but still span 202 and 295 districts. Recordings come from 315 to 582 distinct device models, with no single model exceeding 2.1% of segments in any subset. Corpora collected on standardised hardware overfit to one microphone response; this one cannot.
Indian English here is not one accent: This is English as it is spoken across the country, not the English of one region. All six zones are represented: in the public set, 35% of segments are contributed by southern speakers, 18% from the East, 18% from Central, 16% from the North and 11% from the West. The accent variation that follows from that spread is recorded in the metadata rather than asserted.
Metadata fields
Monsoon ships 18 columns per segment, of which 12 are metadata, where most public ASR test sets ship an identifier, a transcript and a duration. Demographic fields are complete or near-complete; contributors consented to this use.
| Group | Fields |
|---|---|
| Segment | id, audio, audiolengths, language |
| Reference | lattice (Hindi) or text (Indian English) |
| Speaker | speakerid, gender, dateofbirth |
| Background | occupation, educationalbackground, maritalstatus, income |
| Geography | nativedistrict, nativestate, currentcity, yearsspentincurrentdistrict |
| Recording | devicemanufacturer, devicemodel |
The two languages have different geographic shapes, and the shape is informative. The Hindi sets concentrate in the Hindi belt, with Uttar Pradesh accounting for roughly 40% of speakers, which is what a Hindi corpus sampled by population should look like. The Indian English sets are much flatter: no state exceeds 13%, and a third of speakers come from outside the eight largest. Public and private halves match closely on both.
Indian state boundaries were drawn along linguistic lines, so district and state carry real accent signal, which is why these fields are released rather than summarised away. Analysis of this kind has been reported at scale for Indian ASR: district-level error rates spanning roughly 4% to 44%, with underrepresented regions well behind the Hindi belt and the metros, plus disaggregation by audio quality, speaking rate, utterance duration, gender, age and device. Those runs were on a closed benchmark. Monsoon makes the same class of analysis possible on a public leaderboard test set.
Collection and quality control
Broad geographic coverage requires recruitment across hundreds of districts rather than longer sessions from fewer speakers, and distributed recruitment at this scale introduces failure modes that a smaller collection does not face: contributors gaming the task, played-back audio submitted as live speech, and inattentive annotation. Each is addressed by an explicit check.
Recruitment and recording: Contributors were recruited through the Voice Arena community, a global digital platform whose reach extends into the rural and semi-urban districts that speech corpora rarely cover. Pairs then recorded two-person conversations over a peer-to-peer interface, dual-channel, on assigned everyday topics. Contributors used their own handsets and their own connections. Many of those are low-end devices on unstable bandwidth, which is why that condition is present in the released audio rather than filtered out of it. Prospective contributors completed a language proficiency screening before being granted recording access, were compensated, and provided informed consent covering use in training and distribution. A per-speaker duration cap, calibrated per language to the population size and geographic distribution of its speakers, prevented a small number of prolific contributors from dominating a language or region; more than half of the speakers in these sets contribute exactly one segment.
Elicitation: Eliciting spontaneous speech at scale presents its own difficulty, as contributors tend to produce short and sparse responses without structured guidance. Each conversation was therefore seeded with an open-ended narrative cue and progressively revealed follow-up questions, spanning domains including travel, healthcare, agriculture, education and digital services, guiding the exchange toward extended description without scripting it. Candidate topics were generated with large language models, then reviewed and localised by native-speaker linguists.
Quality control: Every recording passed a set of gating checks prior to transcription. The spoken language was verified against the assigned language using language identification models trained on human-annotated data across more than 30 languages. Speaker gender was confirmed against the self-reported label using a dedicated classifier, applied as corroboration of the self-report rather than as a replacement for it. A further model distinguished genuine spontaneous conversation from pre-recorded or played-back audio. Signal-to-noise ratio estimation removed recordings degraded beyond intelligibility, while natural environmental background noise was deliberately preserved so that the acoustic realism of in-the-wild speech is retained. Recordings clearing these checks were segmented by voice activity detection, split at two seconds of continuous silence or at a fifteen-second soft cap closed at the next detected silence. Segmentation was applied independently per channel, so every segment is single-speaker and single-channel. Segments then passed a DNSMOS P.808 check.
Transcription: Reference transcripts are human work. A first draft was generated by internal ASR models trained on in-domain data, none of which appear on any public leaderboard, so no system evaluated on these sets contributed to the references it is scored against. Every subsequent stage was performed by native-speaking linguists under a five-level protocol built on a strict separation of labour, in which each correction round is followed by an independent verification round performed by a different annotator, so no linguist audits their own output. A linguist first corrects the draft segment by segment against the acoustic signal; a second re-verifies it and flags residual disagreements. Subsequent levels iterate this cycle with fresh annotators, progressively resolving ambiguous phonetic realisations, code-switching boundaries, named entities, and orthographic consistency across spelling variants. Numerals are written as words, so that the transcript corresponds directly to what was spoken. Segments still flagged at the final level were returned for re-transcription before admission. Annotator behaviour was monitored automatically throughout, flagging submissions containing characters outside the target script, unnatural character or word repetitions, and unusually low or high edit counts.
Regional variation
What follows is one example, run on the public Indian English split, to show the kind of evaluation the metadata makes possible. It is not the finding the sets exist to deliver; it is an illustration of what becomes answerable once every clip carries a speaker.
Eight models on the leaderboard land between 4.81 and 4.99 WER on this set. That is 0.18 points from best to worst, inside what five hours can resolve. Ranked on the corpus, they are the same model.
Grouping speakers by region tells a different story. Each speaker's native district is rolled up to its zonal council, the Ministry of Home Affairs grouping of Indian states, giving five well-sampled zones. openai/whisper-large-v3-turbo varies by 0.46 points across them. mistralai/Voxtral-Mini-3B-2507, fourteen hundredths of a point behind it on the corpus, varies by 1.68, running 4.38 in the Central zone against 6.06 in the East. Two systems that are indistinguishable on the leaderboard differ almost fourfold in how much their accuracy depends on where the speaker is from.
Which zone is hardest is not fixed either. ibm-granite/granite-speech-3.3-2b is worst in the North, microsoft/VibeVoice-ASR-HF in the South, mistralai/Voxtral-Mini-3B-2507 in the East. If a single region were simply harder to transcribe, every model would rank the zones the same way. They do not, which points at the models rather than the audio.
Region is one of twelve recorded attributes, and the zones above are a coarse rollup of 428 districts. The same breakdown runs on age, education, occupation and handset, and the released files carry everything needed to reproduce it. None of it is available for a test set that records only what was said.
Orthographic variation in Hindi
English orthographic variation is bounded. British against American spelling, punctuation, casing, digits against words: a normaliser can map most of it to a single form, and the leaderboard's does. Hindi is not bounded in the same way. Everyday speech is heavily code-mixed, English-origin words have no settled Devanagari spelling, and compound forms are written joined or separated according to preference. A single phrase can have ten or more valid written forms, and no fixed mapping collapses them, because there is no canonical side to map to.
Scored with a single reference, WER rewards a system for producing the spelling the annotator happened to choose. Two systems that recognised the audio equally well can differ by several points on orthography alone.
The Hindi sets therefore ship a lattice: for each span of the transcript, the set of written forms accepted as correct. Building it is manual work. Candidate variants are drawn from multiple ASR transcripts of the same audio and expanded with language models, then native-speaker linguists decide which are valid for that utterance and prune the rest, so only forms consistent with what was said are admitted.
Thus, for Hindi we report the Orthographically-Informed Word Error Rate (OIWER), introduced by AI4Bharat, instead of WER. A hypothesis is aligned against the accepted set at each span, so any admitted form counts as correct and only genuine recognition errors are charged.
To quantify the effect, the same hypotheses were scored twice. Flattening each lattice to its first variant per span yields a single string reference of the kind a conventional benchmark provides; any of the admitted variants would serve equally well, and a different choice would yield a different reference. Scored against the flattened reference, error rates rise for every system, and they do not rise uniformly. Rankings change as a consequence. The figure below shows two pairs of systems that reverse order between the two references: under a single reference a system is rewarded in part for reproducing the annotator's orthography, whereas the lattice scores only recognition.
We also open source our implementation, voi-oiwer, so every result on these sets can be reproduced directly.
Getting evaluated
For the private splits, get your model on the Open ASR Leaderboard and the Hugging Face team will run the evaluation. As before, the process for adding a model to the leaderboard takes place on the Open ASR Leaderboard GitHub:
- Open a pull request, a model checklist will appear. As before, you should report your results on the public datasets.
- We will verify the results on the public sets and compute the metrics on the private ones.
- Confirm the results we've obtained.
Indian-English joins the main leaderboard as Voice Arena Monsoon, in the default column set rather than as an opt-in toggle, so it contributes to the headline Average WER for every model. The private split feeds the aggregated Private (conversational) column alongside Appen and DataoceanAI data. The public and private Hindi appears in the Multilingual tab, where a model is ranked only if it supports every selected language, making that column a like-for-like comparison. Alternative, select "Hindi" from the "Language dataset breakdown" dropdown menu.
What comes next
Hindi is a sharp case of a general problem. Any language written more than one way, spoken by people a benchmark has not sampled, carries both of the failures described here. These sets do not fix them. What they add is a way to see them: a test set on a leaderboard the field already watches, carrying enough about each speaker and each reference that a difference between two systems can be traced to who was talking and how they write it, instead of disappearing into one number.
These four sets are part of Monsoon, Voice Arena's broader dataset initiative for the Global South.
Models mentioned in this article 4
Spaces mentioned in this article 1
Papers mentioned in this article 2
audio speech benchmark ## Measuring benchmark optimization in speech recognition +3 57 August 21, 2026 tlebryk02, et. al.
Models mentioned in this article 4
Spaces mentioned in this article 1
Papers mentioned in this article 2