# Cultivar：面向数据污染与本地化鲁棒性探究的对比式翻译基准

- 来源：HuggingFace Daily Papers（社区热门论文）
- 发布时间：2026-08-10 08:00
- AIHOT 分数：49
- AIHOT 链接：https://aihot.virxact.com/items/cmsp324kx04i6rorteot48wbl
- 原文链接：https://arxiv.org/abs/2608.09766

## AI 摘要

Cultivar 是 FLORES 的本地化子集，用于评估模型在不同文化语境下的翻译表现，并与非本地化版本对比以探测数据污染和本地化鲁棒性。研究评测了 32 个开源权重模型，发现机器翻译专用模型鲁棒性较差，部分模型可能对 FLORES 过拟合，且模型普遍更擅长翻译美国内容而非其他地区内容。

## 正文

Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.
