We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.
MameLoshnLM:首个开源意第绪语 8B 语言模型与评测基准
AI 导读
MameLoshnLM 是首个专为意第绪语构建的开源 8B 参数语言模型,基于 Llama 3.1 8B 继续预训练而来。研究同时推出 Oytser 高质量意第绪语预训练语料库和 Kashes 多任务评测基准。在基准任务上,MameLoshnLM 优于同规模开源基线,且更好地捕捉了意第绪语特有的词汇与形态模式。
HuggingFace Daily Papers(社区热门论文)
53
AI 编辑部评分,满分 100MameLoshnLM:首个开源意第绪语 8B 语言模型与评测基准
MameLoshnLM 是首个专为意第绪语构建的开源 8B 参数语言模型,基于 Llama 3.1 8B 继续预训练而来。研究同时推出 Oytser 高质量意第绪语预训练语料库和 Kashes 多任务评测基准。在基准任务上,MameLoshnLM 优于同规模开源基线,且更好地捕捉了意第绪语特有的词汇与形态模式。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org