Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameter language model based on the Hierarchical Reasoning Model (HRM) architecture, that is trained from scratch and delivers highly competitive performance for English and sets a new state of the art for Danish using only permissible post-training data. Trained on a mixture of 161 datasets, Mimir v1 outperforms the original HRM-Text 1B and competes with larger frontier models like Qwen 3.5 4B and Gemma 4 E2B, tested across 20 benchmarks for English, Math & Code and Danish. The model is available on the Hugging Face Hub: https://huggingface.co/danish-foundation-models/DFM-Mimir
DFM Mimir v1:仅用合规后训练数据,1B 参数开源 HRM 模型实现前沿性能
AI 导读
DFM 发布 Mimir v1,一个基于 Hierarchical Reasoning Model(HRM)架构的 10 亿参数语言模型,从头训练并仅使用合规后训练数据,在英语上表现极具竞争力,并在丹麦语上创下新 SOTA。
HuggingFace Daily Papers(社区热门论文)
61
AI 编辑部评分,满分 100DFM Mimir v1:仅用合规后训练数据,1B 参数开源 HRM 模型实现前沿性能
DFM 发布 Mimir v1,一个基于 Hierarchical Reasoning Model(HRM)架构的 10 亿参数语言模型,从头训练并仅使用合规后训练数据,在英语上表现极具竞争力,并在丹麦语上创下新 SOTA。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org