Dongxi 东锡 NLP · @dongxi_nlp · X·2026-09-07 20:16·48分钟前
Dongxi 东锡 NLP@dongxi_nlp
34AI 编辑部评分,满分 100
2026-09-07 20:16· 48分钟前

通过与 Claude code 的合作,作者训练的 MoE 模型每个 token 实际只激活约 365 万参数,而且训练仅使用一颗 Intel Xeon CPU 的单个核心。

我们完全可以接受训练阶段花很多资源,因为训练只发生一次,随后模型可能被调用数十亿次。

所以问题来了,如果目标是极低推理成本,我们究竟应该给一个极小模型喂多少数据?

Greg DiamosI think we should revisit outrageously small neural nets. I needed a 10k tok/s CPU model for data processing. So I gave Anthropic claude code a pile of tokens t...

来源:Dongxi 东锡 NLP· x.com