Rohan Paul@rohanpaul_ai
28AI 编辑部评分,满分 100

Meta 新论文:小模型或只是欠调优,而非规模扩展的差预测器

2026-08-17 01:39· 17分钟前
AI 导读

Meta 新论文发现,小模型可能只是欠调优,而非规模扩展的差预测器。规模法则在约 4M 参数处涌现,该规模模型可在 1 块 GPU 上 1 小时内完成训练。每规模需 256 组超参数配置时法则才准确,小模型对超参数异常敏感。

New Meta paper shows, small models may not be bad predictors of scale; they may just be getting under-tuned.

Finds scaling laws emerge around 4M parameters, where models can train in under 1 hour on 1 GPU.

Small models are unusually sensitive to hyperparameters.

With 4 or 16 configurations per scale, the law is basically invisible; at 64 it appears but extrapolates poorly, and at 256 it becomes accurate.

As models get larger, good settings occupy more of the search space, while the effective number of hyperparameters near the optimum drops toward 1.

That helps explain why scaling laws look cleaner at larger sizes: the models are easier to tune.

As a check, small-scale runs recover that pre-norm transformers scale better than post-norm over the tested range.

There is a limit: extrapolate too far beyond the measured scales, and statistical errors can dominate.

For model research, cheap experiments may need more tuning breadth, not more model size.

  • arxiv. org/abs/2608.11859

Title: "Small-Scale Experiments: Are They There Yet?"

来源:Rohan Paul · x.com

Meta 新论文:小模型或只是欠调优,而非规模扩展的差预测器

Rohan Paul · @rohanpaul_ai · X·2026-08-17 01:39·17分钟前
AI 导读

Meta 新论文发现,小模型可能只是欠调优,而非规模扩展的差预测器。规模法则在约 4M 参数处涌现,该规模模型可在 1 块 GPU 上 1 小时内完成训练。每规模需 256 组超参数配置时法则才准确,小模型对超参数异常敏感。

New Meta paper shows, small models may not be bad predictors of scale; they may just be getting under-tuned.

Finds scaling laws emerge around 4M parameters, where models can train in under 1 hour on 1 GPU.

Small models are unusually sensitive to hyperparameters.

With 4 or 16 configurations per scale, the law is basically invisible; at 64 it appears but extrapolates poorly, and at 256 it becomes accurate.

As models get larger, good settings occupy more of the search space, while the effective number of hyperparameters near the optimum drops toward 1.

That helps explain why scaling laws look cleaner at larger sizes: the models are easier to tune.

As a check, small-scale runs recover that pre-norm transformers scale better than post-norm over the tested range.

There is a limit: extrapolate too far beyond the measured scales, and statistical errors can dominate.

For model research, cheap experiments may need more tuning breadth, not more model size.

  • arxiv. org/abs/2608.11859

Title: "Small-Scale Experiments: Are They There Yet?"

来源:Rohan Paul· x.com