# NVIDIA 研究：AdamW 存在规模天花板，SOAP/Muon 更优

- 来源：elvis (@omarsar0)
- 发布时间：2026-07-27 04:09
- AIHOT 分数：48
- AIHOT 链接：https://aihot.virxact.com/items/cms28z1bx00o4ro25guub2vez
- 原文链接：https://x.com/omarsar0/status/2081471875885203791

## AI 摘要

NVIDIA 新研究指出，在高达 1 亿 token 的批大小下，AdamW 优化器训练稳定性下降，而 SOAP 和 Muon 能保持稳定。团队通过每步 QR 正交化消除了 SOAP 在大批大小下的损失尖峰，并在数万亿 token 训练的多十亿参数模型上验证了两种优化器持续优于 AdamW。他们还提出了与 Megatron-LM 兼容的逐层分布式优化器，在不牺牲收敛收益的前提下平衡内存与通信。

## 正文

New research from NVIDIA.

Does AdamW have a scale ceiling？

This work claims yes， and shows where it sits. At batch sizes up to 100M tokens for next-token prediction， SOAP and Muon maintain training stability and quality while AdamW degrades.

Higher-order optimizers have promised faster convergence for a while. The standing objection has been computational cost and numerical stability at scale.

The team identifies instabilities in SOAP at large batch sizes and eliminates the loss spikes with per-step QR orthogonalization and improved preconditioning strategies. They also measure Muon's orthogonalization quality empirically rather than assuming it.

On multi-billion-parameter models trained over trillions of tokens， both optimizers consistently beat AdamW. A layer-wise distributed optimizer compatible with Megatron-LM balances memory and hides communication without approximating the optimizer math， so the convergence benefit survives the systems layer.

Paper： https://arxiv.org/abs/2607.20548

Learn to build effective AI agents in our academy： https://academy.dair.ai/
