Muon optimizer nearly doubled an AI agent's success, showing its reinforcement-learning value depends on credit assignment and learning rate.
With GiGPO, which compares actions from repeated states, Muon raised late success from 0.29 to 0.55.
An optimizer cannot be judged alone because the surrounding reinforcement-learning setup may decide whether it helps or fails.
- arxiv. org/abs/2607.16169