Impressive new paper from Meta.
(bookmark it)
Scaling laws assume model size and training data act on loss independently.
This work introduces Skaling law, which couples capacity and data through a single interaction exponent. The extra term cuts mean absolute percentage error by 1.5x to 3x across both interpolation and extrapolation.
The largest corrections land in the data-scarce and heavy-overtraining regimes where the standard Chinchilla and Kaplan forms drift.
Paired with a sparse grid restricted to low-compute runs, it extrapolates the full grid using roughly 10x less compute than a uniform sweep.
Why does it matter?
Deployment now happens well past compute optimal. A law that stays accurate there, and that can be fit from small runs, changes how a pretraining budget gets planned.
Paper: https://arxiv.org/abs/2608.07222
Track more trending AI papers in our academy: https://academy.dair.ai/