A new Meta FAIR paper finds a blind spot in the Chinchilla scaling law that becomes expensive when you extrapolate.
Chinchilla can look almost perfect inside a training grid and still mispredict what happens at the frontier.
The problem is: Chinchilla assumes model size and training data help separately, but the experiments show that each changes how useful the other one is.
Skaling adds just 1 extra term to capture that connection, cutting prediction error by about 1.5-3× and getting full-grid Chinchilla-level prediction accuracy with roughly 10× less profiling compute.
On Farseer, the difference becomes huge at frontier scale: at 2×10^25 FLOPs, Chinchilla points to ~380 tokens per parameter, while Skaling and the paper's direct estimates land around 20-40.
- arxiv. org/abs/2608.07222
Title: "Skaling: Chinchilla's Exponents Meet Kaplan's Coupling"