A very interesting Google paper.
The main problem with AI is not too little or too much rigor. It has too much engineering rigor, and not enough scientific and philosophical rigor.
It's proven that engineering can work before science can explain it. But AI's rigour isn't balanced, so it's hard to know when it'll fail.
Modern AI is quite demanding on performance but not so confident on explanation and prediction.
The paper identifies three types of rigor: clear ideas, reliable knowledge and reliable performance in the real world.
Conceptual rigor asks whether terms like "smarts" and "understanding" refer to one trait or a cluster of related skills.
They discuss this framework with respect to debates about intelligence, reproducibility, prediction, explanation, benchmarks and already deployed systems.
For science to be rigorous, the results must be able to hold up under new conditions, predict what will happen in the future and explain why something worked or didn't work.
Benchmarks, post-training, tools, monitoring, safety checks: they can all make systems better even without a full theory. This is engineering rigor.
This is why capabilities are rapidly advancing while failure predictions are worsening.
---
- arxiv. org/abs/2607.03634
Title: "The Role of Rigor in AI"