LLMs can't jump, Part II. Mathematicians Timothy Gowers and Peter Sarnak credit large language models with serious math skills but see hard limits for genuinely new ideas. Gowers argues current models are good at combining known methods and trying many search paths but lack the intuition to pick the few productive routes in a vast search space. Sarnak agrees: AI can derive results from existing theory but fails to develop the abstractions that underpin major proofs when starting from an elementary question.
DeepMind researcher Tom Zahavy reached a similar conclusion. In his paper "LLMs Can't Jump," he pins the bottleneck on "manipulative abduction," the ability to invent new foundational assumptions with no linguistic precedent. World models could offer a way forward. These assessments feed into a broader debate about whether LLMs are actually becoming more versatile or "just" getting better at benchmarks and familiar problem spaces.
AMS
Gowers