Dongxi 东锡 NLP · @dongxi_nlp · X·2026-08-18 03:40·8天前
AI 导读

马东锡 NLP 引用 Jason Wei 观点,质疑“1B 认知核心+工具调用”范式:无需工具、快速自然地完成任务至关重要,参数化内化知识有信息上限,追求最高质量智能仍需更大模型。旧范式是预训练算力扩展,当前是推理时扩展,下一扩展轴成谜。

Dongxi 东锡 NLP@dongxi_nlp
41AI 编辑部评分,满分 100
2026-08-18 03:40· 8天前
AI 导读

马东锡 NLP 引用 Jason Wei 观点,质疑“1B 认知核心+工具调用”范式:无需工具、快速自然地完成任务至关重要,参数化内化知识有信息上限,追求最高质量智能仍需更大模型。旧范式是预训练算力扩展,当前是推理时扩展,下一扩展轴成谜。

This is illuminating, but it also raises a deeper question.

The old scaling paradigm: More compute during pretraining -> smarter model

The current scaling paradigm: More compute during inference -> better answer

The next scaling paradigm: ? -> ?

If small models + test-time scaling + tools are beginning to hit a ceiling, and the future shifts back toward much larger models that can solve more problems in a single forward pass, what exactly is the next scaling axis?

Jason WeiWhen language models first started using tools well, I was sympathetic to the narrative that instead of scaling up language models, all we needed was a strong e...

来源:Dongxi 东锡 NLP· x.com