Rohan Paul@rohanpaul_ai
37AI 编辑部评分,满分 100
2026-08-01 21:18· 27分钟前
跳到正文
AI 摘要

耶鲁大学与芝加哥大学论文基于 11,683 篇真实论文构建对照测试,发现 LLM 与人类研究想法的差距不在质量而在范围:模型思维更窄。人类想法涵盖解释机制、检验失败等多种模式,仅 12.1% 主要连接已有工作,而 47.1% 至 64.2% 的 LLM 想法如此,频率约为人类 4 至 5 倍。额外推理反而强化该模式。

This Yale + University of Chicago paper shows that real gap between LLM generated research ideas vs humans is not idea quality, but idea range: LLMs think narrower than human researchers.

The researchers built a controlled test from 11,683 real papers, using each paper's nearby prior work as the shared starting point.

They asked models to propose a new motivation and method from those same prior papers, then compared those ideas with the real human paper ideas.

Instead of asking whether 1 idea looked novel, they labeled each idea by what gap it noticed and what kind of contribution it made.

Human ideas spread across many patterns, such as explaining mechanisms, testing failures, measuring evidence, building systems, and improving efficiency.

Only 12.1% of human ideas were mainly about connecting separate work, but 47.1% to 64.2% of LLM ideas did that, meaning models used this move about 4 to 5 times more often.

Even extra reasoning made this pattern stronger, suggesting models often polish a familiar recipe instead of finding more varied research moves.

---

  • arxiv. org/abs/2607.01233

Title: "Measuring the Gap Between Human and LLM Research Ideas"

Rohan Paul · @rohanpaul_ai · X·2026-08-01 21:18·27分钟前
在 X 看原推· x.com
AI 摘要

耶鲁大学与芝加哥大学论文基于 11,683 篇真实论文构建对照测试,发现 LLM 与人类研究想法的差距不在质量而在范围:模型思维更窄。人类想法涵盖解释机制、检验失败等多种模式,仅 12.1% 主要连接已有工作,而 47.1% 至 64.2% 的 LLM 想法如此,频率约为人类 4 至 5 倍。额外推理反而强化该模式。

This Yale + University of Chicago paper shows that real gap between LLM generated research ideas vs humans is not idea quality, but idea range: LLMs think narrower than human researchers.

The researchers built a controlled test from 11,683 real papers, using each paper's nearby prior work as the shared starting point.

They asked models to propose a new motivation and method from those same prior papers, then compared those ideas with the real human paper ideas.

Instead of asking whether 1 idea looked novel, they labeled each idea by what gap it noticed and what kind of contribution it made.

Human ideas spread across many patterns, such as explaining mechanisms, testing failures, measuring evidence, building systems, and improving efficiency.

Only 12.1% of human ideas were mainly about connecting separate work, but 47.1% to 64.2% of LLM ideas did that, meaning models used this move about 4 to 5 times more often.

Even extra reasoning made this pattern stronger, suggesting models often polish a familiar recipe instead of finding more varied research moves.

---

  • arxiv. org/abs/2607.01233

Title: "Measuring the Gap Between Human and LLM Research Ideas"

在 X 查看原推x.com