显然令人印象深刻。但与许多其他事物一样,我们应当以审慎的态度看待它。
在今天早上发给我的电子邮件中,卡尔·纽波特提出了若干精辟见解,并允许我分享。他既总结了研究发现,也指出了其中一些局限性:
OpenAI 使用了一个尚未发布的新推理模型,成功找到了一个反例,推翻了数学家保罗·埃尔多斯80年前提出的一个离散几何猜想。该模型针对所谓的链式推理进行了调优,在这种推理模式下,模型会针对它试图解决的任何问题无休止地“大声思考”。这种方法让你能够利用大语言模型近似实现类似记忆和动态计算的功能,而大语言模型原本是静态且前馈的。专业数学家从模型推理的长篇记录中识别出了这个反例,然后提取出关键部分,并以更标准的风格将其重写为更简洁的证明。这个模型是如何解决人类数学家未能解决的问题的呢?在OpenAI发布的一篇配套文章中,审阅了模型全部输出的数学家托马斯·布卢姆指出了促成这个反例适合由大语言模型辅助发现的几个因素。他指出,尽管这个猜想年代久远,但研究它的人大多认同埃尔多斯最初认为它正确的观点,因此专注于试图证明它。而基于大语言模型的工具所做的,则是系统地应用和扩展现有技术,寻找该猜想为假的证据。布卢姆说:“(人工智能)在此的成功与之前的成就如出一辙:它常常通过坚持探索人类可能认为不值得花时间研究的路径,结合超乎常人的耐心和对大量技术工具的熟悉,产生最令人惊讶的结果。”我个人的几点观察: (1)非数学专业人士可能不熟悉,近年来大语言模型技术已与现有的计算机辅助数学工具相结合,通过系统且耐心地探索那些令大多数人类数学家感到过于繁琐而失去兴趣的技术和问题空间角落,来寻求新的数学成果。因此,OpenAI这一新成果真正的技术亮点在于,链式推理能够在没有大多数现有工具所使用的更复杂框架的情况下,完成这种系统性的求解。话虽如此,这里使用的内部模型(许多人认为这是OpenAI对规模极其庞大的Mythos大语言模型的回应)的提示词成本可能也同样极其高昂。人工智能辅助数学的未来,很可能将聚焦于更小、更便宜、专为数学调优的大语言模型,并结合更强大的框架。因此,这个实验可能更多是为了营销其新模型的能力,而非真正试图推进计算机辅助数学。 (2)我认为,说这些人工智能辅助数学的例子意味着模型在某种程度上比人类数学家“更聪明”并不准确。我认为一个更好的类比可能是,计算机工具如何帮助建筑师设计出更大胆、更复杂的设计(比如弗兰克·盖里设计的斯塔塔中心,我在麻省理工学院攻读计算机科学博士和博士后期间就在那里工作)。这些工具并非比人类更优秀的建筑师,而是让人类成为了能力更强的建筑师。 (3)从商业角度来看,我实际上认为这一公告对OpenAI来说未必是好消息。几乎没有比专业学术数学更小、利润更低的市场了。OpenAI将部分顶尖技术人才(如诺姆·布朗)投入到这个领域,这一事实突显了其最令人印象深刻的结果仅限于少数非常适合大语言模型的领域(即数学和计算机编程),这就像醉汉在路灯下寻找钥匙一样。如果这个模型在更通用的方面表现出色,那么显然更好的例子应该是解决那些能直接、显著地为它们希望争取的特定类型客户带来巨额收入或节省成本的问题,或实现流程自动化。 总之:人工智能在数学领域的作用确实重要且令人兴奋。我能想到自己职业生涯中研究过的许多成果,如果当时能使用最新一代的工具,我本可以进展得更快或更全面。但人工智能与数学的这种交叉也高度特定于该领域,并且比简单地将人工智能系统想象成越来越聪明的独立数学家更为微妙和复杂。人们应当警惕从数学和编程等领域,对这些模型的其他潜在应用做出过于宏大的概括。
OpenAI 使用了一个尚未发布的新推理模型,帮助识别出一个反例,推翻了数学家保罗·埃尔多斯 80 年前提出的一个离散几何猜想。该模型针对所谓的链式推理进行了调优,在这种推理方式中,模型会针对它试图解决的任何问题无休止地“大声思考”,这种方法让你能够利用大语言模型近似实现类似记忆和动态计算的功能,而大语言模型原本是静态且前馈的。专业数学家从模型推理的长篇记录中识别出了这个反例,然后提取出关键部分,并以更标准的风格将其重写为更简洁的证明。
这个模型是如何解决人类数学家未能解决的问题的呢?在 OpenAI 发布的一篇配套文章中,审阅了模型全部输出的数学家托马斯·布鲁姆指出了促成这一反例适合通过大语言模型辅助发现的各项因素。他指出,尽管这个猜想年代久远,但研究过它的人大多认同埃尔多斯最初认为它正确的观点,因此专注于试图证明它。而基于大语言模型的工具所做的,则是系统地应用和扩展现有技术,寻找该猜想为假的证据。布鲁姆说:“(人工智能的)这一成功与之前的成就相呼应:它常常通过坚持探索人类可能认为不值得花时间而放弃的路径,产生最令人惊讶的结果,将超乎常人的耐心与对大量技术工具的熟悉结合在一起。”
我个人的几点观察:
(1)非数学专业人士可能不太了解,近年来大语言模型技术已与现有的计算机辅助数学工具深度融合,通过系统而耐心地探索那些令大多数人类数学家感到过于繁琐而失去兴趣的技术路径和问题空间角落,来寻求新的数学成果。因此,OpenAI 这项新成果真正的技术亮点在于,链式推理能够在不依赖现有工具中大多数所采用的复杂框架的情况下,实现这种系统性的求解。话虽如此,这里使用的内部模型——许多人认为这是 OpenAI 对规模极其庞大的 Mythos 大语言模型的回应——其提示词成本很可能同样极其高昂。AI 辅助数学的未来,很可能将聚焦于更小、更便宜、专为数学调校的大语言模型,并结合更强大的框架。因此,这项实验或许更多是为了宣传其新模型的能力,而非真正推动计算机辅助数学的进步。
(2)我认为,说这些 AI 辅助数学的例子意味着模型在某种程度上比人类数学家“更聪明”,并不准确。我觉得一个更好的类比是,计算机工具如何帮助建筑师设计出更大胆、更复杂的作品(比如弗兰克·盖里设计的斯塔塔中心,我在麻省理工学院攻读计算机科学博士和从事博士后研究时就在那里)。这些工具并非比人类更优秀的建筑师,而是让人类成为了能力更强的建筑师。
(3)从商业角度来看,我实际上认为这一公告对 OpenAI 未必是好消息。几乎没有比专业学术数学更小、更不赚钱的市场了。OpenAI 将部分顶尖技术人才(如 Noam Brown)投入这一领域,恰恰凸显了其处境——就像醉汉在路灯下找钥匙一样,他们最令人瞩目的成果仅限于少数非常适合大语言模型的领域(即数学和计算机编程)。如果这个模型在更通用的方面表现出色,那么显然更好的例证应该是解决那些能直接、显著地为它们希望争取的客户企业创造巨额收入或节省成本的问题,或实现相关流程自动化。
总之:人工智能在数学领域的角色确实重要且令人振奋。我能想到自己职业生涯中研究过的许多成果,如果当时能使用最新一代工具,本可以进展更快或研究更全面。但人工智能与数学的这一交集也极具领域特殊性,其微妙复杂程度远超简单地将 AI 系统想象成日益卓越的独立数学家。人们应当警惕从数学和编程等领域,对这些模型的其他潜在应用做出过于宽泛的推论。
Kareem Carr 对数学家 Tim Gowers 的论述也给出了审慎的看法,总体方向趋于一致:
除此之外,我们并不真正了解该模型是如何运作的、如何训练的,也不清楚其成果在数学领域内外,乃至更开放、更日常的现实世界中具有多大通用性。我们对该模型在其他基准测试上的表现、能否解决模型幻觉问题,以及运行成本究竟有多高一无所知。
我们同样不知道他们尝试了多少次提示词,又有多少次未能成功;我们只看到了分子,却看不到分母。
这无疑是一个有趣的结果;至于它在现实世界中究竟意味着什么,那就只能任人猜测了。
另一则重大消息来自《华尔街日报》Berber Jin 的独家报道:Anthropic 预计将迎来其历史上首个(略微)盈利的季度。
这确实令人惊叹——前提是它真的能实现——但如果真能实现,很大程度上是因为(正如昨日 SpaceX 在 S-1 首次公开募股文件中披露的那样)Anthropic 在该季度获得了 SpaceX 提供的一次性(非经常性)算力折扣。具体金额未公布,但该折扣很可能超过预计的 5.59 亿美元利润。背景信息很重要,后续季度能否盈利尚不明朗。
Ed Zitron 对此表达了更强烈的质疑,并就相关会计问题提出了一些疑问,详情见此处。
今早我看到的最惊人的数据是:英伟达在循环融资上的投入如此之大,其现金流正趋近于零。
在排除所有这些因素后,要真正了解实际情况实在太难了。
买家自负,IPO 认购者。
Clearly impressive. But as with so much else, it should be viewed with skepticism.
In an email to me this morning, Cal Newport made a number of good points that he said I could share, summarizing both what was found and some limitations:
OpenAI used a new reasoning model (not yet released) to help identify a counterexample that disproved a conjecture from discrete geometry first proposed by Paul Erdos 80 years ago. The model was tuned for so-called chain-of-thought reasoning where the model endlessly “thinks out loud” about whatever it is trying to solve, an approach that lets you approximate something like memory and dynamic computation using LLMs, which are otherwise static and feed-forward. Professional mathematicians identified the counterexample from within a long transcript of the model’s reasoning, and then extracted the key parts and rewrote it as a more succinct proof in a more standard style.How was the model able to solve something that human mathematicians had failed to do? In a companion article released by OpenAI, the mathematician Thomas Bloom, who reviewed the full model output, identified the factors that came together to make this counterexample ripe for LLM-aided discovery. He noted that though the conjecture is old, those who have worked on it have largely shared Erdos’s original belief that it was true and therefore focused on trying to solve it. What the LLM-based tool did instead was to systematically apply and extend existing techniques in search of evidence that the conjecture was false. Here’s Bloom: “[the AI’s] success here echoes previous achievements: it often produces the most surprising results by persevering down the paths that a human may have dismissed as not worth their time to explore, combining superhuman levels of patience with familiarity with a vast array of technical machinery.”A few observations of my own:(1) Non-mathematicians might not be familiar with the degree to which LLM-technology has been combined with existing computer-aided math tools in recent years to seek new math results through the systematic and patient exploration of techniques and corners of problem spaces that are too exhausting to interest most human mathematicians. The real technical headline of the new OpenAI result, therefore, is that chain-of-thought reasoning was able to accomplish this type of systematic solving without the much more intricate scaffolding used in most of these existing tools. That being said, the internal model used here, which many assume is OpenAI’s response to the truly massive Mythos LLM, is likely similarly massively expensive to prompt. The future of AI-assisted math will likely focus on smaller, cheaper, math-tuned LLMs combined with more powerful scaffolding. So, this experiment might be more about marketing the power of their new model than trying to actually advance computer-aided math.(2) I don’t think it’s accurate to say these examples of AI-supported mathematics mean the models are somehow “smarter” than human mathematicians. I think a better analogy might be how computer tools helped architects produce much more daring and complicated designs (like the Frank Gehry-designed Stata Center where I did my CS doctoral and postdoctoral work at MIT). These tools weren’t better architects than humans but made humans more capable architects.(3) From a business perspective, I actually think this announcement isn’t necessarily good news for OpenAI. There are few markets smaller and less lucrative than professional academic mathematics. The fact that this is the area where OpenAI is dedicating some of their top technical talent (like Noam Brown) underscores the degree to which, like the drunk searching for their keys under the streetlight, their most impressive results are limited to the smaller number of areas that are well-suited to LLMs (i.e., math + computer coding). If this model was brilliant in some more general way, obviously the better examples would be solving problems or automating processes that directly and obviously generate massive revenue or savings for the specific types of companies they hope to make their customers.In conclusion: AI’s role in math is genuinely important and exciting. I can think of any number of results I’ve worked on in my career where I could have moved faster or been more comprehensive if I had access to the latest generation of tools. But this intersection of AI and math is also very specific to this field and more nuanced and complicated than simply imagining AI systems as standalone mathematicians who are becoming increasingly brilliant. One should be wary of making ambitious generalizations from fields like math and coding to other potential applications of these models.
OpenAI used a new reasoning model (not yet released) to help identify a counterexample that disproved a conjecture from discrete geometry first proposed by Paul Erdos 80 years ago. The model was tuned for so-called chain-of-thought reasoning where the model endlessly “thinks out loud” about whatever it is trying to solve, an approach that lets you approximate something like memory and dynamic computation using LLMs, which are otherwise static and feed-forward. Professional mathematicians identified the counterexample from within a long transcript of the model’s reasoning, and then extracted the key parts and rewrote it as a more succinct proof in a more standard style.
How was the model able to solve something that human mathematicians had failed to do? In a companion article released by OpenAI, the mathematician Thomas Bloom, who reviewed the full model output, identified the factors that came together to make this counterexample ripe for LLM-aided discovery. He noted that though the conjecture is old, those who have worked on it have largely shared Erdos’s original belief that it was true and therefore focused on trying to solve it. What the LLM-based tool did instead was to systematically apply and extend existing techniques in search of evidence that the conjecture was false. Here’s Bloom: “[the AI’s] success here echoes previous achievements: it often produces the most surprising results by persevering down the paths that a human may have dismissed as not worth their time to explore, combining superhuman levels of patience with familiarity with a vast array of technical machinery.”
A few observations of my own:
(1) Non-mathematicians might not be familiar with the degree to which LLM-technology has been combined with existing computer-aided math tools in recent years to seek new math results through the systematic and patient exploration of techniques and corners of problem spaces that are too exhausting to interest most human mathematicians. The real technical headline of the new OpenAI result, therefore, is that chain-of-thought reasoning was able to accomplish this type of systematic solving without the much more intricate scaffolding used in most of these existing tools. That being said, the internal model used here, which many assume is OpenAI’s response to the truly massive Mythos LLM, is likely similarly massively expensive to prompt. The future of AI-assisted math will likely focus on smaller, cheaper, math-tuned LLMs combined with more powerful scaffolding. So, this experiment might be more about marketing the power of their new model than trying to actually advance computer-aided math.
(2) I don’t think it’s accurate to say these examples of AI-supported mathematics mean the models are somehow “smarter” than human mathematicians. I think a better analogy might be how computer tools helped architects produce much more daring and complicated designs (like the Frank Gehry-designed Stata Center where I did my CS doctoral and postdoctoral work at MIT). These tools weren’t better architects than humans but made humans more capable architects.
(3) From a business perspective, I actually think this announcement isn’t necessarily good news for OpenAI. There are few markets smaller and less lucrative than professional academic mathematics. The fact that this is the area where OpenAI is dedicating some of their top technical talent (like Noam Brown) underscores the degree to which, like the drunk searching for their keys under the streetlight, their most impressive results are limited to the smaller number of areas that are well-suited to LLMs (i.e., math + computer coding). If this model was brilliant in some more general way, obviously the better examples would be solving problems or automating processes that directly and obviously generate massive revenue or savings for the specific types of companies they hope to make their customers.
In conclusion: AI’s role in math is genuinely important and exciting. I can think of any number of results I’ve worked on in my career where I could have moved faster or been more comprehensive if I had access to the latest generation of tools. But this intersection of AI and math is also very specific to this field and more nuanced and complicated than simply imagining AI systems as standalone mathematicians who are becoming increasingly brilliant. One should be wary of making ambitious generalizations from fields like math and coding to other potential applications of these models.
Kareem Carr also gave a tempered view of what the mathematican Tim Gowers had written, converging in the same general direction:
Beyond that, we don’t really know how the model worked, how it was trained, or how general the result is, either within mathematics or outside, in the more open-ended everyday world. We have exactly zero data on how it works on other benchmarks, whether it can solve hallucinations, or how much it costs to run.
We also don’t know how many prompts they tried, and how many didn’t work; we have a numerator but not a denominator.
Definitely it is an interesting result; what it actually means in the real world is anybody’s guess.
The other big news per a scoop from Berber Jin at the WSJ, is that Anthropic is projecting its first (slightly) profitable quarter ever.
That is amazing—assuming it actually happens —but if it does it will be in no small part because (as revealed yesterday in SpaceX’s S-1 IPO filing) Anthropic is getting a one-time (nonrecurring) discount for that quarter on compute from SpaceX. The exact number is not given but that discount may well be bigger than the projected $559M profit. Context matters, and it is unclear whether subsequent quarters will be profitable.
Ed Zitron expresses even more skepticism, and raises some questions about accounting, here.
The wildest stat I saw this morning was this: Nvidia is spending so much on circular financing its cash flow is headed towards zero.
It’s so hard to know what’s really going on, factoring all that out.
Caveat emptor, IPO buyers.