Gary Marcus:The Road to AI We Can Trust(RSS)
59AI 编辑部评分,满分 100

关于 Astra 与数学的两则重要更新:Anthropic 数学家 24 小时复现 OpenAI 半数结果

2026-08-04 00:21· 46分钟前· Gary Marcus
跳到正文
AI 摘要

Gary Marcus 称 OpenAI 的 Astra 可能并非其宣传的突破,Anthropic 数学家 Levent Alpöge 用已公开的 Fable 模型在 24 小时内复现了 OpenAI 半数结果,质疑其真实进展。

First, abstracting from all the words in my long analysis on OpenAI Astra yesterday, the overall point was that Astra may not be the breakthrough that OpenAI wants you to think it was. I expressed concerns about how little they revealed about what they did. Moreover, I thought I didn’t stress it there, we all know they are facing huge economic and competitive pressures; they need a win. I am often left with the feeling that their public facing communications are better marketing than they are science.

My jaw dropped a couple hours after my own post, reading this from Levent Alpöge, an accomplished mathematician who works at Anthropic:

One (obviously smart) guy at Anthropic was able to “replicate” (as best he could given OpenAI’s very sparse report of what they actually did) half of OpenAI’s results within 24 hours. And not with Astra or some other new model, but with the already publicly-released Fable.

I had a couple reactions to this.

a. Alpöge’s fast partial replication made me wonder whether OpenAI’s real advance here was finding (through AI) which open problems were amenable to a certain kind of search and verify technique. OpenAI’s Noam Brown acknowledged, without much detail, that there were failures on some (and maybe many) others:

Unfortunately, beyond that OpenAI does not report which problems they tried and failed at. They have given us a numerator without a denominator, always a worrisome sign.

This doesn’t take away from OpenAI’s accomplishment - if the system solved a bunch of open problems that’s very impressive – but it again raises questions about what the scope is: all math? A certain subset of problems that can be searched in a certain way?

Maybe once that subset has been flagged, maybe many systems can do them. That’s kind of what Levent Alpöge‘s partial replication shows.

And many what was most critical was not some innovation in Astra per se but the idea (probably from a human) to do this particular search.

b. Alpöge‘s fast partial replication undermines the widespread presumption that Astra is a major breakthrough. Astra is an obviously an advance to some degree, but that degree may well turn out to be merely incremental relative to other recent models.

Related to this, I was interested to read in The Information this morning that OpenAI hasn’t decided whether to call Astra (which was an internal code name) GPT 6 or GPT 5.7. If OpenAI truly thought Astra was a quantum leap,, I doubt that would be such a hard choice.

How good Astra is in open-ended real-world problems very much remains to be seen.

Second, a longtime reader of this newletter, Ihor Gowda, pointed me to a fantastic lecture given by Terence Tao, by any measure one of the world’s leading mathematicians; many see Tao as the leading mathematician. His lecture is lucid, accessible, and genuinely important.

What’s more you don’t need to be a professional mathematician to read it. It should be mandatory reading for anyone thinking about how AI will affect mathematics, and strongly recommended for anyone who cares about AI’s impact on the world.

One of the issues that he raises he calls “proof indigestion”; what happens if AI creates a lot of stuff, much of it true, but not necessarily useful:

Tao’s July 26th lecture, written before the Astra announcement, also prefigures another key point that I made yesterday: being good at creating certain sorts of proofs doesn’t mean that an AI system can contribute to all (or even most) of the goals of mathematics. He too distinguishes between solving open problems (which Astra seems strong at, perhaps with some important limits) and building theory (no evidence thus far that Astra can do that):

Tao is no Luddite, or even anti-AI. He is very open to AI, and eager to find a place for it in mathematics. His lecture is a beautiful read on how to do just that — to bring AI into mathematics, as thoughtfully as possible. Highly recommended.

Am working hard to keep you up to date in the most honest way possible and hope you will consider supporting this newsletter.

P.S. Cartoon of the day, illustrating Brandolini’s law:

Noawadays, as the engineer Wouter Vreugdenhil, who brought the cartoon to my attention, points out, “that’s the game.”

关于 Astra 与数学的两则重要更新:Anthropic 数学家 24 小时复现 OpenAI 半数结果

Gary Marcus:The Road to AI We Can Trust(RSS)·2026-08-04 00:21·46分钟前·Gary Marcus
阅读原文· garymarcus.substack.com
AI 摘要

Gary Marcus 称 OpenAI 的 Astra 可能并非其宣传的突破,Anthropic 数学家 Levent Alpöge 用已公开的 Fable 模型在 24 小时内复现了 OpenAI 半数结果,质疑其真实进展。

原文 · 保持原样,未翻译

First, abstracting from all the words in my long analysis on OpenAI Astra yesterday, the overall point was that Astra may not be the breakthrough that OpenAI wants you to think it was. I expressed concerns about how little they revealed about what they did. Moreover, I thought I didn’t stress it there, we all know they are facing huge economic and competitive pressures; they need a win. I am often left with the feeling that their public facing communications are better marketing than they are science.

My jaw dropped a couple hours after my own post, reading this from Levent Alpöge, an accomplished mathematician who works at Anthropic:

One (obviously smart) guy at Anthropic was able to “replicate” (as best he could given OpenAI’s very sparse report of what they actually did) half of OpenAI’s results within 24 hours. And not with Astra or some other new model, but with the already publicly-released Fable.

I had a couple reactions to this.

a. Alpöge’s fast partial replication made me wonder whether OpenAI’s real advance here was finding (through AI) which open problems were amenable to a certain kind of search and verify technique. OpenAI’s Noam Brown acknowledged, without much detail, that there were failures on some (and maybe many) others:

Unfortunately, beyond that OpenAI does not report which problems they tried and failed at. They have given us a numerator without a denominator, always a worrisome sign.

This doesn’t take away from OpenAI’s accomplishment - if the system solved a bunch of open problems that’s very impressive – but it again raises questions about what the scope is: all math? A certain subset of problems that can be searched in a certain way?

Maybe once that subset has been flagged, maybe many systems can do them. That’s kind of what Levent Alpöge‘s partial replication shows.

And many what was most critical was not some innovation in Astra per se but the idea (probably from a human) to do this particular search.

b. Alpöge‘s fast partial replication undermines the widespread presumption that Astra is a major breakthrough. Astra is an obviously an advance to some degree, but that degree may well turn out to be merely incremental relative to other recent models.

Related to this, I was interested to read in The Information this morning that OpenAI hasn’t decided whether to call Astra (which was an internal code name) GPT 6 or GPT 5.7. If OpenAI truly thought Astra was a quantum leap,, I doubt that would be such a hard choice.

How good Astra is in open-ended real-world problems very much remains to be seen.

Second, a longtime reader of this newletter, Ihor Gowda, pointed me to a fantastic lecture given by Terence Tao, by any measure one of the world’s leading mathematicians; many see Tao as the leading mathematician. His lecture is lucid, accessible, and genuinely important.

What’s more you don’t need to be a professional mathematician to read it. It should be mandatory reading for anyone thinking about how AI will affect mathematics, and strongly recommended for anyone who cares about AI’s impact on the world.

One of the issues that he raises he calls “proof indigestion”; what happens if AI creates a lot of stuff, much of it true, but not necessarily useful:

Tao’s July 26th lecture, written before the Astra announcement, also prefigures another key point that I made yesterday: being good at creating certain sorts of proofs doesn’t mean that an AI system can contribute to all (or even most) of the goals of mathematics. He too distinguishes between solving open problems (which Astra seems strong at, perhaps with some important limits) and building theory (no evidence thus far that Astra can do that):

Tao is no Luddite, or even anti-AI. He is very open to AI, and eager to find a place for it in mathematics. His lecture is a beautiful read on how to do just that — to bring AI into mathematics, as thoughtfully as possible. Highly recommended.

Am working hard to keep you up to date in the most honest way possible and hope you will consider supporting this newsletter.

P.S. Cartoon of the day, illustrating Brandolini’s law:

Noawadays, as the engineer Wouter Vreugdenhil, who brought the cartoon to my attention, points out, “that’s the game.”

阅读原文garymarcus.substack.com