看起来相当令人印象深刻。多份报告表明这是一次真正的进步。
作为一个近十年来一直倡导(神经)符号世界模型、且常常遭到强烈敌意的人,看到 OpenAI 的一款产品在其一些最令人印象深刻的计算过程中明确创建并操作符号世界模型,这让我感到极其欣慰。
我们不知道的是这种能力有多稳健。这才是关键问题。
在 ARC-AGI 上取得成功固然很棒、令人印象深刻,但——尽管该任务名为“AGI”——这并不能证明就是 AGI;我怀疑我们会在开放式现实世界任务中看到大量问题。与其他近期模型一样,我预计其在可验证领域表现最佳。
而作为一名科学家,令人失望的是,我们(目前?)对该系统实际如何运作知之甚少。
如果不能更清楚地了解其内部机制,我对它能做什么、不能做什么,以及我们可能遇到哪些新风险,都会感到不那么有信心。我怀疑世界还没有准备好。
一如既往,热衷者提前看到了,怀疑者则没有。这是一种明智的营销策略,但往往具有误导性。我们经常看到的情况是,最初的热情会随着时间推移而降温。我怀疑这次也会如此。
新系统似乎比之前的系统更难监控,从安全角度来看这并不理想。人们确实不希望能力增强的同时可监控性却下降。但另一方面它似乎也更易对齐,原因尚不清楚。
很想知道 Astra 能否在 Miles Brundage 和我在 2024 年底打赌的十项任务中的任何一项上取得进展。(据我所知,迄今为止还没有任何 AI 在任何一项上取得成功。)
这个观点非常初步,有待于获得更多关于该系统工作原理的信息,以及对其局限性进行更详细的考察。
Looks to be pretty impressive. Multiple reports suggest it is a genuine advance.
As someone who has campaigned for nearly a decade for (neuro)symbolic world models, often to exceptional hostility, it is extraordinarily vindicating to see that a product from OpenAI explicitly creates and manipulate symbolic world models in the course of some of its most impressive computations.
What we don’t know is how robust that capability is. That is THE key question.
Success on ARC-AGI is great and impressive, but not —despite the name of the task—proof of AGI; I suspect we will see loads of problems with open-ended real world tasks. As with other recent models I would suspect best performance in verifiable domains.
And as a scientist, it’s disappointing that we don’t (yet?) know much about how the system actually works.
Without a clearer sense of what’s under the hood, I feel less confident about both what it can and can’t do, and what new risks we may encounter. I doubt the world is ready.
As ever, enthusiasts got an advance look; skeptics did not. That’s a sound marketing strategy, but it often turns out to be misleading. What we have often seen is initial enthusiasm that gets tempered over time. I suspect we will see that here as well.
The new system appears to be less monitorable than prior systems, which is not great from a safety perspective. One really doesn’t want more capability in conjunction with less monitorability. But also more alignable, not sure why.
Would be great to see whether Astra can make progress on any of the ten tasks that Miles Brundage and I bet on at the end of 2024. (No AI to date has succeeded on any, AFAIK.)
This hot take is VERY tentative, pending more information about how the systems works and more detailed examination of what its limitations are.