I had to look twice at this table.
I know they are just numbers, but how the heck do you beat a competitor model released just a couple of days ago on some of the toughest benchmarks out there?
Something is different about GPT-6 Astra.
Elvis Saravia 引用 OpenAI 官方内容评论 GPT-6 Astra:该模型在 FrontierMath Tier 4、ARC-AGI 3、TerminalBench-4.0 上达到 SOTA,并在 Terminal-Bench Science 0.1 与 HealthBench Pro 上领先。作者对表格数据表示惊讶,称其击败了几天前刚发布的竞品模型,并说 GPT-6 Astra 有些不同寻常。图中显示 ARC-AGI-3 得分 99.9%、FrontierMath Tier 4 (v2) 97.6%。
Elvis Saravia 引用 OpenAI 官方内容评论 GPT-6 Astra:该模型在 FrontierMath Tier 4、ARC-AGI 3、TerminalBench-4.0 上达到 SOTA,并在 Terminal-Bench Science 0.1 与 HealthBench Pro 上领先。作者对表格数据表示惊讶,称其击败了几天前刚发布的竞品模型,并说 GPT-6 Astra 有些不同寻常。图中显示 ARC-AGI-3 得分 99.9%、FrontierMath Tier 4 (v2) 97.6%。
I had to look twice at this table.
I know they are just numbers, but how the heck do you beat a competitor model released just a couple of days ago on some of the toughest benchmarks out there?
Something is different about GPT-6 Astra.
来源:elvis· x.com