elvis@omarsar0
37AI 编辑部评分,满分 100
2026-08-11 22:56· 32分钟前
AI 导读

这有点不可思议。 BDH-CQ 在 ARC-AGI 1 上得分 29.5%。还行。但每个任务仅花费 0.0007 美元。 它的实现方式尤其有趣。它在潜在空间中进行循环推理,而非使用 CoT。他们还验证了类似 Transformer 的扩展,参数量最高达 600B。

This is kind of wild.

BDH-CQ scored 29.5% on ARC-AGI 1. Fair. But with just $.0007 per task.

How it's done is particularly interesting. It reasons recurrently in latent space rather than using CoT. They've also verified Transformer-like scaling up to 600B params.

Zuzanna StamirowskaPathway's BDH-CQ model redefines the Cost-Efficiency Frontier on ARC-AGI-1: $.0007 at 29.5%. This is made possible by in-context learning and latent reasoning. ...

来源:elvis · x.com

elvis · @omarsar0 · X·2026-08-11 22:56·32分钟前
AI 导读

这有点不可思议。 BDH-CQ 在 ARC-AGI 1 上得分 29.5%。还行。但每个任务仅花费 0.0007 美元。 它的实现方式尤其有趣。它在潜在空间中进行循环推理,而非使用 CoT。他们还验证了类似 Transformer 的扩展,参数量最高达 600B。

This is kind of wild.

BDH-CQ scored 29.5% on ARC-AGI 1. Fair. But with just $.0007 per task.

How it's done is particularly interesting. It reasons recurrently in latent space rather than using CoT. They've also verified Transformer-like scaling up to 600B params.

Zuzanna StamirowskaPathway's BDH-CQ model redefines the Cost-Efficiency Frontier on ARC-AGI-1: $.0007 at 29.5%. This is made possible by in-context learning and latent reasoning. ...

来源:elvis· x.com