Rohan Paul@rohanpaul_ai
64AI 编辑部评分,满分 100
2026-08-06 01:54· 29分钟前
AI 导读

DeepSeek V4 Flash 在魔方与国际象棋双任务中消耗 27.4M tokens、尝试 5 次,却以 $0.557 完成,为四模型中最低价。GPT-5.6 Sol 效率最高(6.7M tokens、16m43s),但象棋 viewer 有缺陷。Paul 认为,推理足够便宜时,token 级低效仍可保持经济实用性,这改变了智能体效率的衡量方式。

4 frontier models built their own chess boards and all lost to Claude Opus 5.

Really Interesting experiments by @thehypedotnews, a 24/7 AI news in a really nice radio format.

In this experiment, I find DeepSeek V4 Flash's performance really interesting.

It used 27.4 mn tokens, needed 5 attempts to build working stands, and still completed both tasks for $0.557.

So kind of changes how agent efficiency should be measured. A model can reason inefficiently at the token level and still remain economically useful if inference is cheap enough to make retries almost free.

thehype.qwen 3.8 max vs deepseek v4 flash 0731 vs kimi k3 vs gpt 5.6 sol - on rubik's cube and chess four frontier models built a rubik's cube stand and solved it, then...

来源:Rohan Paul · x.com

Rohan Paul · @rohanpaul_ai · X·2026-08-06 01:54·29分钟前
AI 导读

DeepSeek V4 Flash 在魔方与国际象棋双任务中消耗 27.4M tokens、尝试 5 次,却以 $0.557 完成,为四模型中最低价。GPT-5.6 Sol 效率最高(6.7M tokens、16m43s),但象棋 viewer 有缺陷。Paul 认为,推理足够便宜时,token 级低效仍可保持经济实用性,这改变了智能体效率的衡量方式。

4 frontier models built their own chess boards and all lost to Claude Opus 5.

Really Interesting experiments by @thehypedotnews, a 24/7 AI news in a really nice radio format.

In this experiment, I find DeepSeek V4 Flash's performance really interesting.

It used 27.4 mn tokens, needed 5 attempts to build working stands, and still completed both tasks for $0.557.

So kind of changes how agent efficiency should be measured. A model can reason inefficiently at the token level and still remain economically useful if inference is cheap enough to make retries almost free.

thehype.qwen 3.8 max vs deepseek v4 flash 0731 vs kimi k3 vs gpt 5.6 sol - on rubik's cube and chess four frontier models built a rubik's cube stand and solved it, then...

来源:Rohan Paul· x.com