# DeepSeek V4 Flash 用 27.4M tokens 完成双任务，成本仅 $0.557

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-08-06 01:54
- AIHOT 分数：64
- AIHOT 链接：https://aihot.virxact.com/items/cmsgeiym703h9ro5qrjf3wx5y
- 原文链接：https://x.com/rohanpaul_ai/status/2085061823108464946

## AI 摘要

DeepSeek V4 Flash 在魔方与国际象棋双任务中消耗 27.4M tokens、尝试 5 次，却以 $0.557 完成，为四模型中最低价。GPT-5.6 Sol 效率最高（6.7M tokens、16m43s），但象棋 viewer 有缺陷。Paul 认为，推理足够便宜时，token 级低效仍可保持经济实用性，这改变了智能体效率的衡量方式。

## 正文

4 frontier models built their own chess boards and all lost to Claude Opus 5.

Really Interesting experiments by @thehypedotnews, a 24/7 AI news in a really nice radio format.

In this experiment, I find DeepSeek V4 Flash's performance really interesting.

It used 27.4 mn tokens, needed 5 attempts to build working stands, and still completed both tasks for $0.557.

So kind of changes how agent efficiency should be measured. A model can reason inefficiently at the token level and still remain economically useful if inference is cheap enough to make retries almost free.

### 引用推文

> thehype.：qwen 3.8 max vs deepseek v4 flash 0731 vs kimi k3 vs gpt 5.6 sol - on rubik's cube and chess four frontier models built a rubik's cube stand and solved it, then...
