# 中国10T参数MoE模型预训练或需约3万块Blackwell GPU

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-08-08 20:51
- AIHOT 分数：51
- AIHOT 链接：https://aihot.virxact.com/items/cmske8etv0az0row95c9u9foq
- 原文链接：https://x.com/rohanpaul_ai/status/2086072724552810954

## AI 摘要

Rohan Paul估算，中国传闻中10T总参数MoE模型预训练约需3万块Blackwell GPU，但参数量只是内存账单而非算力账单。前沿模型稀疏度持续下降，10T规模下激活参数约200–500B。字节跳动已通过马来西亚云运营商租用3.6万块Blackwell GPU（500个GB200机架，$2.5B），预训练窗口为3–6个月。

## 正文

~30K Blackwells sounds like a credible estimate for the rumored 10T-total-param MoE model from China.

but parameter count isn't what sets it.

MoE training cost (FLOPs) ≈ 6 × active params × tokens.

10T total is a memory bill, not a compute bill.

Frontier sparsity is also collapsing:

DeepSeek-V3 671B/37B = 5.5%
-> V4-Pro 1.6T/49B = 3.1%
-> Kimi K3 routes 16 of 896 experts.

At 10T that's ~200-500B active.

Over 15-40T tokens -> 2e25 - 1.2e26 FLOP.

ByteDance already rents 36,000 Blackwell GPUs - 500 GB200 racks, $2.5 B. (through Malaysia cloud operator , deal is in line with US export controls).

And Labs size clusters to the run they're planning.

And FT reports ByteDance's pretraining window at 3-6 months.

### 引用推文

> Zijing Wu：Was told it takes only about 30K GPUs to pre train a mega model. Inference compute is the real monster. But then most likely they'll just serve distilled smalle...
