Rohan Paul@rohanpaul_ai
51AI 编辑部评分,满分 100
2026-08-08 20:51· 41分钟前
AI 导读

Rohan Paul估算,中国传闻中10T总参数MoE模型预训练约需3万块Blackwell GPU,但参数量只是内存账单而非算力账单。前沿模型稀疏度持续下降,10T规模下激活参数约200–500B。字节跳动已通过马来西亚云运营商租用3.6万块Blackwell GPU(500个GB200机架,$2.5B),预训练窗口为3–6个月。

~30K Blackwells sounds like a credible estimate for the rumored 10T-total-param MoE model from China.

but parameter count isn't what sets it.

MoE training cost (FLOPs) ≈ 6 × active params × tokens.

10T total is a memory bill, not a compute bill.

Frontier sparsity is also collapsing:

DeepSeek-V3 671B/37B = 5.5% -> V4-Pro 1.6T/49B = 3.1% -> Kimi K3 routes 16 of 896 experts.

At 10T that's ~200-500B active.

Over 15-40T tokens -> 2e25 - 1.2e26 FLOP.

ByteDance already rents 36,000 Blackwell GPUs - 500 GB200 racks, $2.5 B. (through Malaysia cloud operator , deal is in line with US export controls).

And Labs size clusters to the run they're planning.

And FT reports ByteDance's pretraining window at 3-6 months.

Zijing WuWas told it takes only about 30K GPUs to pre train a mega model. Inference compute is the real monster. But then most likely they'll just serve distilled smalle...

来源:Rohan Paul · x.com

Rohan Paul · @rohanpaul_ai · X·2026-08-08 20:51·41分钟前
AI 导读

Rohan Paul估算,中国传闻中10T总参数MoE模型预训练约需3万块Blackwell GPU,但参数量只是内存账单而非算力账单。前沿模型稀疏度持续下降,10T规模下激活参数约200–500B。字节跳动已通过马来西亚云运营商租用3.6万块Blackwell GPU(500个GB200机架,$2.5B),预训练窗口为3–6个月。

~30K Blackwells sounds like a credible estimate for the rumored 10T-total-param MoE model from China.

but parameter count isn't what sets it.

MoE training cost (FLOPs) ≈ 6 × active params × tokens.

10T total is a memory bill, not a compute bill.

Frontier sparsity is also collapsing:

DeepSeek-V3 671B/37B = 5.5% -> V4-Pro 1.6T/49B = 3.1% -> Kimi K3 routes 16 of 896 experts.

At 10T that's ~200-500B active.

Over 15-40T tokens -> 2e25 - 1.2e26 FLOP.

ByteDance already rents 36,000 Blackwell GPUs - 500 GB200 racks, $2.5 B. (through Malaysia cloud operator , deal is in line with US export controls).

And Labs size clusters to the run they're planning.

And FT reports ByteDance's pretraining window at 3-6 months.

Zijing WuWas told it takes only about 30K GPUs to pre train a mega model. Inference compute is the real monster. But then most likely they'll just serve distilled smalle...

来源:Rohan Paul· x.com