~30K Blackwells sounds like a credible estimate for the rumored 10T-total-param MoE model from China.
but parameter count isn't what sets it.
MoE training cost (FLOPs) ≈ 6 × active params × tokens.
10T total is a memory bill, not a compute bill.
Frontier sparsity is also collapsing:
DeepSeek-V3 671B/37B = 5.5% -> V4-Pro 1.6T/49B = 3.1% -> Kimi K3 routes 16 of 896 experts.
At 10T that's ~200-500B active.
Over 15-40T tokens -> 2e25 - 1.2e26 FLOP.
ByteDance already rents 36,000 Blackwell GPUs - 500 GB200 racks, $2.5 B. (through Malaysia cloud operator , deal is in line with US export controls).
And Labs size clusters to the run they're planning.
And FT reports ByteDance's pretraining window at 3-6 months.