DeepSeek V4 ships with two variants allowing it to hit 1M Context. @kimbochen breaks down the attention changes and the Mega MOE speeding up compute.
AI 导读
DeepSeek V4 发布两个变体,实现百万上下文。@kimbochen 详解注意力机制变化和 Mega MOE 加速计算。
67
AI 编辑部评分,满分 100DeepSeek V4 发布两个变体,实现百万上下文。@kimbochen 详解注意力机制变化和 Mega MOE 加速计算。
来源:SemiAnalysis· x.com