HuggingFace Daily Papers(社区热门论文)
48AI 编辑部评分,满分 100

超越几何互补性:稀疏MoE路由中的相干重叠

2026-07-30 08:00· 1天前
跳到正文
AI 摘要

研究提出专家子空间分离指数等指标,区分路由相干性、候选质量与候选-上下文交互。在六种MoE架构中发现专家子空间高度重叠,但实际路由对token表征的解释力优于匹配替代方案;在OLMoE、Mixtral和DeepSeek的39个因子单元中,选中专家解释力均强于最强未选中候选,但前缀会缩小这一优势。几何重叠不意味功能冗余,24项冻结路由比较显示增加专家可改善预测。

Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often conflates route coherence, candidate quality, and candidate-by-context interaction. We distinguish these quantities using an Expert Subspace Separation Index (ESSI), matched-route residuals, and a prefix-controlled $2\times2$ factorial; frozen-route interventions and a controlled Top-$k$ study assess functional value. Three paired contrasts organize the findings. First, across six MoE architectures, expert subspaces overlap substantially, yet actual routes explain token representations better than matched alternatives. Second, across the 39 factorial cells in OLMoE, Mixtral, and DeepSeek, the selected candidate explains more of the residual representation than the strongest unselected rival in every cell, yet the actual prefix narrows this advantage throughout: all interactions are negative, and every 95% confidence interval lies below zero. Third, this geometric narrowing does not imply functional redundancy: adding later experts improves next-token prediction in 24 of 39 frozen-route comparisons, while the other 15 estimates are inconclusive; a controlled training study also favors Top-2 over Top-1 in all three seeds. We call this joint pattern coherent overlap: routing selects token-relevant experts from a shared geometric neighborhood, while useful multi-expert computation persists without disjoint linear coverage. Separating these quantities clarifies why geometric similarity alone cannot determine redundancy or pruning value.

超越几何互补性:稀疏MoE路由中的相干重叠

HuggingFace Daily Papers(社区热门论文)·2026-07-30 08:00·1天前
阅读原文· arxiv.org
AI 摘要

研究提出专家子空间分离指数等指标,区分路由相干性、候选质量与候选-上下文交互。在六种MoE架构中发现专家子空间高度重叠,但实际路由对token表征的解释力优于匹配替代方案;在OLMoE、Mixtral和DeepSeek的39个因子单元中,选中专家解释力均强于最强未选中候选,但前缀会缩小这一优势。几何重叠不意味功能冗余,24项冻结路由比较显示增加专家可改善预测。

原文 · 保持原样,未翻译

Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often conflates route coherence, candidate quality, and candidate-by-context interaction. We distinguish these quantities using an Expert Subspace Separation Index (ESSI), matched-route residuals, and a prefix-controlled $2\times2$ factorial; frozen-route interventions and a controlled Top-$k$ study assess functional value. Three paired contrasts organize the findings. First, across six MoE architectures, expert subspaces overlap substantially, yet actual routes explain token representations better than matched alternatives. Second, across the 39 factorial cells in OLMoE, Mixtral, and DeepSeek, the selected candidate explains more of the residual representation than the strongest unselected rival in every cell, yet the actual prefix narrows this advantage throughout: all interactions are negative, and every 95% confidence interval lies below zero. Third, this geometric narrowing does not imply functional redundancy: adding later experts improves next-token prediction in 24 of 39 frozen-route comparisons, while the other 15 estimates are inconclusive; a controlled training study also favors Top-2 over Top-1 in all three seeds. We call this joint pattern coherent overlap: routing selects token-relevant experts from a shared geometric neighborhood, while useful multi-expert computation persists without disjoint linear coverage. Separating these quantities clarifies why geometric similarity alone cannot determine redundancy or pruning value.

阅读原文arxiv.org