详细信息
- 来源站点
- ArXiv CS.AI
- 作者
- Bole Ma, Jan Eitzinger, Harald Koestler, Gerhard Wellein
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-07-20
摘要
arXiv:2605.20982v2 Announce Type: replace-cross Abstract: AlltoAll dispatch is the dominant bottleneck of MoE expert parallelism, and the interconnect community has responded with four families of mitigations: predictive sample placement, adaptive expert relayout, hierarchical collectives, and EP-aware topology. All four rest on two assumptions about the workload: that routing imbalance is correctable by the system layer, and that the mock-token benchmarks evaluating them faithfully represent production routing. We introduce DODOCO to test both, instrumenting five open MoE checkpoints that span today's sequence-mixer designs (MHA, MLA, GQA, Gated DeltaNet and Mamba-2 SSM) under a factorial grid of six data conditions and a matched expert-parallelism scan on H100 clusters. Both assumptions fail. Scaling EP leaves per-expert load concentration essentially unchanged: the straggler is intrinsic to the routing decision the model makes, not to how its experts land on ranks.
相关事件
暂无数据
相关公司
暂无数据
相关人物
暂无数据