Diagnosing Overhead in Dispatch Operations: Cross-architecture Observatory 文章

ArXiv CS.AI2026-07-20PAPERen作者: Bole Ma, Jan Eitzinger, Harald Koestler, Gerhard Wellein

详细信息

来源站点
ArXiv CS.AI
作者
Bole Ma, Jan Eitzinger, Harald Koestler, Gerhard Wellein
文章类型
PAPER
语言
en
发布日期
2026-07-20

摘要

arXiv:2605.20982v2 Announce Type: replace-cross Abstract: AlltoAll dispatch is the dominant bottleneck of MoE expert parallelism, and the interconnect community has responded with four families of mitigations: predictive sample placement, adaptive expert relayout, hierarchical collectives, and EP-aware topology. All four rest on two assumptions about the workload: that routing imbalance is correctable by the system layer, and that the mock-token benchmarks evaluating them faithfully represent production routing. We introduce DODOCO to test both, instrumenting five open MoE checkpoints that span today's sequence-mixer designs (MHA, MLA, GQA, Gated DeltaNet and Mamba-2 SSM) under a factorial grid of six data conditions and a matched expert-parallelism scan on H100 clusters. Both assumptions fail. Scaling EP leaves per-expert load concentration essentially unchanged: the straggler is intrinsic to the routing decision the model makes, not to how its experts land on ranks.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据