Fine-grained Computation-Communication Overlap via Tile-level Signaling and Scheduling for Mixture-of-Experts 文章

ArXiv CS.AI2026-07-23PAPERen作者: Minyu Cui, Anna Wingkvist, Morgan Ericsson

详细信息

来源站点
ArXiv CS.AI
作者
Minyu Cui, Anna Wingkvist, Morgan Ericsson
文章类型
PAPER
语言
en
发布日期
2026-07-23

摘要

arXiv:2607.19539v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures increase model capacity without proportionally increasing computation cost and have become a key building block for scaling large language models (LLMs) to trillion-parameter regimes. Efficient deployment of these MoE models relies on distributed execution across multiple GPUs, where each MoE layer involves two all-to-all communications: dispatching tokens to expert ranks and returning the expert outputs to their source ranks. Conventional MoE implementations launch this return all-to-all after expert compute completes, exposing communication latency on the critical path and reducing GPU utilization. We present a fine-grained approach that overlaps expert compute with the second all-to-all via tile-level signaling and scheduling.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据