Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts 文章

ArXiv CS.CV2026-08-04PAPERen作者: Ziang Wu, Peng Jin, Qishen Yin, Munan Ning, Hao Li, Peizhen Zhang, Li Yuan

详细信息

来源站点
ArXiv CS.CV
作者
Ziang Wu, Peng Jin, Qishen Yin, Munan Ning, Hao Li, Peizhen Zhang, Li Yuan
文章类型
PAPER
语言
en
发布日期
2026-08-04

摘要

arXiv:2608.00574v1 Announce Type: new Abstract: Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this token mix. We call the standard token-level Switch auxiliary loss Std-Aux. Std-Aux balances only the mixed load, so large image and text load errors can cancel at one mix. On our main model, the same trained router shows more than a fivefold change in load imbalance across image resolutions. We hold the image and text load profiles fixed and derive the exact load curve as the token mix varies. The image-text load gap controls sensitivity to the token mix. Physical preprocessing can also change the conditional profiles. The fixed-profile law excludes such changes. To design a remedy, we examine the router input structure. Image and text occupy distinct regions, while visual tokens group strongly by source image. The modality boundary motivates separate image and text terms.

相关事件

暂无数据

相关公司查看全部 (1)

A
AMI团队RESEARCH_INSTITUTE

相关人物

暂无数据