Deep Multimodal Fusion Detection through Spatial Mask and Channel Fusion 文章

ArXiv CS.CV2026-08-04PAPERen作者: Guandi Wang, Ming Li, Yunsen Xing, Junle Liu

详细信息

来源站点
ArXiv CS.CV
作者
Guandi Wang, Ming Li, Yunsen Xing, Junle Liu
文章类型
PAPER
语言
en
发布日期
2026-08-04

摘要

arXiv:2608.02092v1 Announce Type: new Abstract: Deep multimodal fusion for object detection has demonstrated good performance through mining modal characteristics. However, existing feature-level fusion methods mainly weigh between two modalities and unify them in a unified representation space. This can lead to overfitting or over-specialization of the statistical properties of a single modality within a dual-backbone architecture. This paper proposes an Attention-Driven Complementarity Resampling framework for robust improvement of cross-modality object detection. Based on a shared channel spatial attention mechanism, we first introduce the semantic mask exchange to actively mix the boundaries of the modalities during the training phase, forcing the backbone network to learn generalized features without relying on fixed modal labels. Then we propose a learnable channel competition to sample and aggregate features in a channel-wise and learnable way.