MambaCount: Efficient Text-guided Open-vocabulary Object Counting with Spatial Sparse State Space Duality Block 文章

ArXiv CS.CV2026-06-17NEWSen作者: Hao-Yuan Ma, Li Zhang, Minjie Qiang, Jie Gao

详细信息

来源站点: ArXiv CS.CV
作者: Hao-Yuan Ma, Li Zhang, Minjie Qiang, Jie Gao
文章类型: NEWS
语言: en
发布日期: 2026-06-17

摘要

arXiv:2606.17650v1 Announce Type: new Abstract: Text-guided Open-vocabulary Object Counting (TOOC) aims to estimate the number of objects described by text prompts, which is particularly challenging in dense scenes with large scale variations. Existing TOOC approaches predominantly rely on Transformers, whose quadratic complexity with respect to image resolution limits their scalability. Mamba offers a promising alternative due to its linear complexity. However, previous Mamba-based methods have two main limitations. On the one hand, the inherent causal formulation of Mamba constrains the bidirectional spatial dependency modeling required by non-causal vision tasks. On the other hand, existing Mamba-based vision models often overlook the unconstrained high entropy in the spatial token responses, which can weaken local details and high-frequency cues.

MambaCount: Efficient Text-guided Open-vocabulary Object Counting with Spatial Sparse State Space Duality Block 文章

详细信息

摘要

相关事件

相关公司

相关人物

相关产品

相关技术查看全部 (4)