Models as Tools: An Agentic Coordination Framework for Unified Multimodal Visual Tracking 文章

ArXiv CS.CV2026-08-04PAPERen作者: Wenrui Cai, Yuzhe Li, Qingjie Liu, Yunhong Wang

详细信息

来源站点
ArXiv CS.CV
作者
Wenrui Cai, Yuzhe Li, Qingjie Liu, Yunhong Wang
文章类型
PAPER
语言
en
发布日期
2026-08-04

摘要

arXiv:2608.00847v1 Announce Type: new Abstract: Most current visual trackers adopt a matching-based architecture trained exclusively on tracking datasets, whose performance gains depend heavily on the length of the input context, and have now reached a bottleneck. While high-performance tracking increasingly relies on foundation models, existing methods use them monolithically, adapting a foundation model into a tracker or modify a segmentation foundation model into a tracking pipeline, which fails to exploit complementary strengths. Matching-based trackers excel at instance-level correspondence but lack semantic discrimination and fine-grained foreground perception, whereas segmentation foundation models produce precise masks yet struggle with instance discrimination and multimodal extension. Both paradigms also lack error-correction capabilities for long-term tracking.