Linear Scaling Video VLMs for Long Video Understanding 事件
PRODUCT_LAUNCH2026-06-01影响: MEDIUM
Linear Scaling Video VLMs for Long Video Understanding arXiv:2605.31598v1 Announce Type: new Abstract: Video vision-language models (VLMs) are increasingly used in long-horizon and streaming settings, yet most video encoders still rely on spatiotemporal self-attention, causing compute and latency to grow quadratically with the number of frames. Existing efficiency methods improve scalability but often lose accuracy relative to full self-attention, for example through aggressive frame/token drop
相关产品查看全部 (10)
相关报道查看全部 (1)
Linear Scaling Video VLMs for Long Video Understanding
ArXiv CS.CV2026-06-01