Omni-Prune: Query-Aware Unified Token Pruning for Efficient Omnimodal Large Language Models 文章

ArXiv CS.CV2026-07-28PAPERen作者: Yiming Zhong, Chang Nie, Caifeng Shan

详细信息

来源站点
ArXiv CS.CV
作者
Yiming Zhong, Chang Nie, Caifeng Shan
文章类型
PAPER
语言
en
发布日期
2026-07-28

摘要

arXiv:2607.23445v1 Announce Type: new Abstract: Omnimodal large language models (OmniLLMs) are rapidly extending multimodal reasoning to cover synchronized audio and video. However, the resulting audio-video token sequences are long, leading to high prefill latency and GPU memory usage at inference time. Existing token pruning methods, designed mainly for vision-only inputs, miss both the cross-modal links between audio and video and the user query that decides which content matters. To bridge this gap, we present Omni-Prune, a training-free, query-aware audio-visual token pruning framework that jointly removes redundancy from both modalities while keeping task-relevant cross-modal evidence. Specifically, Omni-Prune first splits the token sequence into adaptive time windows placed at audio saliency peaks, then scores audio and video tokens on a single scale that combines encoder attention with text-query relevance, and pairs related audio-video tokens so that they are kept together.

相关事件

暂无数据

相关公司查看全部 (3)

A
AT TCOMPANY

相关人物

暂无数据