Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM 文章

ArXiv CS.CL2026-06-30PAPERen作者: Shuvendu Roy, Mengyao Zhai, Hossein Hajimirsadeghi, Golnoosh Samei

详细信息

来源站点
ArXiv CS.CL
作者
Shuvendu Roy, Mengyao Zhai, Hossein Hajimirsadeghi, Golnoosh Samei
文章类型
PAPER
语言
en
发布日期
2026-06-30

摘要

arXiv:2606.29563v1 Announce Type: new Abstract: Large language models (LLMs) excel at complex tasks like question answering and summarization, thanks to their ability to handle long-context inputs. However, deploying LLMs is costly, not only due to the high computational demands of quadratic complexity of self-attention and auto-regressive generation, but also because of the significant memory overhead required for storing the key-value (KV) cache during inference. To reduce the memory cost, existing KV-cache eviction strategies leverage the sparsity in attention to selectively store a subset of tokens. While reducing the memory footprint, such approaches show a considerable drop in performance, especially in tasks that require long-context reasoning. We identify that the drop in performance is linked to a reduction in the coverage of unique tokens.

相关事件

暂无数据

相关公司查看全部 (2)

A
AT TCOMPANY

相关人物

暂无数据