CodeShrink: Adaptive Visual Compression for Efficient Multimodal Code Understanding 文章

ArXiv CS.CV2026-08-03PAPERen作者: Wenxin Tang, Jingyu Xiao, Zhenyu Liu, Zipeng Xie, Junliang Liu, Wang Luo, Yuan Jiang, Yintong Huo, Michael Lyu

详细信息

来源站点
ArXiv CS.CV
作者
Wenxin Tang, Jingyu Xiao, Zhenyu Liu, Zipeng Xie, Junliang Liu, Wang Luo, Yuan Jiang, Yintong Huo, Michael Lyu
文章类型
PAPER
语言
en
发布日期
2026-08-03

摘要

arXiv:2607.29637v1 Announce Type: new Abstract: Rendering source code as images offers a promising way to reduce the input costs of Multimodal Large Language Models (MLLMs). Adjusting image resolution can trade visual token cost against content fidelity. However, resolution scaling alone overlooks two sources of inefficiency: blank regions created by line breaks and indentation, and code regions irrelevant to the current instruction. Moreover, the best compression setting varies across inputs, tasks, and models, limiting fixed-ratio strategies. We propose CodeShrink, an adaptive visual compression framework with three components. Blank-Free Rendering replaces whitespace-dependent layouts with compact layouts and explicit structural markers, removing layout-induced tokens. Adaptive Compression Configuration uses a lightweight agent trained with reinforcement learning to predict a per-input setting that balances token efficiency and readability.

相关事件

暂无数据

相关公司查看全部 (3)

A
ANDINONPROFIT
A
AdjustCOMPANY

相关人物

暂无数据