Joint Structural Pruning and Mixed-Precision Quantization for LLM Compression 文章

ArXiv CS.AI2026-06-09NEWSen作者: Hoang-Loc La, Truong-Thanh Le, Amir Taherkordi, Phuong Hoai Ha

详细信息

来源站点
ArXiv CS.AI
作者
Hoang-Loc La, Truong-Thanh Le, Amir Taherkordi, Phuong Hoai Ha
文章类型
NEWS
语言
en
发布日期
2026-06-09

摘要

arXiv:2606.07819v1 Announce Type: new Abstract: Recently, the efficiency of Large Language Models (LLMs) deployment has become a critical concern in practical applications. While post-training quantization (PTQ) and structural pruning are established techniques for reducing memory footprint and inference latency, most existing PTQ approaches optimize quantization errors on a per-layer basis, overlooking how errors accumulate and propagate through the network, often resulting in suboptimal solutions. Traditional pipelines also tend to apply pruning and quantization in isolation or sequentially, further compounding sub-optimality. We introduce a novel end-to-end framework that addresses these limitations in two key ways. First, we propose a novel mixed-precision PTQ strategy that directly minimizes global error propagation across the entire model, rather than isolating layer-wise errors.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据

相关产品

暂无数据