AAAC: Activation-Aware Adaptive Codebooks for 4-bit LLM Weight Quantization 文章

ArXiv CS.CL2026-06-08NEWSen作者: Beshr IslamBouli, David Jin

详细信息

来源站点
ArXiv CS.CL
作者
Beshr IslamBouli, David Jin
文章类型
NEWS
语言
en
发布日期
2026-06-08

摘要

arXiv:2605.08692v2 Announce Type: replace-cross Abstract: Post-training weight-only quantization to 4 bits is widely used to reduce the memory and compute costs of large language model inference. Existing PTQ methods, such as AWQ and GPTQ, improve how weights are mapped onto a fixed 4-bit grid through scaling, clipping, or error compensation. To further improve accuracy, methods such as OmniQuant and QuIP\# uses gradient-assisted algorithms at the cost of hours of quantization time. In this work, we propose AAAC (Activation-Aware Adaptive Codebooks), a lightweight method for 4-bit LLM weight quantization. AAAC replaces the fixed scalar codebook used in standard quantization with two small learned scalar codebooks (64 bytes) per layer. Each group of weights selects the codebook that minimizes activation-weighted reconstruction error, encoding the choice in the unused sign bit of the group's positive scale and adding zero storage overhead.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据

相关产品

暂无数据