详细信息
- 来源站点
- ArXiv CS.CL
- 作者
- Renuka Oladri, Mohan Vamsi Varadaraju Priya, Jerry Wu
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-07-14
摘要
arXiv:2607.09999v1 Announce Type: new Abstract: We show that post-training quantization can silently alter how large language models reason even when task accuracy is preserved. Using a six-category failure taxonomy validated by two independent human annotators (Cohen's $\kappa$ = 0.906), we classify 30,000 chain-of-thought outputs from five instruction-tuned LLMs (3B--14B parameters) across three quantization precisions (FP32, FP16, NF4) and four reasoning benchmarks. We find that while accuracy is robust across precisions (maximum 3.1 pp drop), Hollow Convergence (correct answers reached through incomplete or unverifiable reasoning) shows a significant size-dependent shift under NF4, dropping sharply for the two smallest models tested but remaining invariant for models at 12B parameters and above. This effect is also benchmark-specific: GSM8K is categorically immune while LogiQA and ARC-Challenge show the largest shifts.
相关事件
暂无数据
相关公司
暂无数据
相关人物
暂无数据