Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization 文章

ArXiv CS.CL2026-08-14PAPERen作者: Palaash Goel, Ayan Sengupta, Akshay Nambi, Tanmoy Chakraborty

详细信息

来源站点
ArXiv CS.CL
作者
Palaash Goel, Ayan Sengupta, Akshay Nambi, Tanmoy Chakraborty
文章类型
PAPER
语言
en
发布日期
2026-08-14

摘要

arXiv:2608.12953v1 Announce Type: new Abstract: Structured pruning is a promising approach for compressing large language models (LLMs), yet existing methods rely heavily on greedy heuristics that produce myopic decisions, and often fail to precisely meet target compression budgets. We present SNIPER, a two-stage structured pruning framework that solves a knapsack optimization over coarse-granularity components to yield conditionally optimal parameter allocations with respect to fixed importance estimates, followed by a fine-grained pruning stage to meet strict budget constraints. We introduce the Compression Ratio Adherence Factor (CRAFT) to quantify budget fidelity, showing that while existing pruners deviate from target compression ratios by up to 33%, SNIPER achieves near-exact adherence with a CRAFT score of 0.98.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据