REVIEW 2 major objections 5 minor 77 references
Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization
T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Structured pruning of large language models can be cast as a 0/1 knapsack problem, and the paper claims dynamic programming with a fine-grained width stage achieves near-exact budget adherence and better task-level stability than greedy…
desk verdict Useful two-stage knapsack-based pruner with broad empirical evaluation; the 'conditional optimality' claim needs tightening because the discretized DP can violate the stated budget and the importance loop is under-specified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a 0/1 knapsack instance built from a mixed-granularity decomposition of a transformer: each layer contributes three mutually exclusive items — attention block, MLP block, or whole layer — grouped so the solver picks at most one branch per layer. The dynamic programming recurrence $f(i,j)$ over the first $i$ items and capacity $j$ finds the maximum-value subset, and a discretization factor $\alpha$ reduces the capacity dimension to keep runtime tractable. A second mechanism, the sensitivity score $\Omega_j = |D_{j,:} \cdot (G_{:,j} \odot U_{:,j})|$ for each MLP column, tells the fine-grained stage which columns to prune after the coarse stage has left a residual capacity. Together these two mechanisms turn pruning into a global, budget-constrained selection problem followed by a precision fill step.
What would settle it
Compare SNIPER's selected component set against a greedy depth pruner that uses the same iterative marginal importance scores on the same model and budget; if the greedy set matches or beats SNIPER's retention, the reported advantage is carried by the scores rather than by the knapsack solver, and the conditional-optimality claim would not be what drives the result.
Extended reading notes
Core claim
The paper's central claim is that the depth-pruning step of an LLM pruner can be set up as Equation (1): choose the subset $S^* \subseteq X$ of coarse components that maximizes total value $\sum_{x \in S} v(x)$ while keeping total parameter weight $\sum_{x \in S} w(x)$ at or below budget $C$. The claimed advance is not global optimality — the authors state plainly that the solution is optimal only with respect to fixed importance estimates — but it is a guarantee that greedy methods lack. The importance estimates themselves are computed iteratively as the marginal logit divergence of removing each component from a partially pruned model, so the scores are meant to reflect joint effects rather than isolated component quality. The follow-up width stage distributes the remaining budget across MLP layers in inverse proportion to their importance and removes the least sensitive columns, which is how the method achieves near-exact budget adherence while preserving tensor structure. The paper also introduces CRAFT, the ratio of observed to target compression, as a way to quantify this budget fidelity.
Load-bearing premise
The knapsack's optimality is conditional on the importance scores being correct, and those scores are heuristic and path-dependent; if the scores fail to capture true performance loss, the optimal selection is optimal for the wrong objective.
Editorial extensions
If this is right
- If the central claim holds, structured pruning becomes budget-exact: a practitioner who asks for a 25% or 35% model gets a model that actually has that parameter count, removing the capacity slack the paper documents in existing pruners.
- The conditional-optimality guarantee means that, for the importance scores SNIPER computes, no other subset of the same coarse components within the budget has higher total value; greedy depth pruners offer no comparable guarantee.
- Because the importance scores transfer across compression ratios with only a small retention loss, the expensive scoring pass can be amortized over multiple target budgets rather than recomputed from scratch.
- The framework's performance on dense, reasoning-tuned, fused-MLP, and mixture-of-experts architectures suggests the knapsack formulation is not bound to one transformer design.
- Lower task-wise standard deviation across 18 tasks implies the method is less likely to over-optimize one metric at the expense of others, addressing a documented failure mode of greedy pruning.
Reading between the lines
- Beyond the paper: the observed transferability of importance scores across budgets suggests the same scores could be reused across checkpoints of a model family, making the dominant scoring cost a one-time investment per family.
- Beyond the paper: because attention and MLP blocks enter the knapsack as exchangeable items, the same solver could be pointed at expert-level components in mixture-of-experts models or at non-uniform per-layer width budgets without changing the core formulation.
- Beyond the paper: the strong overlap of pruned layers across budgets, which the paper presents as interpretable salience, could be read as a diagnostic tool for architectural redundancy that is independent of any particular compression target.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SNIPER, a two-stage structured pruning framework for LLMs. Stage 1 solves a 0/1 knapsack problem over coarse components (attention blocks, MLP blocks, and whole layers) using dynamic programming, with component importance estimated from logit divergence under iterative partial pruning. Stage 2 performs fine-grained MLP column pruning to fill the remaining parameter budget. The paper introduces the Compression Ratio Adherence Factor (CRAFT) as a metric for budget fidelity and evaluates SNIPER on four architectures, six baselines, and 18 tasks, reporting a mean rank of 1.25 and near-exact compression-ratio adherence.
Significance. If the central claims hold, SNIPER would be a practically useful alternative to greedy structured pruners: it offers a non-greedy selection mechanism, a simple budget-fidelity metric, and an unusually broad evaluation suite spanning diverse architectures and tasks. The DP recurrence is standard and the empirical comparison is extensive, including calibration sensitivity, importance-score transfer, and ablation studies. The conditional-optimality framing is honest in that it does not claim a global guarantee over all pruning configurations, and the paper explicitly concedes in Section 8 that no guarantee extends to importance estimation. However, the two concerns below affect load-bearing parts of the paper: feasibility of the claimed capacity constraint, and the circularity/under-specification of the iterative importance estimation loop.
major comments (2)
- [Section 3.2, Eq. (1); Table 8] The DP solution is not guaranteed to satisfy the capacity constraint of Eq. (1). The recurrence discretizes weights as floor(w(x)/alpha) and capacity as floor(C/alpha); rounding weights downward allows the reconstructed set to have true weight exceeding C. Table 8 shows this is not hypothetical: taking the compression ratio as the fraction of parameters removed, at target CR 25% SNIPER's observed CR is 24.62%, so the final model retains 75.38% of parameters while the budget allows 75%; since Stage 2 only removes parameters, Stage 1 must have already exceeded the budget. The same pattern occurs for every target CR in Table 8 (observed CR is below target in all five rows). Consequently, the statement that SNIPER yields a 'conditionally optimal' allocation with respect to Eq. (1) is not supported for the actual parameter budget; at best the solver is optimal for a discretized surrogate problem. The paper should either use a feasibility-preserving discretization (e.g., rounding weights up and capacity down), state the guarantee for the discretized problem and provide a feasible repair step, or explicitly re-define Eq. (1) in terms of the discretized weights.
- [Section 3.3 and Section 5.4] The importance scores v(x_i) in Eq. (3) are defined relative to a partially pruned model M^{(i-1)}, but Section 3.3 never specifies how M^{(i-1)} is constructed. Section 5.4 states that its pruning configuration is 'determined by our dynamic programming solver,' i.e., by the same S* that Eq. (1) computes from v. This makes the definition of v circular unless an explicit iterative fixed-point procedure (initialization, update rule, convergence criterion) is provided. Because the paper's optimality guarantee is explicitly conditional on fixed importance estimates, the missing specification means the guarantee does not currently apply to the actual algorithm used in the experiments. The ablation against leave-one-out scoring in Table 4 shows that the choice of scoring procedure materially changes results, so this is a central issue rather than a cosmetic one.
minor comments (5)
- [Section 3.2; Section 6.4] Cross-references are inconsistent: Section 3.2 cites 'Appendix 6.5' and Section 6.4 cites 'Appendix 6.2'; both should point to the corresponding sections (6.5 and 6.2, respectively) rather than to an appendix.
- [Abstract; Section 5.3; Table 8] The headline CRAFT value of 0.98 appears in the abstract, Section 5.3, and Figure 1, but Table 8's SNIPER rows average to approximately 0.97. The paper should reconcile the reported number or indicate explicitly that 0.98 refers to the values in Table 7 rather than Table 8.
- [Section 5.1; Table 1] The text states that at 35% compression on Qwen3-8B with RFT, SNIPER reduces task-wise variance to 15.90% 'compared to 17–30% for competing pruners,' but Table 1 reports 2SSP with Std RP 15.80% in the same configuration, which is lower than SNIPER's 15.90%.
- [Section 5; Table 1] The sentence 'Across all configurations except one, SNIPER achieves the highest average retention performance' is contradicted by the same paragraph: on LLaMA-3.1-8B-Instruct at 25% compression with RFT, both ReplaceMe (88.90) and ShortGPT (88.47) outperform SNIPER (87.34). The count should be corrected.
- [Appendix B] LAMBADA appears in both 'Generative Performance' and 'NLU and Inference' task lists. Please clarify whether it is evaluated once and categorized twice, or whether two different task formats (log-perplexity vs. next-word prediction accuracy) are used.
Circularity Check
Importance scores are defined using the DP solution they are supposed to determine, making the claimed conditional optimum self-referential.
-
self definitional
[Section 3.3 (Eq. 3) and Section 5.4 (Iterative Pruning for Importance Estimation)]
"To estimate component importance, SNIPER computes the marginal degradation in performance incurred by dropping a component from a partially pruned model, whose pruning configuration is determined by our dynamic programming solver (Section 3.3)."
In Eq. 3, v(x_i) is the logit divergence from removing x_i from M(i-1), where Section 3.3 defines M(i-1) as 'the model pruned up to component i-1.' Section 5.4 states that this partially pruned model's configuration is 'determined by our dynamic programming solver.' But the DP solver (Eqs. 1-2) takes v(x_i) as fixed inputs and produces the selected set S*. Thus the optimization's inputs are defined in terms of the optimization's own output: v(x_i) depends on which components the DP already pruned, while the DP's choice depends on v(x_i). The paper provides no external initialization or fixed-point iteration, so the claimed 'conditional optimum with respect to fixed importance estimates' is conditional on scores that are not fixed independent inputs.
full rationale
The knapsack recurrence (Eq. 2) genuinely optimizes the stated objective for fixed values v(x_i), and the 18-task evaluation is an external benchmark that does not reduce to the method's inputs. However, the paper's own ablation text makes the importance estimates load-bearing and self-referential: v(x_i) is computed from a partially pruned model whose configuration is 'determined by our dynamic programming solver,' while the solver's selection is in turn a function of v(x_i). As written, Eq. 3's input depends on Eq. 1's output, so the advertised 'conditional optimum with respect to fixed importance estimates' lacks a fixed, independent input; Section 8 explicitly declines to extend guarantees to importance estimation. Separately, the reported CRAFT is a designed property of Stage 2 (columns are pruned until the residual budget is filled), not an independent prediction, and the alpha-discretized DP can return sets whose true weight exceeds the budget (Table 8 at target 25% shows observed CR 24.62%, implying Stage 1 retained more than 75% of parameters), so feasibility of the claimed optimum is also not established. These issues are partly correctness concerns, but the self-referential scoring loop is a genuine circularity in the derivation of the main optimality claim.
Assumptions & free parameters
free parameters (3)
- Discretizing factor alpha =
32
- Softmax temperature tau =
1
- Calibration sample count =
50 samples
assumptions (6)
- standard math The 0/1 knapsack recurrence in Equation 2 with grouped mutually exclusive items returns the exact optimum for the stated objective in Equation 1.
- domain assumption Logit-divergence importance in Equation 3 is a faithful proxy for how removing a component affects downstream task performance.
- domain assumption Transformer components can be deterministically sequenced so the iterative marginal contributions in Section 3.3 are well-defined and meaningful.
- domain assumption The column sensitivity score in Equation 6 identifies the least salient MLP columns.
- domain assumption Baselines re-engineered for new architectures behave as their original implementations intended.
- domain assumption LoRA recovery fine-tuning with a single epoch restores all methods comparably.
invented entities (1)
-
Compression Ratio Adherence Factor (CRAFT)
Cite this review
Pith. "Pith review of Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization." pith.science (2026). https://pith.science/paper/INTJUYIC
@misc{pith2026260812953,
author = {Pith},
title = {Pith review of: Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization},
year = {2026},
howpublished = {\url{https://pith.science/paper/INTJUYIC}},
note = {Machine review of arXiv:2608.12953}
}
read the original abstract
Structured pruning is a promising approach for compressing large language models (LLMs), yet existing methods rely heavily on greedy heuristics that produce myopic decisions, and often fail to precisely meet target compression budgets. We present SNIPER, a two-stage structured pruning framework that solves a knapsack optimization over coarse-granularity components to yield conditionally optimal parameter allocations with respect to fixed importance estimates, followed by a fine-grained pruning stage to meet strict budget constraints. We introduce the Compression Ratio Adherence Factor (CRAFT) to quantify budget fidelity, showing that while existing pruners deviate from target compression ratios by up to 33%, SNIPER achieves near-exact adherence with a CRAFT score of 0.98. Evaluations across four diverse architectures over a set of 18 tasks spanning five domains demonstrate SNIPER's consistent improvements in average performance retention and task-level stability over six state-of-the-art pruners. Across all pruning configurations, SNIPER achieves an excellent mean rank of 1.25, indicating its robust cross-architectural generalizability and excellent reliability.
Reference graph
Works this paper leans on
-
[1]
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Y ang, A. Fan, A. Goyal, A. Hartshorn, A. Y ang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. M...
arXiv 2024
-
[2]
A. Y ang, A. Li, B. Y ang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Y ang, J. Tu, J. Zhang, J. Y ang, J. Y ang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Y ang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P . Zhang, P . Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. ...
arXiv 2025
-
[3]
DeepSeek-AI, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Guo, D. Y ang, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, J. L. Cai, J. Liang, J. Guo, J. Ni, J. Li, J. Wang, J. Chen, J. Chen, J. Yuan, ...
arXiv 2024
-
[4]
gpt-oss-120b & gpt-oss-20b model card,
OpenAI, :, S. Agarwal, L. Ahmad, J. Ai, S. Altman, A. Applebaum, E. Arbus, R. K. Arora, Y . Bai, B. Baker, H. Bao, B. Barak, A. Bennett, T. Bertao, N. Brett, E. Brevdo, G. Brockman, S. Bubeck, C. Chang, K. Chen, M. Chen, E. Cheung, A. Clark, D. Cook, M. Dukhan, C. Dvorak, K. Fives, V . Fomenko, T. Garipov, K. Georgiev, M. Glaese, T. Gogineni, A. Goucher, ...
arXiv 2025
-
[5]
A survey on model compression for large language models,
X. Zhu, J. Li, Y . Liu, C. Ma, and W. Wang, “A survey on model compression for large language models,” 2024. [Online]. Available: https://arxiv.org/abs/2308.07633
arXiv 2024
-
[6]
A survey of small language models,
C. V . Nguyen, X. Shen, R. Aponte, Y . Xia, S. Basu, Z. Hu, J. Chen, M. Parmar, S. Kunapuli, J. Barrow, J. Wu, A. Singh, Y . Wang, J. Gu, F. Dernoncourt, N. K. Ahmed, N. Lipka, R. Zhang, X. Chen, T. Yu, S. Kim, H. Deilamsalehy, N. Park, M. Rimer, Z. Zhang, H. Y ang, R. A. Rossi, and T. H. Nguyen, “A survey of small language models,” 2024. [Online]. Availa...
arXiv 2024
-
[7]
Efficient 8-bit quantization of transformer neural machine language translation model,
A. Bhandare, V . Sripathi, D. Karkada, V . Menon, S. Choi, K. Datta, and V . Saletore, “Efficient 8-bit quantization of transformer neural machine language translation model,” 2019. [Online]. Available: https://arxiv.org/abs/1906.00532
arXiv 2019
-
[8]
Qlora: efficient finetuning of quantized llms,
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “Qlora: efficient finetuning of quantized llms,” in Proceedings of the 37th International Conference on Neural Information Processing Systems , ser. NIPS ’23. Red Hook, NY, USA: Curran Associates Inc., 2023
work page 2023
Show all 77 references
-
[9]
CBQ: Cross-block quantization for large language models,
X. Ding, X. Liu, Z. Tu, Y . Zhang, W. Li, J. Hu, H. Chen, Y . Tang, Z. Xiong, B. Yin, and Y . Wang, “CBQ: Cross-block quantization for large language models,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview...
2025
-
[10]
Minillm: Knowledge distillation of large language models,
Y . Gu, L. Dong, F. Wei, and M. Huang, “Minillm: Knowledge distillation of large language models,” 2025. [Online]. Available: https://arxiv.org/abs/2306.08543
2025 arXiv
-
[11]
A good learner can teach better: Teacher-student collaborative knowledge distillation,
A. Sengupta, S. Dixit, M. S. Akhtar, and T. Chakraborty, “A good learner can teach better: Teacher-student collaborative knowledge distillation,” in Proceedings of the The Twelfth International Conference on Learning Representations, Virtual Event, 2023, pp. 25–29
2023
-
[12]
Replaceme: Network simplification via depth pruning and transformer block linearization,
D. Shopkhoev, A. Ali, M. Zhussip, V . Malykh, S. Lefkimmiatis, N. Komodakis, and S. Zagoruyko, “Replaceme: Network simplification via depth pruning and transformer block linearization,” 2025. [Online]. Available: https: //arxiv.org/abs/2505.02819
2025
-
[13]
Shortgpt: Layers in large language models are more redundant than you expect,
X. Men, M. Xu, Q. Zhang, B. Wang, H. Lin, Y . Lu, X. Han, and W. Chen, “Shortgpt: Layers in large language models are more redundant than you expect,” 2024. [Online]. Available: https://arxiv.org/abs/2403.03853
2024 arXiv
-
[14]
SLEB: Streamlining LLMs through redundancy verification and elimination of transformer blocks,
J. Song, K. Oh, T. Kim, H. Kim, Y . Kim, and J.-J. Kim, “SLEB: Streamlining LLMs through redundancy verification and elimination of transformer blocks,” in Proceedings of the 41st International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, R. S...
2024
-
[15]
The unreasonable ineffectiveness of the deeper layers,
A. Gromov, K. Tirumala, H. Shapourian, P . Glorioso, and D. Roberts, “The unreasonable ineffectiveness of the deeper layers,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=ngmEcEer8a
2025
-
[16]
Llm-pruner: on the structural pruning of large language models,
X. Ma, G. Fang, and X. Wang, “Llm-pruner: on the structural pruning of large language models,” in Proceedings of the 37th International Conference on Neural Information Processing Systems , ser. NIPS ’23. Red Hook, NY, USA: Curran Associates Inc., 2023
2023
-
[17]
Y ou only prune once: Designing calibration-free model compression with policy learning,
A. Sengupta, S. Chaudhary, and T. Chakraborty, “Y ou only prune once: Designing calibration-free model compression with policy learning,” in The Thirteenth International Conference on Learning Representations, 2025. [Online]. Available: https://openreview.net/forum?id=5RZoYIT3u6
2025
-
[18]
Slicegpt: Compress large language models by deleting rows and columns,
S. Ashkboos, M. L. Croci, M. G. do Nascimento, T. Hoefler, and J. Hensman, “Slicegpt: Compress large language models by deleting rows and columns,” 2024. [Online]. Available: https://arxiv.org/abs/2401.15024
2024 arXiv
-
[19]
Shortened llama: Depth pruning for large language models with comparison of retraining methods,
B.-K. Kim, G. Kim, T.-H. Kim, T. Castells, S. Choi, J. Shin, and H.-K. Song, “Shortened llama: Depth pruning for large language models with comparison of retraining methods,” 2024. [Online]. Available: https://arxiv.org/abs/2402.02834 16 P . Goel, A. Sengupta, A. Nambi and T. ...
2024 arXiv
-
[20]
Sliding-window merging for compacting patch-redundant layers in llms,
X. Ding, R. Sun, Y . Zhang, X. Y an, Y . Zhou, K. Huang, S. Fu, A. I. Aviles-Rivero, C. Xie, and Y . Zhu, “Sliding-window merging for compacting patch-redundant layers in llms,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 40, no. 25, p. 20826–20834, Ma...
2026 doi
-
[21]
Beware of calibration data for pruning large language models,
Y . Ji, Y . Xiang, J. Li, Q. Xia, P . Li, X. Duan, Z. Wang, and M. Zhang, “Beware of calibration data for pruning large language models,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=x83w6yGIWb
2025
-
[22]
I-bert: Integer-only bert quantization,
S. Kim, A. Gholami, Z. Y ao, M. W. Mahoney, and K. Keutzer, “I-bert: Integer-only bert quantization,” in Proceedings of the 38th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul...
2021
-
[23]
Llm-fp4: 4-bit floating-point quantized transformers,
S.-y. Liu, Z. Liu, X. Huang, P . Dong, and K.-T. Cheng, “Llm-fp4: 4-bit floating-point quantized transformers,” in Pro- ceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2023, p. 592–605. [Online]. A...
2023 doi
-
[24]
Aptq: Attention-aware post-training mixed-precision quantization for large language models,
Z. Guan, H. Huang, Y . Su, H. Huang, N. Wong, and H. Yu, “Aptq: Attention-aware post-training mixed-precision quantization for large language models,” in Proceedings of the 61st ACM/IEEE Design Automation Conference , ser. DAC ’24. ACM, Jun. 2024, p. 1–6. [Online]. Available: ...
2024
-
[25]
Distilling the knowledge in a neural network,
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” 2015. [Online]. Available: https://arxiv.org/abs/1503.02531
2015 arXiv
-
[26]
TinyBERT: Distilling BERT for natural language understanding,
X. Jiao, Y . Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu, “TinyBERT: Distilling BERT for natural language understanding,” in Findings of the Association for Computational Linguistics: EMNLP 2020 , T. Cohn, Y . He, and Y . Liu, Eds. Online: Association for Comp...
2020
-
[27]
Meta-learned modality-weighted knowledge distillation for robust multi-modal learning with missing data,
H. Wang, S. Hassan, Y . Liu, C. Ma, Y . Chen, Q. Li, J. Geng, B. Wang, Y . Tian, Y . Xie, J. Avery, L. Hull, I. Reid, M. Y aqub, and G. Carneiro, “Meta-learned modality-weighted knowledge distillation for robust multi-modal learning with missing data,” 2025. [Online]. Availabl...
2025 arXiv
-
[28]
Activation sparsity opportunities for compressing general large language models,
N. Dhar, B. Deng, M. R. Islam, K. F. Ahmad Nasif, L. Zhao, and K. Suo, “Activation sparsity opportunities for compressing general large language models,” in 2024 IEEE International Performance, Computing, and Communications Conference (IPCCC). IEEE, Nov. 2024, p. 1–9. [Online]...
2024
-
[29]
Sparsing law: Towards large language models with greater activation sparsity,
Y . Luo, C. Song, X. Han, Y . Chen, C. Xiao, X. Meng, L. Deng, J. Wei, Z. Liu, and M. Sun, “Sparsing law: Towards large language models with greater activation sparsity,” 2025. [Online]. Available: https://arxiv.org/abs/2411.02335
2025 arXiv
-
[30]
Training-free activation sparsity in large language models,
J. Liu, P . Ponnusamy, T. Cai, H. Guo, Y . Kim, and B. Athiwaratkun, “Training-free activation sparsity in large language models,” 2025. [Online]. Available: https://arxiv.org/abs/2408.14690
2025 arXiv
-
[31]
R-sparse: Rank-aware activation sparsity for efficient llm inference,
Z. Zhang, Z. Liu, Y . Tian, H. Khaitan, Z. Wang, and S. Li, “R-sparse: Rank-aware activation sparsity for efficient llm inference,” 2025. [Online]. Available: https://arxiv.org/abs/2504.19449
2025 arXiv
-
[32]
Learning both weights and connections for efficient neural networks,
S. Han, J. Pool, J. Tran, and W. J. Dally, “Learning both weights and connections for efficient neural networks,” 2015. [Online]. Available: https://arxiv.org/abs/1506.02626
2015 arXiv
-
[33]
A simple and effective pruning approach for large language models,
M. Sun, Z. Liu, A. Bair, and J. Z. Kolter, “A simple and effective pruning approach for large language models,” 2024. [Online]. Available: https://arxiv.org/abs/2306.11695
2024 arXiv
-
[34]
Second order derivatives for network pruning: Optimal brain surgeon,
B. Hassibi and D. Stork, “Second order derivatives for network pruning: Optimal brain surgeon,” in Advances in Neural Information Processing Systems, S. Hanson, J. Cowan, and C. Giles, Eds., vol. 5. Morgan-Kaufmann, 1992. [Online]. Available: https://proceedings.neurips.cc/pap...
1992
-
[35]
Sparsegpt: massive language models can be accurately pruned in one-shot,
E. Frantar and D. Alistarh, “Sparsegpt: massive language models can be accurately pruned in one-shot,” in Proceedings of the 40th International Conference on Machine Learning , ser. ICML ’23. JMLR.org, 2023
2023
-
[36]
The LLM surgeon,
T. F. A. van der Ouderaa, M. Nagel, M. V . Baalen, and T. Blankevoort, “The LLM surgeon,” in The Twelfth International Conference on Learning Representations, 2024. [Online]. Available: https://openreview.net/forum?id=DYIIRgwg2i 17 Unifying Depth and Width Pruning for LLMs via...
2024
-
[37]
Accelerating sparse deep neural networks,
A. Mishra, J. A. Latorre, J. Pool, D. Stosic, D. Stosic, G. Venkatesh, C. Yu, and P . Micikevicius, “Accelerating sparse deep neural networks,” 2021. [Online]. Available: https://arxiv.org/abs/2104.08378
2021 arXiv
-
[38]
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned,
E. Voita, D. Talbot, F. Moiseev, R. Sennrich, and I. Titov, “Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, A. Korhonen, D. Traum,...
2019
-
[39]
Are sixteen heads really better than one?
P . Michel, O. Levy, and G. Neubig, “Are sixteen heads really better than one?” in Advances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, Eds., vol. 32. Curran Associates, Inc., 2019. [Online]. Avai...
2019
-
[40]
SlimLLM: Accurate structured pruning for large language models,
J. Guo, X. Chen, Y . Tang, and Y . Wang, “SlimLLM: Accurate structured pruning for large language models,” in Forty-second International Conference on Machine Learning , 2025. [Online]. Available: https://openreview.net/forum? id=2xjUkU7FDb
2025
-
[41]
Language model compression with weighted low-rank factorization,
Y .-C. Hsu, T. Hua, S. Chang, Q. Lou, Y . Shen, and H. Jin, “Language model compression with weighted low-rank factorization,” 2022. [Online]. Available: https://arxiv.org/abs/2207.00112
2022 arXiv
-
[42]
Asvd: Activation-aware singular value decomposition for compressing large language models,
Z. Yuan, Y . Shang, Y . Song, D. Y ang, Q. Wu, Y . Y an, and G. Sun, “Asvd: Activation-aware singular value decomposition for compressing large language models,” 2025. [Online]. Available: https://arxiv.org/abs/2312.05821
2025 arXiv
-
[43]
Svd-llm: Truncation-aware singular value decomposition for large language model compression,
X. Wang, Y . Zheng, Z. Wan, and M. Zhang, “Svd-llm: Truncation-aware singular value decomposition for large language model compression,” 2025. [Online]. Available: https://arxiv.org/abs/2403.07378
2025 arXiv
-
[44]
Streamlining redundant layers to compress large language models,
X. Chen, Y . Hu, J. Zhang, Y . Wang, C. Li, and H. Chen, “Streamlining redundant layers to compress large language models,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=IC5RJvRoMp
2025
-
[45]
DLP: Dynamic layerwise pruning in large language models,
Y . Chen, B. Cheng, J. Han, Y . Zhang, Y . Li, and S. Zhang, “DLP: Dynamic layerwise pruning in large language models,” in Forty-second International Conference on Machine Learning , 2025. [Online]. Available: https: //openreview.net/forum?id=11id5ppGZ8
2025
-
[46]
Prompt-based depth pruning of large language models,
J. Wee, M. Park, and J. Lee, “Prompt-based depth pruning of large language models,” in Forty-second International Conference on Machine Learning, 2025. [Online]. Available: https://openreview.net/forum?id=hRxHF1xPYB
2025
-
[47]
Less is more: Towards green code large language models via unified structural pruning,
G. Y ang, Y . Zhou, X. Zhang, W. Cheng, K. Liu, X. Chen, T. Y . Zhuo, and T. Chen, “Less is more: Towards green code large language models via unified structural pruning,” 2025. [Online]. Available: https://arxiv.org/abs/2412.15921
2025 arXiv
-
[48]
Knapsack problems,
D. Pisinger and P . Toth, “Knapsack problems,” inHandbook of Combinatorial Optimization: Volume1–3. Springer, 1998, pp. 299–428
1998
-
[49]
Qwen2.5 technical report,
Qwen, :, A. Y ang, B. Y ang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Y ang, J. Tu, J. Zhang, J. Y ang, J. Y ang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Y ang, L. Yu, M. Li, M. Xue, P . Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T...
2025 arXiv
-
[50]
Phi-4 technical report,
M. Abdin, J. Aneja, H. Behl, S. Bubeck, R. Eldan, S. Gunasekar, M. Harrison, R. J. Hewett, M. Javaheripi, P . Kauffmann, J. R. Lee, Y . T. Lee, Y . Li, W. Liu, C. C. T. Mendes, A. Nguyen, E. Price, G. de Rosa, O. Saarikivi, A. Salim, S. Shah, X. Wang, R. Ward, Y . Wu, D. Yu, C...
2024 arXiv
-
[51]
2SSP: A two-stage framework for structured pruning of LLMs,
F. Sandri, E. Cunegatti, and G. Iacca, “2SSP: A two-stage framework for structured pruning of LLMs,” Transactions on Machine Learning Research, 2025. [Online]. Available: https://openreview.net/forum?id=Qd7LzJBg21
2025
-
[52]
Pointer sentinel mixture models,
S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer sentinel mixture models,” 2016. [Online]. Available: https://arxiv.org/abs/1609.07843 18 P . Goel, A. Sengupta, A. Nambi and T. Chakraborty
2016 arXiv
-
[53]
The LAMBADA dataset: Word prediction requiring a broad discourse context,
D. Paperno, G. Kruszewski, A. Lazaridou, N. Q. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández, “The LAMBADA dataset: Word prediction requiring a broad discourse context,” in Proceedings of the 54th Annual Meeting of the Association for Computational Lin...
2016
-
[54]
Piqa: Reasoning about physical commonsense in natural language,
Y . Bisk, R. Zellers, R. L. Bras, J. Gao, and Y . Choi, “Piqa: Reasoning about physical commonsense in natural language,” in Thirty-Fourth AAAI Conference on Artificial Intelligence, 2020
2020
-
[55]
PROST: Physical reasoning about objects through space and time,
S. Aroca-Ouellette, C. Paik, A. Roncone, and K. Kann, “PROST: Physical reasoning about objects through space and time,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , C. Zong, F. Xia, W. Li, and R. Navigli, Eds. Online: Association for Computat...
2021
-
[56]
CommonsenseQA: A question answering challenge targeting commonsense knowledge,
A. Talmor, J. Herzig, N. Lourie, and J. Berant, “CommonsenseQA: A question answering challenge targeting commonsense knowledge,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, V...
2019
-
[57]
Think you have solved question answering? try arc, the ai2 reasoning challenge,
P . Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord, “Think you have solved question answering? try arc, the ai2 reasoning challenge,” 2018. [Online]. Available: https://arxiv.org/abs/1803.05457
2018 arXiv
-
[58]
MathQA: Towards interpretable math word problem solving with operation-based formalisms,
A. Amini, S. Gabriel, S. Lin, R. Koncel-Kedziorski, Y . Choi, and H. Hajishirzi, “MathQA: Towards interpretable math word problem solving with operation-based formalisms,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational ...
2019
-
[59]
Can a suit of armor conduct electricity? a new dataset for open book question answering,
T. Mihaylov, P . Clark, T. Khot, and A. Sabharwal, “Can a suit of armor conduct electricity? a new dataset for open book question answering,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, E. Riloff, D. Chiang, J. Hockenmaier, and J....
2018
-
[60]
What disease does this patient have? a large-scale open domain question answering dataset from medical exams,
D. Jin, E. Pan, N. Oufattole, W.-H. Weng, H. Fang, and P . Szolovits, “What disease does this patient have? a large-scale open domain question answering dataset from medical exams,” Applied Sciences , vol. 11, no. 14, 2021. [Online]. Available: https://www.mdpi.com/2076-3417/1...
2021
-
[61]
Blimp: The benchmark of linguistic minimal pairs for english,
A. Warstadt, A. Parrish, H. Liu, A. Mohananey, W. Peng, S.-F. Wang, and S. R. Bowman, “Blimp: The benchmark of linguistic minimal pairs for english,” Transactions of the Association for Computational Linguistics , vol. 8, pp. 377–392,
-
[62]
BoolQ: Exploring the surprising difficulty of natural yes/no questions,
C. Clark, K. Lee, M.-W . Chang, T. Kwiatkowski, M. Collins, and K. Toutanova, “BoolQ: Exploring the surprising difficulty of natural yes/no questions,” in Proceedings of the 2019 Conference of the No rth American Chapter of the Association for Computational Linguistics: Human L...
2019
-
[63]
Winogrande: An adversarial winograd schema challenge at scale,
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y . Choi, “Winogrande: An adversarial winograd schema challenge at scale,” Communications of the ACM, vol. 64, no. 9, pp. 99–106, 2021
2021
-
[64]
CoQA: A conversational question answering challenge,
S. Reddy, D. Chen, and C. D. Manning, “CoQA: A conversational question answering challenge,” Transactions of the Association for Computational Linguistics , vol. 7, pp. 249–266, 2019. [Online]. Available: https://aclanthology.org/ Q19-1016/
2019
-
[65]
TruthfulQA: Measuring how models mimic human falsehoods,
S. Lin, J. Hilton, and O. Evans, “TruthfulQA: Measuring how models mimic human falsehoods,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , S. Muresan, P . Nakov, and A. Villavicencio, Eds. Dublin, Ireland: A...
2022
-
[66]
Gender bias in coreference resolution,
R. Rudinger, J. Naradowsky, B. Leonard, and B. Van Durme, “Gender bias in coreference resolution,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), M. Wal...
2018
-
[67]
Moral stories: Situated reasoning about norms, intents, actions, and their consequences,
D. Emelin, R. Le Bras, J. D. Hwang, M. Forbes, and Y . Choi, “Moral stories: Situated reasoning about norms, intents, actions, and their consequences,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M.-F. Moens, X. Huang, L. Specia, ...
2021
-
[68]
Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification,
W. Lian, G. Wang, B. Goodson, E. Pentland, A. Cook, C. Vong, and ”Teknium”, “Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification,” 2023. [Online]. Available: https://huggingface.co/datasets/ Open-Orca/SlimOrca
2023
-
[69]
Stanford alpaca: An instruction-following llama model,
R. Taori, I. Gulrajani, T. Zhang, Y . Dubois, X. Li, C. Guestrin, P . Liang, and T. B. Hashimoto, “Stanford alpaca: An instruction-following llama model,” https://github.com/tatsu-lab/stanford_alpaca, 2023
2023
-
[70]
Exploring the limits of transfer learning with a unified text-to-text transformer,
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P . J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res. , vol. 21, no. 1, Jan. 2020
2020
-
[71]
Pytorch: An imperative style, high-performance deep learning library,
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Y ang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-p...
2019
-
[72]
Transformers: State-of-the-art natural language processing,
T. Wolf, L. Debut, V . Sanh, J. Chaumond, C. Delangue, A. Moi, P . Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P . von Platen, C. Ma, Y . Jernite, J. Plu, C. Xu, T. Le Scao, S. Gugger, M. Drame, Q. Lhoest, and A. Rush, “Transformers: State-of-the-art natu...
2020
-
[73]
A framework for few-shot language model evaluation,
L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noac’h, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou, “A framework...
2024
-
[74]
Orca: Progressive learning from complex explanation traces of gpt-4,
S. Mukherjee, A. Mitra, G. Jawahar, S. Agarwal, H. Palangi, and A. Awadallah, “Orca: Progressive learning from complex explanation traces of gpt-4,” 2023. [Online]. Available: https://arxiv.org/abs/2306.02707
2023 arXiv
-
[75]
LoRA: Low-rank adaptation of large language models,
E. J. Hu, Y . Shen, P . Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum?id=nZeVKeeFYf9 20 P . Go...
2022
-
[77]
he” changed to “she
which contains complex reasoning questions that require models to refer to a given corpus of salient facts to answer correctly. Natural Language Understanding and Inference. We gauge a model’s semantic understanding on natural understand- ing with the help of the BLIMP [ 61], ...
-
[2020]
Available: https://doi.org/10.1162/tacl_a_00321
[Online]. Available: https://doi.org/10.1162/tacl_a_00321
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.