Pith. sign in

REVIEW 2 major objections 5 minor 77 references

Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization

T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Structured pruning of large language models can be cast as a 0/1 knapsack problem, and the paper claims dynamic programming with a fine-grained width stage achieves near-exact budget adherence and better task-level stability than greedy…

desk verdict Useful two-stage knapsack-based pruner with broad empirical evaluation; the 'conditional optimality' claim needs tightening because the discretized DP can violate the stated budget and the importance loop is under-specified. read the letter →

arxiv 2608.12953 v1 pith:INTJUYIC submitted 2026-08-13 cs.CL cs.LG

classification cs.CLcs.LG
keywords LLMcompressionstructuredpruningknapsackoptimizationdepthwidthbudgetadherenceimportanceestimationCRAFT
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that structured pruning of LLMs should be treated as a global budget-allocation problem instead of a sequence of greedy local choices. It decomposes each transformer layer into coarse items — the attention block, the MLP block, or the whole layer — assigns each a parameter weight and a marginal importance value, and solves the resulting 0/1 knapsack with dynamic programming. A second, fine-grained stage prunes MLP columns to spend the leftover budget, which the paper credits for hitting target compression ratios almost exactly: a CRAFT of 0.98 where, by its measurements, baseline pruners deviate by up to 33%. Across four architectures and 18 tasks, SNIPER reports the best average retention in nearly every configuration, a mean rank of 1.25, and lower task-wise variance than six baselines. A sympathetic reader takes away that non-greedy optimization can turn pruning from a myopic heuristic into a budget-exact, conditionally optimal procedure.

What carries the argument

The load-bearing object is a 0/1 knapsack instance built from a mixed-granularity decomposition of a transformer: each layer contributes three mutually exclusive items — attention block, MLP block, or whole layer — grouped so the solver picks at most one branch per layer. The dynamic programming recurrence $f(i,j)$ over the first $i$ items and capacity $j$ finds the maximum-value subset, and a discretization factor $\alpha$ reduces the capacity dimension to keep runtime tractable. A second mechanism, the sensitivity score $\Omega_j = |D_{j,:} \cdot (G_{:,j} \odot U_{:,j})|$ for each MLP column, tells the fine-grained stage which columns to prune after the coarse stage has left a residual capacity. Together these two mechanisms turn pruning into a global, budget-constrained selection problem followed by a precision fill step.

What would settle it

Compare SNIPER's selected component set against a greedy depth pruner that uses the same iterative marginal importance scores on the same model and budget; if the greedy set matches or beats SNIPER's retention, the reported advantage is carried by the scores rather than by the knapsack solver, and the conditional-optimality claim would not be what drives the result.

Watch

Extended reading notes

Core claim

The paper's central claim is that the depth-pruning step of an LLM pruner can be set up as Equation (1): choose the subset $S^* \subseteq X$ of coarse components that maximizes total value $\sum_{x \in S} v(x)$ while keeping total parameter weight $\sum_{x \in S} w(x)$ at or below budget $C$. The claimed advance is not global optimality — the authors state plainly that the solution is optimal only with respect to fixed importance estimates — but it is a guarantee that greedy methods lack. The importance estimates themselves are computed iteratively as the marginal logit divergence of removing each component from a partially pruned model, so the scores are meant to reflect joint effects rather than isolated component quality. The follow-up width stage distributes the remaining budget across MLP layers in inverse proportion to their importance and removes the least sensitive columns, which is how the method achieves near-exact budget adherence while preserving tensor structure. The paper also introduces CRAFT, the ratio of observed to target compression, as a way to quantify this budget fidelity.

Load-bearing premise

The knapsack's optimality is conditional on the importance scores being correct, and those scores are heuristic and path-dependent; if the scores fail to capture true performance loss, the optimal selection is optimal for the wrong objective.

Editorial extensions

If this is right

  • If the central claim holds, structured pruning becomes budget-exact: a practitioner who asks for a 25% or 35% model gets a model that actually has that parameter count, removing the capacity slack the paper documents in existing pruners.
  • The conditional-optimality guarantee means that, for the importance scores SNIPER computes, no other subset of the same coarse components within the budget has higher total value; greedy depth pruners offer no comparable guarantee.
  • Because the importance scores transfer across compression ratios with only a small retention loss, the expensive scoring pass can be amortized over multiple target budgets rather than recomputed from scratch.
  • The framework's performance on dense, reasoning-tuned, fused-MLP, and mixture-of-experts architectures suggests the knapsack formulation is not bound to one transformer design.
  • Lower task-wise standard deviation across 18 tasks implies the method is less likely to over-optimize one metric at the expense of others, addressing a documented failure mode of greedy pruning.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the observed transferability of importance scores across budgets suggests the same scores could be reused across checkpoints of a model family, making the dominant scoring cost a one-time investment per family.
  • Beyond the paper: because attention and MLP blocks enter the knapsack as exchangeable items, the same solver could be pointed at expert-level components in mixture-of-experts models or at non-uniform per-layer width budgets without changing the core formulation.
  • Beyond the paper: the strong overlap of pruned layers across budgets, which the paper presents as interpretable salience, could be read as a diagnostic tool for architectural redundancy that is independent of any particular compression target.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper proposes SNIPER, a two-stage structured pruning framework for LLMs. Stage 1 solves a 0/1 knapsack problem over coarse components (attention blocks, MLP blocks, and whole layers) using dynamic programming, with component importance estimated from logit divergence under iterative partial pruning. Stage 2 performs fine-grained MLP column pruning to fill the remaining parameter budget. The paper introduces the Compression Ratio Adherence Factor (CRAFT) as a metric for budget fidelity and evaluates SNIPER on four architectures, six baselines, and 18 tasks, reporting a mean rank of 1.25 and near-exact compression-ratio adherence.

Significance. If the central claims hold, SNIPER would be a practically useful alternative to greedy structured pruners: it offers a non-greedy selection mechanism, a simple budget-fidelity metric, and an unusually broad evaluation suite spanning diverse architectures and tasks. The DP recurrence is standard and the empirical comparison is extensive, including calibration sensitivity, importance-score transfer, and ablation studies. The conditional-optimality framing is honest in that it does not claim a global guarantee over all pruning configurations, and the paper explicitly concedes in Section 8 that no guarantee extends to importance estimation. However, the two concerns below affect load-bearing parts of the paper: feasibility of the claimed capacity constraint, and the circularity/under-specification of the iterative importance estimation loop.

major comments (2)
  1. [Section 3.2, Eq. (1); Table 8] The DP solution is not guaranteed to satisfy the capacity constraint of Eq. (1). The recurrence discretizes weights as floor(w(x)/alpha) and capacity as floor(C/alpha); rounding weights downward allows the reconstructed set to have true weight exceeding C. Table 8 shows this is not hypothetical: taking the compression ratio as the fraction of parameters removed, at target CR 25% SNIPER's observed CR is 24.62%, so the final model retains 75.38% of parameters while the budget allows 75%; since Stage 2 only removes parameters, Stage 1 must have already exceeded the budget. The same pattern occurs for every target CR in Table 8 (observed CR is below target in all five rows). Consequently, the statement that SNIPER yields a 'conditionally optimal' allocation with respect to Eq. (1) is not supported for the actual parameter budget; at best the solver is optimal for a discretized surrogate problem. The paper should either use a feasibility-preserving discretization (e.g., rounding weights up and capacity down), state the guarantee for the discretized problem and provide a feasible repair step, or explicitly re-define Eq. (1) in terms of the discretized weights.
  2. [Section 3.3 and Section 5.4] The importance scores v(x_i) in Eq. (3) are defined relative to a partially pruned model M^{(i-1)}, but Section 3.3 never specifies how M^{(i-1)} is constructed. Section 5.4 states that its pruning configuration is 'determined by our dynamic programming solver,' i.e., by the same S* that Eq. (1) computes from v. This makes the definition of v circular unless an explicit iterative fixed-point procedure (initialization, update rule, convergence criterion) is provided. Because the paper's optimality guarantee is explicitly conditional on fixed importance estimates, the missing specification means the guarantee does not currently apply to the actual algorithm used in the experiments. The ablation against leave-one-out scoring in Table 4 shows that the choice of scoring procedure materially changes results, so this is a central issue rather than a cosmetic one.
minor comments (5)
  1. [Section 3.2; Section 6.4] Cross-references are inconsistent: Section 3.2 cites 'Appendix 6.5' and Section 6.4 cites 'Appendix 6.2'; both should point to the corresponding sections (6.5 and 6.2, respectively) rather than to an appendix.
  2. [Abstract; Section 5.3; Table 8] The headline CRAFT value of 0.98 appears in the abstract, Section 5.3, and Figure 1, but Table 8's SNIPER rows average to approximately 0.97. The paper should reconcile the reported number or indicate explicitly that 0.98 refers to the values in Table 7 rather than Table 8.
  3. [Section 5.1; Table 1] The text states that at 35% compression on Qwen3-8B with RFT, SNIPER reduces task-wise variance to 15.90% 'compared to 17–30% for competing pruners,' but Table 1 reports 2SSP with Std RP 15.80% in the same configuration, which is lower than SNIPER's 15.90%.
  4. [Section 5; Table 1] The sentence 'Across all configurations except one, SNIPER achieves the highest average retention performance' is contradicted by the same paragraph: on LLaMA-3.1-8B-Instruct at 25% compression with RFT, both ReplaceMe (88.90) and ShortGPT (88.47) outperform SNIPER (87.34). The count should be corrected.
  5. [Appendix B] LAMBADA appears in both 'Generative Performance' and 'NLU and Inference' task lists. Please clarify whether it is evaluated once and categorized twice, or whether two different task formats (log-perplexity vs. next-word prediction accuracy) are used.

Circularity Check

1 steps flagged · score 6.0 of 10

Importance scores are defined using the DP solution they are supposed to determine, making the claimed conditional optimum self-referential.

  1. self definitional [Section 3.3 (Eq. 3) and Section 5.4 (Iterative Pruning for Importance Estimation)]
    "To estimate component importance, SNIPER computes the marginal degradation in performance incurred by dropping a component from a partially pruned model, whose pruning configuration is determined by our dynamic programming solver (Section 3.3)."

    In Eq. 3, v(x_i) is the logit divergence from removing x_i from M(i-1), where Section 3.3 defines M(i-1) as 'the model pruned up to component i-1.' Section 5.4 states that this partially pruned model's configuration is 'determined by our dynamic programming solver.' But the DP solver (Eqs. 1-2) takes v(x_i) as fixed inputs and produces the selected set S*. Thus the optimization's inputs are defined in terms of the optimization's own output: v(x_i) depends on which components the DP already pruned, while the DP's choice depends on v(x_i). The paper provides no external initialization or fixed-point iteration, so the claimed 'conditional optimum with respect to fixed importance estimates' is conditional on scores that are not fixed independent inputs.

full rationale

The knapsack recurrence (Eq. 2) genuinely optimizes the stated objective for fixed values v(x_i), and the 18-task evaluation is an external benchmark that does not reduce to the method's inputs. However, the paper's own ablation text makes the importance estimates load-bearing and self-referential: v(x_i) is computed from a partially pruned model whose configuration is 'determined by our dynamic programming solver,' while the solver's selection is in turn a function of v(x_i). As written, Eq. 3's input depends on Eq. 1's output, so the advertised 'conditional optimum with respect to fixed importance estimates' lacks a fixed, independent input; Section 8 explicitly declines to extend guarantees to importance estimation. Separately, the reported CRAFT is a designed property of Stage 2 (columns are pruned until the residual budget is filled), not an independent prediction, and the alpha-discretized DP can return sets whose true weight exceeds the budget (Table 8 at target 25% shows observed CR 24.62%, implying Stage 1 retained more than 75% of parameters), so feasibility of the claimed optimum is also not established. These issues are partly correctness concerns, but the self-referential scoring loop is a genuine circularity in the derivation of the main optimality claim.

Assumptions & free parameters 3 free parameters · 6 assumptions · 1 invented entities

The knapsack machinery itself is standard and parameter-light. The load-bearing assumptions are all about the heuristic importance scores: logit-divergence fidelity, deterministic sequencing, column sensitivity, fair baseline re-engineering, and comparable LoRA recovery. CRAFT is the only new ledger entry and it is a metric, not an entity with independent predictive evidence.

free parameters (3)
  • Discretizing factor alpha = 32
    Controls DP capacity granularity and runtime in Section 3.2; Table 7 shows invariant behavior from 8 to 32768 and the authors choose 32 as a conservative value.
  • Softmax temperature tau = 1
    Used in Equation 4 to distribute fine-stage pruning budget across layers; default value, not tuned or justified.
  • Calibration sample count = 50 samples
    Importance estimation in Section 3.3 uses 50 randomly drawn SlimOrca samples; robustness from 50 to 1000 is tested in Section 6.1, so it is an experimenter-chosen input rather than a fitted constant.
assumptions (6)
  • standard math The 0/1 knapsack recurrence in Equation 2 with grouped mutually exclusive items returns the exact optimum for the stated objective in Equation 1.
    Standard DP correctness; invoked in Section 3.2 and used as the paper's 'conditional optimality' guarantee.
  • domain assumption Logit-divergence importance in Equation 3 is a faithful proxy for how removing a component affects downstream task performance.
    The paper offers an ablation against leave-one-out but no proof that the marginal logit divergence ranks components correctly for the 18 evaluation tasks; Section 3.3 and Limitations in Section 8 both concede importance estimation is heuristic.
  • domain assumption Transformer components can be deterministically sequenced so the iterative marginal contributions in Section 3.3 are well-defined and meaningful.
    Section 3.2 says components can be deterministically sequenced; the path dependence of scores computed on a partially pruned model is never analyzed.
  • domain assumption The column sensitivity score in Equation 6 identifies the least salient MLP columns.
    Used in Stage 2; Section 5.4 compares against magnitude and activation baselines, but no formal justification is given.
  • domain assumption Baselines re-engineered for new architectures behave as their original implementations intended.
    Section 4 notes ReplaceMe and ShortGPT do not natively support newer models; the re-engineering details are not released, so the comparison assumes faithful adaptation.
  • domain assumption LoRA recovery fine-tuning with a single epoch restores all methods comparably.
    Appendix A fixes one RFT recipe for all methods; no analysis shows the recipe is equally suitable across pruners.
invented entities (1)
  • Compression Ratio Adherence Factor (CRAFT)
    purpose: Quantify budget fidelity as observed over target compression ratio; used in Figure 1, Table 8, and the abstract's 0.98 claim.
    CRAFT is a definition rather than a falsifiable prediction; it is transparently computable from the paper's tables but provides no external handle.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization." pith.science (2026). https://pith.science/paper/INTJUYIC

@misc{pith2026260812953,
  author       = {Pith},
  title        = {Pith review of: Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/INTJUYIC}},
  note         = {Machine review of arXiv:2608.12953}
}
read the original abstract

Structured pruning is a promising approach for compressing large language models (LLMs), yet existing methods rely heavily on greedy heuristics that produce myopic decisions, and often fail to precisely meet target compression budgets. We present SNIPER, a two-stage structured pruning framework that solves a knapsack optimization over coarse-granularity components to yield conditionally optimal parameter allocations with respect to fixed importance estimates, followed by a fine-grained pruning stage to meet strict budget constraints. We introduce the Compression Ratio Adherence Factor (CRAFT) to quantify budget fidelity, showing that while existing pruners deviate from target compression ratios by up to 33%, SNIPER achieves near-exact adherence with a CRAFT score of 0.98. Evaluations across four diverse architectures over a set of 18 tasks spanning five domains demonstrate SNIPER's consistent improvements in average performance retention and task-level stability over six state-of-the-art pruners. Across all pruning configurations, SNIPER achieves an excellent mean rank of 1.25, indicating its robust cross-architectural generalizability and excellent reliability.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

77 extracted references · 44 canonical work pages

  1. [1]

    The llama 3 herd of models,

    A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Y ang, A. Fan, A. Goyal, A. Hartshorn, A. Y ang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. M...

  2. [2]

    Qwen3 technical report,

    A. Y ang, A. Li, B. Y ang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Y ang, J. Tu, J. Zhang, J. Y ang, J. Y ang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Y ang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P . Zhang, P . Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. ...

  3. [3]

    Deepseek-v3 technical report,

    DeepSeek-AI, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Guo, D. Y ang, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, J. L. Cai, J. Liang, J. Guo, J. Ni, J. Li, J. Wang, J. Chen, J. Chen, J. Yuan, ...

  4. [4]

    gpt-oss-120b & gpt-oss-20b model card,

    OpenAI, :, S. Agarwal, L. Ahmad, J. Ai, S. Altman, A. Applebaum, E. Arbus, R. K. Arora, Y . Bai, B. Baker, H. Bao, B. Barak, A. Bennett, T. Bertao, N. Brett, E. Brevdo, G. Brockman, S. Bubeck, C. Chang, K. Chen, M. Chen, E. Cheung, A. Clark, D. Cook, M. Dukhan, C. Dvorak, K. Fives, V . Fomenko, T. Garipov, K. Georgiev, M. Glaese, T. Gogineni, A. Goucher, ...

  5. [5]

    A survey on model compression for large language models,

    X. Zhu, J. Li, Y . Liu, C. Ma, and W. Wang, “A survey on model compression for large language models,” 2024. [Online]. Available: https://arxiv.org/abs/2308.07633

  6. [6]

    A survey of small language models,

    C. V . Nguyen, X. Shen, R. Aponte, Y . Xia, S. Basu, Z. Hu, J. Chen, M. Parmar, S. Kunapuli, J. Barrow, J. Wu, A. Singh, Y . Wang, J. Gu, F. Dernoncourt, N. K. Ahmed, N. Lipka, R. Zhang, X. Chen, T. Yu, S. Kim, H. Deilamsalehy, N. Park, M. Rimer, Z. Zhang, H. Y ang, R. A. Rossi, and T. H. Nguyen, “A survey of small language models,” 2024. [Online]. Availa...

  7. [7]

    Efficient 8-bit quantization of transformer neural machine language translation model,

    A. Bhandare, V . Sripathi, D. Karkada, V . Menon, S. Choi, K. Datta, and V . Saletore, “Efficient 8-bit quantization of transformer neural machine language translation model,” 2019. [Online]. Available: https://arxiv.org/abs/1906.00532

  8. [8]

    Qlora: efficient finetuning of quantized llms,

    T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “Qlora: efficient finetuning of quantized llms,” in Proceedings of the 37th International Conference on Neural Information Processing Systems , ser. NIPS ’23. Red Hook, NY, USA: Curran Associates Inc., 2023

Show all 77 references
  1. [9]

    CBQ: Cross-block quantization for large language models,

    X. Ding, X. Liu, Z. Tu, Y . Zhang, W. Li, J. Hu, H. Chen, Y . Tang, Z. Xiong, B. Yin, and Y . Wang, “CBQ: Cross-block quantization for large language models,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview...

  2. [10]

    Minillm: Knowledge distillation of large language models,

    Y . Gu, L. Dong, F. Wei, and M. Huang, “Minillm: Knowledge distillation of large language models,” 2025. [Online]. Available: https://arxiv.org/abs/2306.08543

  3. [11]

    A good learner can teach better: Teacher-student collaborative knowledge distillation,

    A. Sengupta, S. Dixit, M. S. Akhtar, and T. Chakraborty, “A good learner can teach better: Teacher-student collaborative knowledge distillation,” in Proceedings of the The Twelfth International Conference on Learning Representations, Virtual Event, 2023, pp. 25–29

  4. [12]

    Replaceme: Network simplification via depth pruning and transformer block linearization,

    D. Shopkhoev, A. Ali, M. Zhussip, V . Malykh, S. Lefkimmiatis, N. Komodakis, and S. Zagoruyko, “Replaceme: Network simplification via depth pruning and transformer block linearization,” 2025. [Online]. Available: https: //arxiv.org/abs/2505.02819

  5. [13]

    Shortgpt: Layers in large language models are more redundant than you expect,

    X. Men, M. Xu, Q. Zhang, B. Wang, H. Lin, Y . Lu, X. Han, and W. Chen, “Shortgpt: Layers in large language models are more redundant than you expect,” 2024. [Online]. Available: https://arxiv.org/abs/2403.03853

  6. [14]

    SLEB: Streamlining LLMs through redundancy verification and elimination of transformer blocks,

    J. Song, K. Oh, T. Kim, H. Kim, Y . Kim, and J.-J. Kim, “SLEB: Streamlining LLMs through redundancy verification and elimination of transformer blocks,” in Proceedings of the 41st International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, R. S...

  7. [15]

    The unreasonable ineffectiveness of the deeper layers,

    A. Gromov, K. Tirumala, H. Shapourian, P . Glorioso, and D. Roberts, “The unreasonable ineffectiveness of the deeper layers,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=ngmEcEer8a

  8. [16]

    Llm-pruner: on the structural pruning of large language models,

    X. Ma, G. Fang, and X. Wang, “Llm-pruner: on the structural pruning of large language models,” in Proceedings of the 37th International Conference on Neural Information Processing Systems , ser. NIPS ’23. Red Hook, NY, USA: Curran Associates Inc., 2023

  9. [17]

    Y ou only prune once: Designing calibration-free model compression with policy learning,

    A. Sengupta, S. Chaudhary, and T. Chakraborty, “Y ou only prune once: Designing calibration-free model compression with policy learning,” in The Thirteenth International Conference on Learning Representations, 2025. [Online]. Available: https://openreview.net/forum?id=5RZoYIT3u6

  10. [18]

    Slicegpt: Compress large language models by deleting rows and columns,

    S. Ashkboos, M. L. Croci, M. G. do Nascimento, T. Hoefler, and J. Hensman, “Slicegpt: Compress large language models by deleting rows and columns,” 2024. [Online]. Available: https://arxiv.org/abs/2401.15024

  11. [19]

    Shortened llama: Depth pruning for large language models with comparison of retraining methods,

    B.-K. Kim, G. Kim, T.-H. Kim, T. Castells, S. Choi, J. Shin, and H.-K. Song, “Shortened llama: Depth pruning for large language models with comparison of retraining methods,” 2024. [Online]. Available: https://arxiv.org/abs/2402.02834 16 P . Goel, A. Sengupta, A. Nambi and T. ...

  12. [20]

    Sliding-window merging for compacting patch-redundant layers in llms,

    X. Ding, R. Sun, Y . Zhang, X. Y an, Y . Zhou, K. Huang, S. Fu, A. I. Aviles-Rivero, C. Xie, and Y . Zhu, “Sliding-window merging for compacting patch-redundant layers in llms,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 40, no. 25, p. 20826–20834, Ma...

  13. [21]

    Beware of calibration data for pruning large language models,

    Y . Ji, Y . Xiang, J. Li, Q. Xia, P . Li, X. Duan, Z. Wang, and M. Zhang, “Beware of calibration data for pruning large language models,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=x83w6yGIWb

  14. [22]

    I-bert: Integer-only bert quantization,

    S. Kim, A. Gholami, Z. Y ao, M. W. Mahoney, and K. Keutzer, “I-bert: Integer-only bert quantization,” in Proceedings of the 38th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul...

  15. [23]

    Llm-fp4: 4-bit floating-point quantized transformers,

    S.-y. Liu, Z. Liu, X. Huang, P . Dong, and K.-T. Cheng, “Llm-fp4: 4-bit floating-point quantized transformers,” in Pro- ceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2023, p. 592–605. [Online]. A...

  16. [24]

    Aptq: Attention-aware post-training mixed-precision quantization for large language models,

    Z. Guan, H. Huang, Y . Su, H. Huang, N. Wong, and H. Yu, “Aptq: Attention-aware post-training mixed-precision quantization for large language models,” in Proceedings of the 61st ACM/IEEE Design Automation Conference , ser. DAC ’24. ACM, Jun. 2024, p. 1–6. [Online]. Available: ...

  17. [25]

    Distilling the knowledge in a neural network,

    G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” 2015. [Online]. Available: https://arxiv.org/abs/1503.02531

  18. [26]

    TinyBERT: Distilling BERT for natural language understanding,

    X. Jiao, Y . Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu, “TinyBERT: Distilling BERT for natural language understanding,” in Findings of the Association for Computational Linguistics: EMNLP 2020 , T. Cohn, Y . He, and Y . Liu, Eds. Online: Association for Comp...

  19. [27]

    Meta-learned modality-weighted knowledge distillation for robust multi-modal learning with missing data,

    H. Wang, S. Hassan, Y . Liu, C. Ma, Y . Chen, Q. Li, J. Geng, B. Wang, Y . Tian, Y . Xie, J. Avery, L. Hull, I. Reid, M. Y aqub, and G. Carneiro, “Meta-learned modality-weighted knowledge distillation for robust multi-modal learning with missing data,” 2025. [Online]. Availabl...

  20. [28]

    Activation sparsity opportunities for compressing general large language models,

    N. Dhar, B. Deng, M. R. Islam, K. F. Ahmad Nasif, L. Zhao, and K. Suo, “Activation sparsity opportunities for compressing general large language models,” in 2024 IEEE International Performance, Computing, and Communications Conference (IPCCC). IEEE, Nov. 2024, p. 1–9. [Online]...

  21. [29]

    Sparsing law: Towards large language models with greater activation sparsity,

    Y . Luo, C. Song, X. Han, Y . Chen, C. Xiao, X. Meng, L. Deng, J. Wei, Z. Liu, and M. Sun, “Sparsing law: Towards large language models with greater activation sparsity,” 2025. [Online]. Available: https://arxiv.org/abs/2411.02335

  22. [30]

    Training-free activation sparsity in large language models,

    J. Liu, P . Ponnusamy, T. Cai, H. Guo, Y . Kim, and B. Athiwaratkun, “Training-free activation sparsity in large language models,” 2025. [Online]. Available: https://arxiv.org/abs/2408.14690

  23. [31]

    R-sparse: Rank-aware activation sparsity for efficient llm inference,

    Z. Zhang, Z. Liu, Y . Tian, H. Khaitan, Z. Wang, and S. Li, “R-sparse: Rank-aware activation sparsity for efficient llm inference,” 2025. [Online]. Available: https://arxiv.org/abs/2504.19449

  24. [32]

    Learning both weights and connections for efficient neural networks,

    S. Han, J. Pool, J. Tran, and W. J. Dally, “Learning both weights and connections for efficient neural networks,” 2015. [Online]. Available: https://arxiv.org/abs/1506.02626

  25. [33]

    A simple and effective pruning approach for large language models,

    M. Sun, Z. Liu, A. Bair, and J. Z. Kolter, “A simple and effective pruning approach for large language models,” 2024. [Online]. Available: https://arxiv.org/abs/2306.11695

  26. [34]

    Second order derivatives for network pruning: Optimal brain surgeon,

    B. Hassibi and D. Stork, “Second order derivatives for network pruning: Optimal brain surgeon,” in Advances in Neural Information Processing Systems, S. Hanson, J. Cowan, and C. Giles, Eds., vol. 5. Morgan-Kaufmann, 1992. [Online]. Available: https://proceedings.neurips.cc/pap...

  27. [35]

    Sparsegpt: massive language models can be accurately pruned in one-shot,

    E. Frantar and D. Alistarh, “Sparsegpt: massive language models can be accurately pruned in one-shot,” in Proceedings of the 40th International Conference on Machine Learning , ser. ICML ’23. JMLR.org, 2023

  28. [36]

    The LLM surgeon,

    T. F. A. van der Ouderaa, M. Nagel, M. V . Baalen, and T. Blankevoort, “The LLM surgeon,” in The Twelfth International Conference on Learning Representations, 2024. [Online]. Available: https://openreview.net/forum?id=DYIIRgwg2i 17 Unifying Depth and Width Pruning for LLMs via...

  29. [37]

    Accelerating sparse deep neural networks,

    A. Mishra, J. A. Latorre, J. Pool, D. Stosic, D. Stosic, G. Venkatesh, C. Yu, and P . Micikevicius, “Accelerating sparse deep neural networks,” 2021. [Online]. Available: https://arxiv.org/abs/2104.08378

  30. [38]

    Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned,

    E. Voita, D. Talbot, F. Moiseev, R. Sennrich, and I. Titov, “Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, A. Korhonen, D. Traum,...

  31. [39]

    Are sixteen heads really better than one?

    P . Michel, O. Levy, and G. Neubig, “Are sixteen heads really better than one?” in Advances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, Eds., vol. 32. Curran Associates, Inc., 2019. [Online]. Avai...

  32. [40]

    SlimLLM: Accurate structured pruning for large language models,

    J. Guo, X. Chen, Y . Tang, and Y . Wang, “SlimLLM: Accurate structured pruning for large language models,” in Forty-second International Conference on Machine Learning , 2025. [Online]. Available: https://openreview.net/forum? id=2xjUkU7FDb

  33. [41]

    Language model compression with weighted low-rank factorization,

    Y .-C. Hsu, T. Hua, S. Chang, Q. Lou, Y . Shen, and H. Jin, “Language model compression with weighted low-rank factorization,” 2022. [Online]. Available: https://arxiv.org/abs/2207.00112

  34. [42]

    Asvd: Activation-aware singular value decomposition for compressing large language models,

    Z. Yuan, Y . Shang, Y . Song, D. Y ang, Q. Wu, Y . Y an, and G. Sun, “Asvd: Activation-aware singular value decomposition for compressing large language models,” 2025. [Online]. Available: https://arxiv.org/abs/2312.05821

  35. [43]

    Svd-llm: Truncation-aware singular value decomposition for large language model compression,

    X. Wang, Y . Zheng, Z. Wan, and M. Zhang, “Svd-llm: Truncation-aware singular value decomposition for large language model compression,” 2025. [Online]. Available: https://arxiv.org/abs/2403.07378

  36. [44]

    Streamlining redundant layers to compress large language models,

    X. Chen, Y . Hu, J. Zhang, Y . Wang, C. Li, and H. Chen, “Streamlining redundant layers to compress large language models,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=IC5RJvRoMp

  37. [45]

    DLP: Dynamic layerwise pruning in large language models,

    Y . Chen, B. Cheng, J. Han, Y . Zhang, Y . Li, and S. Zhang, “DLP: Dynamic layerwise pruning in large language models,” in Forty-second International Conference on Machine Learning , 2025. [Online]. Available: https: //openreview.net/forum?id=11id5ppGZ8

  38. [46]

    Prompt-based depth pruning of large language models,

    J. Wee, M. Park, and J. Lee, “Prompt-based depth pruning of large language models,” in Forty-second International Conference on Machine Learning, 2025. [Online]. Available: https://openreview.net/forum?id=hRxHF1xPYB

  39. [47]

    Less is more: Towards green code large language models via unified structural pruning,

    G. Y ang, Y . Zhou, X. Zhang, W. Cheng, K. Liu, X. Chen, T. Y . Zhuo, and T. Chen, “Less is more: Towards green code large language models via unified structural pruning,” 2025. [Online]. Available: https://arxiv.org/abs/2412.15921

  40. [48]

    Knapsack problems,

    D. Pisinger and P . Toth, “Knapsack problems,” inHandbook of Combinatorial Optimization: Volume1–3. Springer, 1998, pp. 299–428

  41. [49]

    Qwen2.5 technical report,

    Qwen, :, A. Y ang, B. Y ang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Y ang, J. Tu, J. Zhang, J. Y ang, J. Y ang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Y ang, L. Yu, M. Li, M. Xue, P . Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T...

  42. [50]

    Phi-4 technical report,

    M. Abdin, J. Aneja, H. Behl, S. Bubeck, R. Eldan, S. Gunasekar, M. Harrison, R. J. Hewett, M. Javaheripi, P . Kauffmann, J. R. Lee, Y . T. Lee, Y . Li, W. Liu, C. C. T. Mendes, A. Nguyen, E. Price, G. de Rosa, O. Saarikivi, A. Salim, S. Shah, X. Wang, R. Ward, Y . Wu, D. Yu, C...

  43. [51]

    2SSP: A two-stage framework for structured pruning of LLMs,

    F. Sandri, E. Cunegatti, and G. Iacca, “2SSP: A two-stage framework for structured pruning of LLMs,” Transactions on Machine Learning Research, 2025. [Online]. Available: https://openreview.net/forum?id=Qd7LzJBg21

  44. [52]

    Pointer sentinel mixture models,

    S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer sentinel mixture models,” 2016. [Online]. Available: https://arxiv.org/abs/1609.07843 18 P . Goel, A. Sengupta, A. Nambi and T. Chakraborty

  45. [53]

    The LAMBADA dataset: Word prediction requiring a broad discourse context,

    D. Paperno, G. Kruszewski, A. Lazaridou, N. Q. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández, “The LAMBADA dataset: Word prediction requiring a broad discourse context,” in Proceedings of the 54th Annual Meeting of the Association for Computational Lin...

  46. [54]

    Piqa: Reasoning about physical commonsense in natural language,

    Y . Bisk, R. Zellers, R. L. Bras, J. Gao, and Y . Choi, “Piqa: Reasoning about physical commonsense in natural language,” in Thirty-Fourth AAAI Conference on Artificial Intelligence, 2020

  47. [55]

    PROST: Physical reasoning about objects through space and time,

    S. Aroca-Ouellette, C. Paik, A. Roncone, and K. Kann, “PROST: Physical reasoning about objects through space and time,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , C. Zong, F. Xia, W. Li, and R. Navigli, Eds. Online: Association for Computat...

  48. [56]

    CommonsenseQA: A question answering challenge targeting commonsense knowledge,

    A. Talmor, J. Herzig, N. Lourie, and J. Berant, “CommonsenseQA: A question answering challenge targeting commonsense knowledge,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, V...

  49. [57]

    Think you have solved question answering? try arc, the ai2 reasoning challenge,

    P . Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord, “Think you have solved question answering? try arc, the ai2 reasoning challenge,” 2018. [Online]. Available: https://arxiv.org/abs/1803.05457

  50. [58]

    MathQA: Towards interpretable math word problem solving with operation-based formalisms,

    A. Amini, S. Gabriel, S. Lin, R. Koncel-Kedziorski, Y . Choi, and H. Hajishirzi, “MathQA: Towards interpretable math word problem solving with operation-based formalisms,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational ...

  51. [59]

    Can a suit of armor conduct electricity? a new dataset for open book question answering,

    T. Mihaylov, P . Clark, T. Khot, and A. Sabharwal, “Can a suit of armor conduct electricity? a new dataset for open book question answering,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, E. Riloff, D. Chiang, J. Hockenmaier, and J....

  52. [60]

    What disease does this patient have? a large-scale open domain question answering dataset from medical exams,

    D. Jin, E. Pan, N. Oufattole, W.-H. Weng, H. Fang, and P . Szolovits, “What disease does this patient have? a large-scale open domain question answering dataset from medical exams,” Applied Sciences , vol. 11, no. 14, 2021. [Online]. Available: https://www.mdpi.com/2076-3417/1...

  53. [61]

    Blimp: The benchmark of linguistic minimal pairs for english,

    A. Warstadt, A. Parrish, H. Liu, A. Mohananey, W. Peng, S.-F. Wang, and S. R. Bowman, “Blimp: The benchmark of linguistic minimal pairs for english,” Transactions of the Association for Computational Linguistics , vol. 8, pp. 377–392,

  54. [62]

    BoolQ: Exploring the surprising difficulty of natural yes/no questions,

    C. Clark, K. Lee, M.-W . Chang, T. Kwiatkowski, M. Collins, and K. Toutanova, “BoolQ: Exploring the surprising difficulty of natural yes/no questions,” in Proceedings of the 2019 Conference of the No rth American Chapter of the Association for Computational Linguistics: Human L...

  55. [63]

    Winogrande: An adversarial winograd schema challenge at scale,

    K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y . Choi, “Winogrande: An adversarial winograd schema challenge at scale,” Communications of the ACM, vol. 64, no. 9, pp. 99–106, 2021

  56. [64]

    CoQA: A conversational question answering challenge,

    S. Reddy, D. Chen, and C. D. Manning, “CoQA: A conversational question answering challenge,” Transactions of the Association for Computational Linguistics , vol. 7, pp. 249–266, 2019. [Online]. Available: https://aclanthology.org/ Q19-1016/

  57. [65]

    TruthfulQA: Measuring how models mimic human falsehoods,

    S. Lin, J. Hilton, and O. Evans, “TruthfulQA: Measuring how models mimic human falsehoods,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , S. Muresan, P . Nakov, and A. Villavicencio, Eds. Dublin, Ireland: A...

  58. [66]

    Gender bias in coreference resolution,

    R. Rudinger, J. Naradowsky, B. Leonard, and B. Van Durme, “Gender bias in coreference resolution,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), M. Wal...

  59. [67]

    Moral stories: Situated reasoning about norms, intents, actions, and their consequences,

    D. Emelin, R. Le Bras, J. D. Hwang, M. Forbes, and Y . Choi, “Moral stories: Situated reasoning about norms, intents, actions, and their consequences,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M.-F. Moens, X. Huang, L. Specia, ...

  60. [68]

    Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification,

    W. Lian, G. Wang, B. Goodson, E. Pentland, A. Cook, C. Vong, and ”Teknium”, “Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification,” 2023. [Online]. Available: https://huggingface.co/datasets/ Open-Orca/SlimOrca

  61. [69]

    Stanford alpaca: An instruction-following llama model,

    R. Taori, I. Gulrajani, T. Zhang, Y . Dubois, X. Li, C. Guestrin, P . Liang, and T. B. Hashimoto, “Stanford alpaca: An instruction-following llama model,” https://github.com/tatsu-lab/stanford_alpaca, 2023

  62. [70]

    Exploring the limits of transfer learning with a unified text-to-text transformer,

    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P . J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res. , vol. 21, no. 1, Jan. 2020

  63. [71]

    Pytorch: An imperative style, high-performance deep learning library,

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Y ang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-p...

  64. [72]

    Transformers: State-of-the-art natural language processing,

    T. Wolf, L. Debut, V . Sanh, J. Chaumond, C. Delangue, A. Moi, P . Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P . von Platen, C. Ma, Y . Jernite, J. Plu, C. Xu, T. Le Scao, S. Gugger, M. Drame, Q. Lhoest, and A. Rush, “Transformers: State-of-the-art natu...

  65. [73]

    A framework for few-shot language model evaluation,

    L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noac’h, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou, “A framework...

  66. [74]

    Orca: Progressive learning from complex explanation traces of gpt-4,

    S. Mukherjee, A. Mitra, G. Jawahar, S. Agarwal, H. Palangi, and A. Awadallah, “Orca: Progressive learning from complex explanation traces of gpt-4,” 2023. [Online]. Available: https://arxiv.org/abs/2306.02707

  67. [75]

    LoRA: Low-rank adaptation of large language models,

    E. J. Hu, Y . Shen, P . Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum?id=nZeVKeeFYf9 20 P . Go...

  68. [77]

    he” changed to “she

    which contains complex reasoning questions that require models to refer to a given corpus of salient facts to answer correctly. Natural Language Understanding and Inference. We gauge a model’s semantic understanding on natural understand- ing with the help of the BLIMP [ 61], ...

  69. [2020]

    Available: https://doi.org/10.1162/tacl_a_00321

    [Online]. Available: https://doi.org/10.1162/tacl_a_00321

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.