Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Low-bit Model Quantization for Deep Neural Networks: A Survey

T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper claims that 179 papers from 2020–2025 low-bit quantization can be organized into eight methodological families and 24 subfamilies.

desk verdict A genuinely useful survey map of 2020-2025 low-bit quantization, softened by an inconsistent scope boundary and the absence of quantitative comparison. read the letter →

arxiv 2505.05530 v1 pith:YPLGILLR submitted 2025-05-08 cs.LG cs.AI

classification cs.LGcs.AI
keywords modelquantizationlow-bitpost-trainingquantization-awaretrainingmixedprecisiondata-freediffusionneuralnetworkcompression
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey tries to establish an organizing map of the last five years of low-bit neural-network quantization, covering 179 papers across 8 main families and 24 subfamilies. It argues that nearly all current methods reduce to choices about how to set quantization scales and zero-points, what loss or metric guides the conversion, which layers get which bit-width, how data distributions are reshaped, and which number format is used. If the map is right, a practitioner can locate any new quantization method within a technical family and see which open problems remain. The authors also state plainly in the supplementary file that they could not provide a quantitative comparison across the surveyed methods, because models, datasets, and quantization schemes differ too much, so the survey's value is the organization rather than a leaderboard.

What carries the argument

The organizing instrument is the taxonomy itself: eight main categories plus 24 subcategories, built on top of the formal quantization operator $X_{\mathrm{int}} = \mathrm{Clamp}(\mathrm{Round}(X_{\mathrm{FP}}/s)+z, n, p)$ with dequantization $\hat{X}=s(X_{\mathrm{int}}-z)$. That operator supplies the survey's vocabulary—bit-width $b$, scale $s$, zero-point $z$—and the taxonomy groups methods by which of these knobs they turn, turning a scattered literature into a decision tree for a practitioner.

What would settle it

Take the 179 cited papers, strip away the survey's own classifications, and have an independent reader assign each paper to one of the eight families using only its abstract and method description; if agreement is no better than chance, or if major methods clearly straddle or fall outside all eight families, the taxonomy is not a stable description of the field. A second check is to count 2020–2025 publications: if extreme 1-bit and 1.58-bit work, which the survey explicitly excludes, dominates the period, then the survey's scope claim omits a major branch of the literature.

Watch

Extended reading notes

Core claim

The paper's central discovery is taxonomic: it claims that the recent quantization literature divides along eight recognizable methodological fault lines—scale and zero-point optimization, metrics and training mechanisms, mixed precision, redistribution of weights and activations, data-free quantization, advanced numeric formats, diffusion-model-specific methods, and a residual 'other' bucket—with 24 finer subcategories. It formalizes quantization as mapping floating-point tensors through clamp, round, scale, and zero-point operations, and presents post-training quantization (PTQ) and quantization-aware training (QAT) as a spectrum rather than a strict dichotomy. The paper also reports that the field has moved beyond plain integer formats into float formats, fixed-point formats, learned rotations, and adaptive rounding, while deliberately excluding 1-bit and 1.58-bit methods as a methodologically separate branch.

Load-bearing premise

The whole map is only as good as the selection and classification of the 179 papers; if the papers were chosen with a hidden bias or assigned to the wrong families, the survey would mislead rather than orient.

Editorial extensions

If this is right

  • A newcomer can use the taxonomy to locate any major 2020–2025 quantization method and identify its core technique without reading the full literature.
  • Because PTQ and QAT are presented as a continuum, the boundary between calibration-only methods and retraining-based methods is expected to keep blurring.
  • The four future directions named by the paper—multimodal deployment, combining quantization with pruning and low-rank compression, software-hardware co-optimization, and task-specific quantization—are where the paper expects the next progress.
  • The survey implies that extreme 1-bit and 1.58-bit quantization is a separate research vein, not a subcase of the eight families, so its conclusions should not be read as covering that line of work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the taxonomy is accurate, a useful next step is to use it as a checklist for designing new methods: a paper that combines scale optimization, rotation-based outlier removal, and adaptive bit allocation would span three families, which the survey itself permits since some papers appear in multiple categories.
  • The absence of comparative benchmarks also suggests an opportunity: building a standardized quantization benchmark with fixed models, calibration sets, and bit-width budgets would let future surveys replace prose organization with reproducible measurement.
  • The eight-way split may over-weight recent LLM and diffusion-model work, since the field's center of gravity moved there in 2023–2025; a reader should expect older CNN-only methods to be concentrated in earlier families such as 'better s and z' and 'redistribution.'
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The manuscript surveys low-bit model quantization for deep neural networks over roughly the past five years. Section 2 introduces the quantization formalism, basic quantizer designs, and a foundational taxonomy; Section 3 then proposes an organizing scheme of eight main categories and twenty-four sub-categories and assigns 179 papers to them (Fig. 4). Section 4 lists future research directions, and an accompanying curated repository is advertised. The abstract promises that state-of-the-art methods are discussed and compared; the main text provides qualitative discussion of each category, while the supplementary file discloses that a quantitative performance comparison was attempted but not completed.

Significance. If the coverage boundary and the comparison claim are made consistent, the survey would be a useful reference map for the low-bit quantization community: it aggregates 179 papers, organizes them by technique rather than by task, includes recently active areas such as diffusion-model quantization and data-free quantization, and provides an accompanying curated list. The paper is particularly valuable for newcomers who need to locate methodological families. I also credit the authors for explicitly stating in the supplementary material that the attempted quantitative comparison failed rather than hiding the limitation; however, that disclosure is not reflected in the main-text claims.

major comments (3)
  1. [Section 1 and Section 3 (3.4.3, 3.7.3, 3.7.4)] The stated scope exclusion of extreme quantization is contradicted by the included methods. The Introduction says "we omit extreme quantization techniques (e.g., 1-bit or 1.58-bit quantization) as they involve substantially different methodologies," yet Section 3.7.3 presents BiDM [159] as employing "a dynamical binary quantizer" (1-bit), Section 3.7.4 includes BitsFusion [161], whose title is "1.99 bits weight quantization of diffusion model," and Section 3.4.3 includes QuIP [104] and QuIP# [105], which are 2-bit lattice-codebook methods. These are not peripheral mentions: they are described as representative techniques in their subsections. A reader therefore cannot infer from the stated scope which methods belong in the survey. The fix is to relax the exclusion statement to match the actual coverage or to move or explicitly mark the extreme-bit methods as out-of-scope but discussed for contrast.
  2. [Abstract and Supplementary Section 1] The abstract claims that the paper "discuss[es] and compare[s] the state-of-the-art quantization methods," and Section 3 promises a "comprehensive analysis and discussion," but Supplementary Section 1 states: "We tried to provide the performance comparison of different quantization methods on these benchmarks, but failed. This is because the models, datasets, and quantification schemes adopted by the recent methods are all different." No accuracy, latency, or memory comparison table appears in the main text. The comparison actually delivered is qualitative only, so the abstract should either remove the word "compare" or the main text should contain a limitations paragraph explicitly stating that no quantitative comparison is provided.
  3. [Section 3 opening and Fig. 4] No selection or classification methodology is reported for the 179 surveyed papers. The text does not state which databases were searched, which keywords or time window were used, what inclusion/exclusion criteria were applied, or how the eight-way assignment was performed and validated. Because the paper's central contribution is a comprehensive and correctly organized map of the field, the absence of this information makes the completeness claim unverifiable. Please add a methodology paragraph describing the paper collection and classification procedure, or rescope the claim to "a curated selection" rather than a systematic survey.
minor comments (6)
  1. [Section 3.2.2] The sentence "Recent works [12], [43] have indicated that traditional loss functions, such as MSE and Exponential Moving Average (EMA), such as MSE and Cross-Entropy (CE), may not be sufficient" contains a duplicated "such as MSE"; it should be rewritten as a single list.
  2. [References and Supplementary Table 1] Reference [143] is truncated to "70" and should be completed, and Supplementary Table 1 contains "Mixral" (should be "Mixtral"), "VICUNA-V1.5 []" with an empty citation, and "SQAI" where ScienceQA appears to be meant.
  3. [Section 3.3.1 and reference list] OWQ appears twice, as [50] and as [193], and Q-BERT appears as both [53] and [194]; these duplicate entries for the same methods should be consolidated into single citations.
  4. [Section 3.6.3, Eq. (15)] The typesetting of the bi-exponent representation "2en|eo" is difficult to parse; please rewrite with proper superscripts or parentheses so the shared exponents are unambiguous.
  5. [Section 3.8 and Fig. 4] The heading in the text is "Other" while Fig. 4 labels the category "Others"; the terminology should be made consistent.
  6. [Section 2.3b and Section 1] The phrase "extreme quantization" appears in Section 2.3b before any definition and in a context that is excluded from the survey's scope; a parenthetical definition or a cross-reference to the scope statement in the Introduction would avoid confusion.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey's taxonomy is an organizational claim over external literature, and its self-citations are not load-bearing.

full rationale

This paper is a survey: it does not perform a predictive or derivational chain, so the main circularity failure modes (fitting a parameter and then predicting a closely related quantity, or defining an output in terms of an input) do not apply. The claimed result is an organizational taxonomy of 179 quantization papers into 8 main categories, which is a classification of external literature rather than a value derived from fitted inputs. Self-citations are present (e.g., QuantSR [41], DSG [125], BiDM [159], MPQ-DM [164], PassionSR [169]), but the descriptions track the published contributions and none is used as a load-bearing premise to justify the eight-way division; removing or replacing these entries would not change the classification structure. The equations in Section 2 are standard quantization formalisms or quoted method objectives, not apparatus that converts inputs into a predicted output. The stated exclusion of extreme quantization in Section 1 conflicts with the inclusion of binary and 1.99-bit methods in Section 3.7 (BiDM [159] uses a dynamical binary quantizer; BitsFusion [161] is described as 1.99 bits), and the supplementary admits that no quantitative comparison across methods is provided; these are scope and completeness limitations, not circular reasoning. No specific reduction by construction or by self-citation can be exhibited, so the circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters appear in this survey; there is no derivation or fitting. The claims that matter are the faithfulness of the 179 paper summaries, the representative selection, and the standard background that quantization accelerates inference. The verifiability of those claims is the entire epistemic load, and the authors themselves disclose that the central comparison goal was not achieved.

assumptions (3)
  • domain assumption The 179 surveyed papers are accurately described and correctly classified into the 8 main categories and 24 sub-categories.
    The survey's utility rests on the faithfulness of its summaries and taxonomy assignments; no inter-annotator agreement or external validation is provided. Invoked throughout Section 3.
  • domain assumption The selected papers are representative of the recent five-year progress in low-bit quantization.
    The survey claims to cover "the most representative and influential studies" (Section 1) but gives no selection criteria, and explicitly excludes 1-bit and 1.58-bit methods as "substantially different methodologies."
  • domain assumption Quantization accelerates inference through memory-access savings and vectorization.
    Supplementary Section 4 asserts these acceleration mechanisms without citation; the survey treats them as established background rather than deriving them.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Low-bit Model Quantization for Deep Neural Networks: A Survey." pith.science (2026). https://pith.science/paper/YPLGILLR

@misc{pith2026250505530,
  author       = {Pith},
  title        = {Pith review of: Low-bit Model Quantization for Deep Neural Networks: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YPLGILLR}},
  note         = {Machine review of arXiv:2505.05530}
}
read the original abstract

With unprecedented rapid development, deep neural networks (DNNs) have deeply influenced almost all fields. However, their heavy computation costs and model sizes are usually unacceptable in real-world deployment. Model quantization, an effective weight-lighting technique, has become an indispensable procedure in the whole deployment pipeline. The essence of quantization acceleration is the conversion from continuous floating-point numbers to discrete integer ones, which significantly speeds up the memory I/O and calculation, i.e., addition and multiplication. However, performance degradation also comes with the conversion because of the loss of precision. Therefore, it has become increasingly popular and critical to investigate how to perform the conversion and how to compensate for the information loss. This article surveys the recent five-year progress towards low-bit quantization on DNNs. We discuss and compare the state-of-the-art quantization methods and classify them into 8 main categories and 24 sub-categories according to their core techniques. Furthermore, we shed light on the potential research opportunities in the field of model quantization. A curated list of model quantization is provided at https://github.com/Kai-Liu001/Awesome-Model-Quantization.

Figures

Figures reproduced from arXiv: 2505.05530 by the authors.

Figure 1
Figure 1. Illustration of different quantization schemes, including symmetric and asymmetric quantization, uniform and non-uniform quantization. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. QAT usually takes large datasets to retrain the weights and [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Distribution of activation and weight in common models. For CNN, the distributions of activations and weights are mostly symmetrical and [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Taxonomy for recent quantization methods. We classify 179 papers into 8 categories and further into 24 sub-categories. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Two paradigms of mixed precision. The left one is input irrelevant, and the bit-width is allocated during the calibration process. The right one is [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Diagram of redistribution. A widely used approach is to leverage [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: Three data-free quantization paradigms. (a) The generator takes [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: Four schematic diagrams of diffusion model quantization. There are relevant studies on the input, model structure, output and calibration [PITH_FULL_IMAGE:figures/full_fig_p013_8.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    JuZhou 1.0 is a 0.387B-parameter T2I diffusion model with 4-step inference achieving 0.69 GenEval, trained on 9M Chinese pairs using Sugon K100 accelerators and deployable on Android/iOS devices.

Reference graph

Works this paper leans on

115 extracted references · 61 canonical work pages · cited by 1 Pith paper

  1. [104]

    One less reason for filter pruning: Gaining free adversarial robustness with structured grouped kernel pruning,

    S. Zhong, Z. You, J. Zhang, S. Zhao, Z. LeClaire, Z. Liu, D. Zha, V . Chaudhary, S. Xu, and X. Hu, “One less reason for filter pruning: Gaining free adversarial robustness with structured grouped kernel pruning,” in Proc. Adv. Neural Inform. Process. Syst. , 2023

  2. [105]

    Storage efficient and dynamic flexible runtime channel pruning via deep reinforcement learning,

    J. Chen, S. Chen, and S. J. Pan, “Storage efficient and dynamic flexible runtime channel pruning via deep reinforcement learning,” in Proc. Int. Conf. Learn. Represent. , vol. 33, 2020, pp. 14 747–14 758

  3. [1]

    Pointer sentinel mixture models,

    S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer sentinel mixture models,” in Proc. Int. Conf. Learn. Represent. , 2017

  4. [2]

    Exploring the limits of transfer learning with a unified text-to-text transformer,

    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P . J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research, vol. 21, no. 140, pp. 1–67, 2020

  5. [3]

    Training a helpful and harmless assistant with reinforcement learning from human feedback,

    Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighanet al., “Training a helpful and harmless assistant with reinforcement learning from human feedback,” arXiv preprint arXiv:2204.05862 , 2022

  6. [4]

    Stanford alpaca: An instruction- following llama model,

    R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P . Liang, and T. B. Hashimoto, “Stanford alpaca: An instruction- following llama model,” 2023

  7. [5]

    The flan collection: Designing data and methods for effective instruction tuning,

    S. Longpre, L. Hou, T. Vu, A. Webson, H. W. Chung, Y. Tay, D. Zhou, Q. V . Le, B. Zoph, J. Wei et al. , “The flan collection: Designing data and methods for effective instruction tuning,” in Proc. Int. Conf. Mach. Learn. PMLR, 2023, pp. 22 631–22 648

  8. [6]

    Squad: 100,000+ questions for machine comprehension of text,

    P . Rajpurkar, J. Zhang, K. Lopyrev, and P . Liang, “Squad: 100,000+ questions for machine comprehension of text,” in Proc. Conf. Empir. Methods Nat. Lang. Process. , 2016

Show all 115 references
  1. [7]

    The penn treebank: annotating predicate argument structure,

    M. Marcus, G. Kim, M. A. Marcinkiewicz, R. MacIntyre, A. Bies, M. Ferguson, K. Katz, and B. Schasberger, “The penn treebank: annotating predicate argument structure,” in Proceedings of the Workshop on Human Language T echnology, 1994, p. 114–119

  2. [8]

    Glue: A multi-task benchmark and analysis platform for natural language understanding,

    A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, “Glue: A multi-task benchmark and analysis platform for natural language understanding,” in Proc. Int. Conf. Learn. Represent. , 2019

  3. [9]

    Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks,

    Y. Wang, S. Mishra, P . Alipoormolabashi, Y. Kordi, A. Mirzaei, A. Arunkumar, A. Ashok, A. S. Dhanasekaran, A. Naik, D. Stap et al., “Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks,” in Proc. Conf. Empir. Methods Nat. Lang. Process., 2022

  4. [10]

    Piqa: Reasoning about physical commonsense in natural language,

    Y. Bisk, R. Zellers, R. L. Bras, J. Gao, and Y. Choi, “Piqa: Reasoning about physical commonsense in natural language,” in Proc. AAAI Conf. Artif. Intell. , 2020

  5. [11]

    Hellaswag: Can a machine really finish your sentence?

    R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi, “Hellaswag: Can a machine really finish your sentence?” in Proc. Annu. Meeting Assoc. Comput. Linguistics , 2019

  6. [12]

    Wino- grande: An adversarial winograd schema challenge at scale,

    K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi, “Wino- grande: An adversarial winograd schema challenge at scale,” CoRR, 2019

  7. [13]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. Adv. Neural Inform. Process. Syst. , vol. 30, 2017

  8. [14]

    Opt: Open pre-trained transformer language models,

    S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V . Lin et al., “Opt: Open pre-trained transformer language models,” arXiv preprint arXiv:2205.01068 , 2022

  9. [15]

    Bloom: A 176b-parameter open-access multilingual language model,

    B. Workshop, T. L. Scao, A. Fan, C. Akiki, E. Pavlick, S. Ili ´c, D. Hesslow, R. Castagn´e, A. S. Luccioni, F. Yvon et al., “Bloom: A 176b-parameter open-access multilingual language model,” arXiv preprint arXiv:2211.05100, 2022

  10. [16]

    Glm-130b: An open bilingual pre-trained model,

    A. Zeng, X. Liu, Z. Du, Z. Wang, H. Lai, M. Ding, Z. Yang, Y. Xu, W. Zheng, X. Xia et al., “Glm-130b: An open bilingual pre-trained model,” in Proc. Int. Conf. Learn. Represent. , 2023

  11. [17]

    Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model,

    S. Smith, M. Patwary, B. Norick, P . LeGresley, S. Rajbhandari, J. Casper, Z. Liu, S. Prabhumoye, G. Zerveas, V . Korthikantiet al., “Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model,” arXiv preprint arXiv:2201.11990, 2022

  12. [18]

    Llama 2: Open foundation and fine-tuned chat models,

    H. Touvron, L. Martin, K. Stone, P . Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P . Bhargava, S. Bhosale et al. , “Llama 2: Open foundation and fine-tuned chat models,” arXiv preprint arXiv:2307.09288, 2023

  13. [19]

    The falcon series of open language models,

    E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojocaru, M. Debbah, ´E. Goffinet, D. Hesslow, J. Launay, Q. Malartic et al., “The falcon series of open language models,” arXiv preprint arXiv:2311.16867, 2023

  14. [20]

    Mixtral of experts,

    A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand et al., “Mixtral of experts,” arXiv preprint arXiv:2401.04088 , 2024

  15. [21]

    Bert: Pre- training of deep bidirectional transformers for language under- standing,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre- training of deep bidirectional transformers for language under- standing,” in Proc. 2019 NAACL - HLT, Vol. 1 (Long & Short Papers) , 2019, pp. 4171–4186

  16. [22]

    Roberta: A ro- bustly optimized bert pretraining approach,

    Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V . Stoyanov, “Roberta: A ro- bustly optimized bert pretraining approach,” arXiv preprint arXiv:1907.11692, 2019

  17. [23]

    Xlnet: Generalized autoregressive pretraining for language understanding,

    Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V . Le, “Xlnet: Generalized autoregressive pretraining for language understanding,” in Proc. Adv. Neural Inform. Process. Syst. , vol. 32, 2019

  18. [24]

    Gpt-j-6b: A 6 billion parameter autoregressive language model,

    B. Wang and A. Komatsuzaki, “Gpt-j-6b: A 6 billion parameter autoregressive language model,” 2021

  19. [25]

    Gpt-neox-20b: An open-source autoregressive language model,

    S. Black, S. Biderman, E. Hallahan, Q. Anthony, L. Gao, L. Golding, H. He, C. Leahy, K. McDonell, J. Phang et al. , “Gpt-neox-20b: An open-source autoregressive language model,” arXiv preprint arXiv:2204.06745, 2022

  20. [26]

    Mistral 7b,

    A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P . Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed, “Mistral 7b,” arXiv preprint arXiv:2310.06...

  21. [27]

    Starcoder: may the source be with you!

    R. Li, L. B. Allal, Y. Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chim, Q. Liu, E. Zheltonozhskii, T. Y. Zhuo, T. Wang, O. Dehaene, M. Davaadorj, J. Lamy-Poirier, J. Monteiro, O. Shliazhko, N. Gontier, N. Meade, A. Zebaze, M.-H. Yee, L. K. Umapathi...

  22. [28]

    Gemma: Open models based on gemini research and technology,

    G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivi `ere, M. S. Kale, J. Love et al., “Gemma: Open models based on gemini research and technology,” arXiv preprint arXiv:2403.08295, 2024

  23. [29]

    Training language models to follow instructions with human feedback,

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P . Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P . Welinder, P . Christiano, J. Leike, and R. Lowe, “Training language models to follow instructi...

  24. [30]

    Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,

    Z. Yao, R. Y. Aminabadi, M. Zhang, X. Wu, C. Li, and Y. He, “Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,” in Proc. Adv. Neural Inform. Process. Syst. , vol. 35, 2022, pp. 27 168–27 183

  25. [31]

    Smoothquant: Accurate and efficient post-training quantization for large language models,

    G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “Smoothquant: Accurate and efficient post-training quantization for large language models,” in Proc. Int. Conf. Mach. Learn. , 2023

  26. [32]

    Squeezellm: Dense-and-sparse quanti- zation,

    S. Kim, C. Hooper, A. Gholami, Z. Dong, X. Li, S. Shen, M. W. Mahoney, and K. Keutzer, “Squeezellm: Dense-and-sparse quanti- zation,” in Proc. Int. Conf. Mach. Learn. , 2024

  27. [33]

    Q-bert: Hessian based ultra low precision quantization of bert,

    S. Shen, Z. Dong, J. Ye, L. Ma, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer, “Q-bert: Hessian based ultra low precision quantization of bert,” in Proc. AAAI Conf. Artif. Intell. , 2020

  28. [34]

    Qlora: Efficient finetuning of quantized llms,

    T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “Qlora: Efficient finetuning of quantized llms,” in Proc. Adv. Neural Inform. Process. Syst., 2023

  29. [35]

    Gptq: Accurate post-training quantization for generative pre-trained transformers,

    E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh, “Gptq: Accurate post-training quantization for generative pre-trained transformers,” in Proc. Int. Conf. Learn. Represent. , 2023

  30. [36]

    Spqr: A sparse-quantized representation for near-lossless llm weight compression,

    T. Dettmers, R. Svirschevski, V . Egiazarian, D. Kuznedelev, E. Fran- tar, S. Ashkboos, A. Borzunov, T. Hoefler, and D. Alistarh, “Spqr: A sparse-quantized representation for near-lossless llm weight compression,” in Proc. Int. Conf. Learn. Represent. , 2024

  31. [37]

    Awq: Activation-aware weight quantization for llm compression and acceleration,

    J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han, “Awq: Activation-aware weight quantization for llm compression and acceleration,” in Proc. Mach. Learn. Syst. Conf. , 2024

  32. [38]

    Training with quantization noise for extreme model compression,

    A. Fan, P . Stock, B. Graham, E. Grave, R. Gribonval, H. Jegou, and A. Joulin, “Training with quantization noise for extreme model compression,” in Proc. Int. Conf. Learn. Represent. , 2021

  33. [39]

    Obelics: An open web-scale filtered dataset of interleaved image-text documents,

    H. Lauren c ¸on, L. Saulnier, L. Tronchon, S. Bekman, A. Singh, A. Lozhkov, T. Wang, S. Karamcheti, A. M. Rush, D. Kiela, M. Cord, and V . Sanh, “Obelics: An open web-scale filtered dataset of interleaved image-text documents,” in Proc. Adv. Neural Inform. Process. Syst., 2023

  34. [40]

    Laion-5b: An open large-scale dataset for training next generation image-text models,

    C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P . Schramowski, S. Kundurthy, K. Crowson, L. Schmidt, R. Kacz- marczyk, and J. Jitsev, “Laion-5b: An open large-scale dataset for training next generation ima...

  35. [41]

    Wit: Wikipedia-based image text dataset for multimodal multilin- gual machine learning,

    K. Srinivasan, K. Raman, J. Chen, M. Bendersky, and M. Najork, “Wit: Wikipedia-based image text dataset for multimodal multilin- gual machine learning,” in Proc. Int. ACM SIGIR Conf. Res. Dev. Inf. Retrieval, 2021, pp. 2443–2449

  36. [42]

    Gqa: A new dataset for real- world visual reasoning and compositional question answering,

    D. A. Hudson and C. D. Manning, “Gqa: A new dataset for real- world visual reasoning and compositional question answering,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 6700–6709

  37. [43]

    Towards vqa models that can read,

    A. Singh, V . Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards vqa models that can read,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 8317–8326

  38. [44]

    Learn to explain: Multimodal reasoning via thought chains for science question answering,

    P . Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P . Clark, and A. Kalyan, “Learn to explain: Multimodal reasoning via thought chains for science question answering,” in Proc. Adv. Neural Inform. Process. Syst. , vol. 35, 2022, pp. 2507–2521

  39. [45]

    Vizwiz grand challenge: Answering visual questions from blind people,

    D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P . Bigham, “Vizwiz grand challenge: Answering visual questions from blind people,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 3608–3617

  40. [46]

    Visual instruction tuning,

    H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Proc. Adv. Neural Inform. Process. Syst. , vol. 36, 2023, pp. 34 892– 34 916

  41. [47]

    Openflamingo: An open-source framework for training large autoregressive vision-language models,

    A. Awadalla, I. Gao, J. Gardner, J. Hessel, Y. Hanafy, W. Zhu, K. Marathe, Y. Bitton, S. Gadre, S. Sagawa, J. Jitsev, S. Korn- blith, P . W. Koh, G. Ilharco, M. Wortsman, and L. Schmidt, “Openflamingo: An open-source framework for training large autoregressive vision-language ...

  42. [48]

    Noisyquant: Noisy bias-enhanced post-training activation quanti- zation for vision transformers,

    Y. Liu, H. Yang, Z. Dong, K. Keutzer, L. Du, and S. Zhang, “Noisyquant: Noisy bias-enhanced post-training activation quanti- zation for vision transformers,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog., 2023

  43. [49]

    Imagenet large scale visual recognition challenge,

    O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al., “Imagenet large scale visual recognition challenge,” International journal of computer vision, vol. 115, pp. 211–252, 2015

  44. [50]

    Learning multiple layers of features from tiny images,

    A. Krizhevsky, G. Hinton et al. , “Learning multiple layers of features from tiny images,” T oronto, ON, Canada, 2009

  45. [51]

    Microsoft coco captions: Data collection and evaluation server,

    X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P . Doll ´ar, and C. L. Zitnick, “Microsoft coco captions: Data collection and evaluation server,” arXiv preprint arXiv:1504.00325 , 2015

  46. [52]

    Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop,

    F. Yu, A. Seff, Y. Zhang, S. Song, T. Funkhouser, and J. Xiao, “Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop,” arXiv preprint arXiv:1506.03365 , 2016

  47. [53]

    Mak- ing the v in vqa matter: Elevating the role of image understanding in visual question answering,

    Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Mak- ing the v in vqa matter: Elevating the role of image understanding in visual question answering,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 6904–6913

  48. [54]

    Improved denoising diffusion probabilistic models,

    A. Q. Nichol and P . Dhariwal, “Improved denoising diffusion probabilistic models,” in Proc. Int. Conf. Mach. Learn. , 2021, pp. 8162–8171

  49. [55]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P . Abbeel, “Denoising diffusion probabilistic models,” in Proc. Adv. Neural Inform. Process. Syst. , vol. 33, 2020, pp. 6840–6851

  50. [56]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 770–778

  51. [57]

    Mobilenetv2: Inverted residuals and linear bottlenecks,

    M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 4510–4520

  52. [58]

    A style-based generator archi- tecture for generative adversarial networks,

    T. Karras, S. Laine, and T. Aila, “A style-based generator archi- tecture for generative adversarial networks,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 4401–4410

  53. [59]

    Vila: On pre-training for visual language models,

    J. Lin, H. Yin, W. Ping, P . Molchanov, M. Shoeybi, and S. Han, “Vila: On pre-training for visual language models,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2024, pp. 26 689–26 699

  54. [60]

    Swin transformer: Hierarchical vision transformer using shifted windows,

    Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2021

  55. [61]

    Qdrop: Randomly dropping quantization for extremely low-bit post-training quanti- zation,

    X. Wei, R. Gong, Y. Li, X. Liu, and F. Yu, “Qdrop: Randomly dropping quantization for extremely low-bit post-training quanti- zation,” in Proc. Int. Conf. Learn. Represent. , 2022

  56. [62]

    Post-training quantization on diffusion models,

    Y. Shang, Z. Yuan, B. Xie, B. Wu, and Y. Yan, “Post-training quantization on diffusion models,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog., 2023

  57. [63]

    Q-DM: An efficient low-bit quantized diffusion model,

    Y. Li, S. Xu, X. Cao, X. Sun, and B. Zhang, “Q-DM: An efficient low-bit quantized diffusion model,” in Proc. Adv. Neural Inform. Process. Syst., 2023

  58. [64]

    Accurate post training quantization with small calibration sets,

    I. Hubara, Y. Nahshan, Y. Hanani, R. Banner, and D. Soudry, “Accurate post training quantization with small calibration sets,” in Proc. Int. Conf. Mach. Learn. , 2021

  59. [65]

    Ntire 2017 challenge on single image super-resolution: Dataset and study,

    E. Agustsson and R. Timofte, “Ntire 2017 challenge on single image super-resolution: Dataset and study,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 126–135

  60. [66]

    Component divide-and-conquer for real-world image super- resolution,

    P . Wei, Z. Xie, H. Lu, Z. Zhan, Q. Ye, W. Zuo, and L. Lin, “Component divide-and-conquer for real-world image super- resolution,” in Proc. Eur. Conf. Comput. Vis. Springer, 2020, pp. 101–117

  61. [67]

    Low-complexity single-image super-resolution based on nonnegative neighbor embedding,

    M. Bevilacqua, A. Roumy, C. Guillemot, and M. L. Alberi- Morel, “Low-complexity single-image super-resolution based on nonnegative neighbor embedding,” in Proc. IEEE Int. Conf. Acoust. Speech Signal Process. , 2012

  62. [68]

    On single image scale-up using sparse-representations,

    R. Zeyde, M. Elad, and M. Protter, “On single image scale-up using sparse-representations,” in Proc. Int. Conf. Curves Surfaces , 2010, pp. 711–730

  63. [69]

    A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics,

    D. Martin, C. Fowlkes, D. Tal, and J. Malik, “A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics,” in Proc. Int. Conf. Comput. Vis. , vol. 2, 2001, pp. 416–423

  64. [70]

    Single image super- resolution from transformed self-exemplars,

    J.-B. Huang, A. Singh, and N. Ahuja, “Single image super- resolution from transformed self-exemplars,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2015, pp. 5197–5206

  65. [71]

    Accurate image super-resolution using very deep convolutional networks,

    J. Kim, J. K. Lee, and K. M. Lee, “Accurate image super-resolution using very deep convolutional networks,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 1646–1654. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 5

  66. [72]

    Enhanced deep residual networks for single image super-resolution,

    B. Lim, S. Son, H. Kim, S. Nah, and K. Mu Lee, “Enhanced deep residual networks for single image super-resolution,” inProc. IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 136–144

  67. [73]

    Residual dense network for image super-resolution,

    Y. Zhang, Y. Tian, Y. Kong, B. Zhong, and Y. Fu, “Residual dense network for image super-resolution,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 2472–2481

  68. [74]

    Photo- realistic single image super-resolution using a generative adver- sarial network,

    C. Ledig, L. Theis, F. Husz ´ar, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang et al. , “Photo- realistic single image super-resolution using a generative adver- sarial network,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 4681–4690

  69. [75]

    Daq: Channel-wise distribution-aware quantization for deep image super-resolution networks,

    C. Hong, H. Kim, S. Baik, J. Oh, and K. M. Lee, “Daq: Channel-wise distribution-aware quantization for deep image super-resolution networks,” in Proc. IEEE Winter Conf. Appl. Comput. Vis. , 2022, pp. 2675–2684

  70. [76]

    Searching for low-bit weights in quantized neural networks,

    Z. Yang, Y. Wang, K. Han, C. Xu, C. Xu, D. Tao, and C. Xu, “Searching for low-bit weights in quantized neural networks,” in Proc. Adv. Neural Inform. Process. Syst. , 2020

  71. [77]

    Learnable lookup table for neural network quantization,

    L. Wang, X. Dong, Y. Wang, L. Liu, W. An, and Y. Guo, “Learnable lookup table for neural network quantization,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2022, pp. 12 423–12 433

  72. [78]

    Encoder-decoder with atrous separable convolution for semantic image segmentation,

    L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder-decoder with atrous separable convolution for semantic image segmentation,” in Proc. Eur. Conf. Comput. Vis. , 2018, pp. 801–818

  73. [79]

    Repvgg: Making vgg-style convnets great again,

    X. Ding, X. Zhang, N. Ma, J. Han, G. Ding, and J. Sun, “Repvgg: Making vgg-style convnets great again,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2021, pp. 13 733–13 742

  74. [80]

    Up or down? adaptive rounding for post-training quantization,

    M. Nagel, R. A. Amjad, M. van Baalen, C. Louizos, and T. Blankevoort, “Up or down? adaptive rounding for post-training quantization,” in Proc. Int. Conf. Mach. Learn. , 2020

  75. [81]

    Designing network design spaces,

    I. Radosavovic, R. P . Kosaraju, R. Girshick, K. He, and P . Doll´ar, “Designing network design spaces,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 10 428–10 436

  76. [82]

    Mnasnet: Platform-aware neural architecture search for mobile,

    M. Tan, B. Chen, R. Pang, V . Vasudevan, M. Sandler, A. Howard, and Q. V . Le, “Mnasnet: Platform-aware neural architecture search for mobile,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 2820–2828

  77. [83]

    Very deep convolutional net- works for large-scale image recognition,

    K. Simonyan and A. Zisserman, “Very deep convolutional net- works for large-scale image recognition,” in Proc. Int. Conf. Learn. Represent., 2015

  78. [84]

    Rethinking the inception architecture for computer vision,

    C. Szegedy, V . Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 2818–2826

  79. [85]

    Brecq: Pushing the limit of post-training quantization by block reconstruction,

    Y. Li, R. Gong, X. Tan, Y. Yang, P . Hu, Q. Zhang, F. Yu, W. Wang, and S. Gu, “Brecq: Pushing the limit of post-training quantization by block reconstruction,” in Proc. Int. Conf. Learn. Represent. , 2021

  80. [86]

    Hawq: Hessian aware quantization of neural networks with mixed-precision,

    Z. Dong, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer, “Hawq: Hessian aware quantization of neural networks with mixed-precision,” in Proc. Int. Conf. Comput. Vis., 2019, pp. 293–302

  81. [87]

    Overcoming oscillations in quantization-aware training,

    M. Nagel, M. Fournarakis, Y. Bondarenko, and T. Blankevoort, “Overcoming oscillations in quantization-aware training,” in Proc. Int. Conf. Mach. Learn. , 2022

  82. [88]

    Focal loss for dense object detection,

    T.-Y. Lin, P . Goyal, R. Girshick, K. He, and P . Doll´ar, “Focal loss for dense object detection,” in Proc. Int. Conf. Comput. Vis. , 2017, pp. 2980–2988

  83. [89]

    A comprehensive study on post-training quantization for large language models,

    Z. Yao, C. Li, X. Wu, S. Youn, and Y. He, “A comprehensive study on post-training quantization for large language models,” arXiv preprint arXiv:2303.08302, 2023

  84. [90]

    A comprehen- sive survey on model quantization for deep neural networks in image classification,

    B. Rokh, A. Azarpeyvand, and A. Khanteymoori, “A comprehen- sive survey on model quantization for deep neural networks in image classification,” ACM T rans. Intell. Syst. T echnol., vol. 14, no. 6, pp. 1–50, 2023

  85. [91]

    A survey of quantization methods for efficient neural network inference,

    A. Gholami, S. Kim, Z. Dong, Z. Yao, M. W. Mahoney, and K. Keutzer, “A survey of quantization methods for efficient neural network inference,” in Low-power computer vision . Chapman and Hall/CRC, 2022, pp. 291–326

  86. [92]

    Evaluating quantized large language models,

    S. Li, X. Ning, L. Wang, T. Liu, X. Shi, S. Yan, G. Dai, H. Yang, and Y. Wang, “Evaluating quantized large language models,” arXiv preprint arXiv:2402.18158, 2024

  87. [93]

    Exploiting llm quantization,

    K. Egashira, M. Vero, R. Staab, J. He, and M. Vechev, “Exploiting llm quantization,” arXiv preprint arXiv:2405.18137 , 2024

  88. [94]

    Exploring post-training quantization in llms from comprehensive study to low rank compensation,

    Z. Yao, X. Wu, C. Li, S. Youn, and Y. He, “Exploring post-training quantization in llms from comprehensive study to low rank compensation,” in Proc. AAAI Conf. Artif. Intell. , vol. 38, no. 17, 2024, pp. 19 377–19 385

  89. [95]

    Exploring quantization techniques for large-scale language models: Methods, challenges and future directions,

    A. Shen, Z. Lai, and D. Li, “Exploring quantization techniques for large-scale language models: Methods, challenges and future directions,” in Proc. Int. Conf. Cyber Secur. Inf. Eng. , 2024, pp. 783–790

  90. [96]

    The case for 4-bit precision: k-bit inference scaling laws,

    T. Dettmers and L. Zettlemoyer, “The case for 4-bit precision: k-bit inference scaling laws,” in Proc. Int. Conf. Mach. Learn. PMLR, 2023, pp. 7750–7774

  91. [97]

    Only train once: A one-shot neural network training and pruning framework,

    T. Chen, B. Ji, T. Ding, B. Fang, G. Wang, Z. Zhu, L. Liang, Y. Shi, S. Yi, and X. Tu, “Only train once: A one-shot neural network training and pruning framework,” in Proc. Adv. Neural Inform. Process. Syst., 2021

  92. [98]

    Deephoyer: Learning sparser neural network with differentiable scale-invariant sparsity measures,

    H. Yang, W. Wen, and H. Li, “Deephoyer: Learning sparser neural network with differentiable scale-invariant sparsity measures,” in Proc. Int. Conf. Learn. Represent. , 2020

  93. [99]

    Towards compact convnets via structure-sparsity regularized filter pruning,

    S. Lin, R. Ji, Y. Li, C. Deng, and X. Li, “Towards compact convnets via structure-sparsity regularized filter pruning,” arXiv preprint arXiv:1901.07827, 2019

  94. [100]

    Accelerate cnn via recursive bayesian pruning,

    Y. Zhou, Y. Zhang, Y. Wang, and Q. Tian, “Accelerate cnn via recursive bayesian pruning,” in Proc. Int. Conf. Comput. Vis. , 2019, pp. 3306–3315

  95. [101]

    Bayesian bits: Unifying quan- tization and pruning,

    M. van Baalen, C. Louizos, M. Nagel, R. A. Amjad, Y. Wang, T. Blankevoort, and M. Welling, “Bayesian bits: Unifying quan- tization and pruning,” in Proc. Adv. Neural Inform. Process. Syst. , 2020

  96. [102]

    Otov2: Auto- matic, generic, user-friendly,

    T. Chen, L. Liang, T. Ding, Z. Zhu, and I. Zharkov, “Otov2: Auto- matic, generic, user-friendly,” in Proc. Int. Conf. Learn. Represent. , 2023

  97. [103]

    Eagleeye: Fast sub-net evaluation for efficient neural network pruning,

    B. Li, B. Wu, J. Su, G. Wang, and L. Lin, “Eagleeye: Fast sub-net evaluation for efficient neural network pruning,” in Proc. Eur. Conf. Comput. Vis., 2020

  98. [106]

    Linear mode connectivity and the lottery ticket hypothesis,

    J. Frankle, G. K. Dziugaite, D. M. Roy, and M. Carbin, “Linear mode connectivity and the lottery ticket hypothesis,” in Proc. Int. Conf. Mach. Learn. , 2020

  99. [107]

    Neuron-level structured pruning using polarization regularizer,

    Z. Tao, Z. Zhang, Y. Huang, X. Zeng, K. Shuang, and X. Li, “Neuron-level structured pruning using polarization regularizer,” in Proc. Adv. Neural Inform. Process. Syst. , 2020

  100. [108]

    Pruning filters for efficient convnets,

    H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P . Graf, “Pruning filters for efficient convnets,” in Proc. Int. Conf. Learn. Represent. , 2017

  101. [109]

    Learning both weights and connections for efficient neural networks,

    S. Han, J. Pool, J. Tran, and W. J. Dally, “Learning both weights and connections for efficient neural networks,” in Proc. Adv. Neural Inform. Process. Syst. , vol. 28, 2015

  102. [110]

    Distilling the knowledge in a neural network,

    G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” in Proc. Adv. Neural Inform. Process. Syst. , 2014

  103. [111]

    Temporal separation with entropy regularization for knowledge distillation in spiking neural networks,

    K. Yu, C. Yu, T. Zhang, X. Zhao, S. Yang, H. Wang, Q. Zhang, and Q. Xu, “Temporal separation with entropy regularization for knowledge distillation in spiking neural networks,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2025

  104. [112]

    Dkdm: Data-free knowledge distillation for diffusion models with any architecture,

    Q. Xiang, M. Zhang, Y. Shang, J. Wu, Y. Yan, and L. Nie, “Dkdm: Data-free knowledge distillation for diffusion models with any architecture,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2025

  105. [113]

    Customkd: Customizing large vision foundation for edge model improvement via knowledge distillation,

    J. Lee, D. Das, M. Hayat, S. Choi, K. Hwang, and F. Porikli, “Customkd: Customizing large vision foundation for edge model improvement via knowledge distillation,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2025

  106. [114]

    Cat: Cross attention in vision transformer,

    H. Lin, X. Cheng, X. Wu, F. Yang, D. Shen, Z. Wang, Q. Song, and W. Yuan, “Cat: Cross attention in vision transformer,” Proc. IEEE Int. Conf. Multimedia Expo , 2021

  107. [115]

    Vision transformer with deformable attention,

    Z. Xia, X. Pan, S. Song, L. E. Li, and G. Huang, “Vision transformer with deformable attention,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog., 2022

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.