REVIEW 3 major objections 6 minor 63 references
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper claims that 2-bit quantization errors, although high-rank, can be offset by small LoRA adapters if the adapters are trained against the whole model's output rather than layer-by-layer.
desk verdict RILQ is a useful empirical recipe for 2-bit LoRA-based quantization error compensation, but the paper's central rank-insensitivity story is not actually isolated in the experiments. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the model-wise activation discrepancy loss, called Model-Loss, defined as $\|Y_N - Y_N^q\|_F$, the Frobenius norm between the output activations of the final Transformer layer of the full-precision model and the quantized model with LoRA adapters. This loss performs the argument's load-bearing work: because it is applied after all Transformer layers, errors in individual layers are free to drift as long as the final output aligns, which makes the loss rank-insensitive and lets adapters with rank as low as 16 compensate errors created by 2-bit weights. The paper measures rank sensitivity through the relative error $E = |(Y - Y^q)/Y|$ at the LM-head output and supports the cooperative-compensation explanation by showing that Model-Loss raises the singular-vector magnitudes of rank-critical FFN1 modules relative to Q-Proj.
What would settle it
Quantize a held-out model not used in calibration, such as Mistral-7B, to 2-bit, train adapters with Model-Loss at rank 16 and rank 256, and compare the LM-head relative error; if the rank-16 error is substantially worse than rank-256 (comparable to the SVD spread in Fig. 4(a)), the central claim is wrong.
Extended reading notes
Core claim
The paper's central claim is that the rank requirement of LQEC is a property of the discrepancy scope, not of the quantization error alone. When the discrepancy loss is defined at a single linear module (Linear-Loss) or a single Transformer layer (Layer-Loss), 2-bit quantization error demands high-rank adapters, contradicting LoRA's low-rank premise. When the loss is defined at the output activation of the final Transformer layer (Model-Loss, $\|Y_N - Y_N^q\|_F$), the relative error at the LM-head stays low even at rank 16, because internal activation drift is allowed to happen freely and the adapters cooperate across layers to align the final output. The paper interprets this as balancing rank-critical modules like FFN1 against rank-redundant modules like Q-Proj, and combines Model-Loss with a causal language modeling loss (GT-Loss) for further gains.
Load-bearing premise
The whole method depends on the assumption that what the paper observes on LLaMA-2-7B with one calibration setup—that the model-level loss keeps working at rank 16—also holds on other models and quantizers, and that low output error really means better answers.
Editorial extensions
If this is right
- RILQ makes low-rank LQEC practical at 2-bit precision, so serving a 2-bit quantized LLM can regain most of the accuracy lost to quantization while keeping the memory savings of 2-bit weights.
- Because RILQ works as initialization, it can replace SVD-based initialization in LoRA fine-tuning pipelines, improving downstream task accuracy for models that are already quantized to 2 bits.
- The rank-insensitivity result implies that rank selection for LQEC becomes far less critical: rank 16 nearly matches rank 256, removing the need for per-module rank allocation.
- The gains transfer across different quantizers (OmniQuant, QuIP#, QuaRot) and both LLaMA-2 and LLaMA-3, indicating the effect depends on the loss scope rather than on a specific quantization scheme.
- Since the adapters can be merged into quantized weights, RILQ adds no inference-time overhead relative to adapter-less quantized inference, including in the QA-LoRA setting.
Reading between the lines
- The output-level compensation principle could generalize to other high-rank structured perturbations beyond 2-bit weights, such as aggressive pruning or extreme activation quantization, where a model-wide loss might again remove the need for high-rank corrections.
- The internal-drift observation suggests a testable design principle for low-rank error compensation: allow intermediate feature mismatch as long as the final output is aligned, which is the opposite of layer-wise distillation objectives.
- If the effect is a general property of transformer error propagation, then the optimal loss scope may scale with model depth, implying that larger models need even broader compensation scopes—this could be checked by measuring rank sensitivity across 7B, 13B, and 70B models.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses why LoRA-based quantization error compensation (LQEC) fails for 2-bit weight quantization and proposes RILQ, a method that trains low-rank adapters by minimizing a model-wise activation discrepancy loss (Model-Loss, Eq. 5) together with a ground-truth causal language modeling loss (GT-Loss, Eq. 6). The authors first present a rank sensitivity analysis showing that the relative error of LM-head outputs grows with rank as discrepancy scope narrows, and that the model-level loss is less rank-sensitive than linear- or layer-level losses. They then evaluate RILQ on LLaMA-2-7B and LLaMA-3-8B with OmniQuant, QuIP#, QuaRot, and LoftQ, reporting consistent perplexity and zero-shot CSQA/GSM8K improvements at 2-bit, plus gains in task-specific fine-tuning and QA-LoRA settings. Ablations examine rank sensitivity, loss scope, GT-Loss contribution, calibration size, and model scale up to 70B.
Significance. If the central claim holds, the paper makes a practically useful contribution: it shows that a simple final-layer output discrepancy objective can make 2-bit LQEC effective with rank-16 adapters, enabling adapter-merged weight-quantized inference without large accuracy loss. The strengths of the paper are its broad experimental coverage (two model families, four quantizers, direct correction and fine-tuning, QA-LoRA integration, model scales up to 70B) and the promised public code. The main weakness is that the causal attribution of rank-insensitivity to Model-Loss is not fully isolated from the auxiliary GT-Loss and from the larger optimization budget. Because RILQ is defined as Model-Loss plus GT-Loss, and Table 7 shows GT-Loss alone already improves accuracy, the rank-sweep evidence in Table 5 does not by itself establish the proposed mechanism. This underdetermination limits, but does not destroy, the significance of the empirical gains; the issue is addressable with targeted ablations.
major comments (3)
- [Sec. 'Rank-Insensitive LQEC'; Eqs. (5)-(6); Tables 5 and 7] The central claim is that Model-Loss mitigates the high-rank requirements of 2-bit quantization error and that this is what makes RILQ rank-insensitive. However, the rank-sweep evidence in Table 5 (sigma = 0.07 for RILQ vs 0.69 for SVD at W2A16) is obtained with RILQ, which is the combination of Model-Loss and GT-Loss. Table 7 shows that GT-Loss alone reaches 53.28 average CSQA accuracy, nearly matching Model-Loss alone (52.77) at a single rank, but Table 7 does not report a rank sweep. Consequently, the observed flatness across ranks could in principle be driven by GT-Loss, by the interaction of the two losses, or by the substantially larger optimization budget (up to 10,000 Adam steps with early stopping) rather than by the wider discrepancy scope per se. Please provide rank sweeps (e.g., ranks 16, 32, 64, 128, 256) for Model-Loss alone, GT-Loss alone, Model-Loss+GT-Loss, and the SVD baseline, with the same optimization budget and early-stopping criterion, on at least LLaMA-2-7B with OmniQuant and one additional quantizer. In addition, state explicitly which objective was used to produce Fig. 4(a) and Fig. 4(b); if those figures use Model-Loss alone, say so and reconcile with Table 5, which uses RILQ.
- [Sec. 'Rank Sensitivity Analysis'; Fig. 4] The rank sensitivity analysis that motivates the central claim is performed on a single model (LLaMA-2-7B) with a single quantizer (OmniQuant for Fig. 4). The paper's broader conclusion, that Model-Loss is rank-insensitive for 2-bit LQEC, is used to justify the method across LLaMA-3-8B, QuIP#, QuaRot, and LoftQ. Given that Table 1 shows very different error magnitudes across models and quantizers (e.g., LoftQ W2A16 perplexity is 56168 on LLaMA-3-8B vs 1078 on LLaMA-2-7B), the rank-sensitivity phenomenon could be model- or quantizer-specific. Please replicate the rank sensitivity comparison (relative error E vs rank) for LLaMA-3-8B and for at least one additional 2-bit quantizer (e.g., QuIP# or QuaRot), or explicitly qualify the generalizability of the proposed mechanism.
- [Tables 1, 7, and 9] All reported accuracy and perplexity numbers appear to come from single runs without error bars or multiple seeds. RILQ involves stochastic optimization (Adam, random calibration sampling of 256 C4 sentences, early stopping), and several reported differences are small in absolute terms: for example, in Table 7 the difference between GT-Loss alone (53.28) and Model-Loss alone (52.77) is 0.51 points, and the difference between Model-Loss alone and RILQ (54.71) is 1.94 points. Without variance estimates, it is not possible to tell which of these differences are significant. Please report mean and standard deviation over at least three seeds for the central tables (in particular Table 1, Table 7, and the rank sweep in Table 5), or provide a justification for why run-to-run variance is negligible for this setup.
minor comments (6)
- [Sec. 'Rank Sensitivity Analysis'] The relative error metric E = |(Y - Y^q)/Y| is not defined precisely; please specify the norm, the averaging (over tokens, batch, and hidden dimensions), and the exact computation procedure for the values plotted in Fig. 4(a).
- [Fig. 3(a) caption] The caption of Fig. 3(a) does not state the model or quantizer used; the text mentions LLaMA-2-7B, but the figure should include this information for self-containedness.
- [Table 3 and surrounding text] The layout of Table 3 is confusing: it is unclear whether 'RILQ Error Compensation' and 'Fine-Tuning' are separate rows or columns, and the text does not explain how the W2A16 OmniQuant baseline (47.42 CSQA) differs from the Table 1 OmniQuant result (51.88). Please clarify the experimental protocol and the relationship between the two tables.
- [Related Work and References] There is a citation inconsistency for LoftQ: it is attributed to (Guo et al. 2024) in one place in the Related Work section and to (Li et al. 2024) elsewhere; please ensure LoftQ is consistently cited as Li et al. 2024 and LQ-LoRA as Guo et al. 2024.
- [Appendix, 'Procedure of RILQ'] The appendix step 3 states that LoRA is initialized 'using gradient descent on Model-Loss (Eq. 5) and GT-Loss (Eq. 6)', which means RILQ always uses both losses; the terminology in the main text sometimes refers to RILQ and sometimes to Model-Loss alone, causing ambiguity. Please introduce distinct terms (e.g., 'Model-Loss-only' vs 'RILQ') and use them consistently.
- [Appendix, 'Memory Cost Analysis'] The comparison to Apple's accuracy-recovery adapter (Gunter et al. 2024) with '200x sample efficiency' is not substantiated with a precise token count or a direct comparison under matched conditions; please either add the necessary details or soften the claim.
Circularity Check
Rank-insensitivity claim reduces to the training objective: Model-Loss minimizes the LM-head output discrepancy used as the rank-sensitivity metric.
-
self definitional
[Section 'Rank Sensitivity Analysis'; definition of rank sensitivity E and Eq. (5) Model-Loss; Fig. 4(a)]
"We introduce a new metric called rank sensitivity , which measures the relative error ( E = |(Y − Y q)/Y |) of logits at the LM-Head output. ... We extend the discrepancy scope to encompass all Transformer layers, proposing a new discrepancy loss at the output activation of the final ( N’th) Transformer layer (Model-Loss, Fig. 2(e)): arg min L1,L2 ∥YN − Y q N ∥F . (5)"
The paper's evidence that Model-Loss is rank-insensitive is the rank-sensitivity metric E, which is the relative discrepancy at the LM-Head output. Model-Loss (Eq. 5) directly minimizes the final-layer output discrepancy that feeds that same metric: since the LM-Head is a fixed linear projection, the logit discrepancy measured by E is bounded by, and effectively proportional to, the minimized term. The observed decrease in E as the discrepancy scope widens to the model level is therefore a consequence of aligning the training objective with the evaluation metric, not an independent empirical discovery about 2-bit quantization errors.
full rationale
The paper's practical evaluations are not circular: RILQ is calibrated on C4 training sentences and then measured on held-out benchmarks (WikiText-2, C4 validation, and zero-shot CSQA tasks in Table 1), so the reported accuracy and perplexity improvements are genuine empirical results. The self-citations to TSLD (Kim et al. 2023a) and RA-LoRA (Kim et al. 2024) provide context and a supporting hypothesis about rank-critical modules, but the core comparisons do not reduce to those citations, so they are not load-bearing. The one genuine circular step is the rank-sensitivity analysis surrounding Eq. 5 and Fig. 4(a): the metric E is the relative LM-head output error, and Model-Loss minimizes exactly the final-layer output discrepancy that drives that error, making the 'rank-insensitive nature' of Model-Loss largely a definitional consequence of optimizing the measured quantity. This gives the central understanding claim a partial by-construction character. Independent content remains in the C4 perplexity flatness of Table 5 and the held-out task improvements of Table 1, although those tables use RILQ (Model-Loss plus GT-Loss with up to 10,000 steps), so the specific attribution of rank-insensitivity to Model-Loss alone is also confounded by the auxiliary GT-Loss and training budget; that confound is an internal-validity concern rather than a further circular reduction. Overall, the method is not forced end-to-end by its inputs, but one central 'discovery' does reduce to its own objective, warranting a partial-circularity score of 6.
Assumptions & free parameters
free parameters (4)
- rank =
64 (default; tested 16-256)
- loss weights =
0.5 for Model-Loss, 0.5 for GT-Loss
- calibration set size =
256 samples, sequence length 512 (default)
- learning rate =
1e-4
assumptions (4)
- domain assumption Pre-trained LLaMA-2 and LLaMA-3 weights are a representative base for evaluating quantization error compensation.
- domain assumption The C4 calibration set is representative enough that adapters fit on 256 sentences generalize to the evaluation benchmarks.
- ad hoc to paper The relative error metric E at the LM-head output is a valid proxy for token prediction accuracy.
- domain assumption The official implementations of baseline quantizers (LoftQ, OmniQuant, QuIP#, QuaRot) are correct and their settings are representative.
Cite this review
Pith. "Pith review of RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy." pith.science (2026). https://pith.science/paper/3BVTHADI
@misc{pith2026241201129,
author = {Pith},
title = {Pith review of: RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy},
year = {2026},
howpublished = {\url{https://pith.science/paper/3BVTHADI}},
note = {Machine review of arXiv:2412.01129}
}
read the original abstract
Low-rank adaptation (LoRA) has become the dominant method for parameter-efficient LLM fine-tuning, with LoRA-based quantization error compensation (LQEC) emerging as a powerful tool for recovering accuracy in compressed LLMs. However, LQEC has underperformed in sub-4-bit scenarios, with no prior investigation into understanding this limitation. We propose RILQ (Rank-Insensitive LoRA-based Quantization Error Compensation) to understand fundamental limitation and boost 2-bit LLM accuracy. Based on rank analysis revealing model-wise activation discrepancy loss's rank-insensitive nature, RILQ employs this loss to adjust adapters cooperatively across layers, enabling robust error compensation with low-rank adapters. Evaluations on LLaMA-2 and LLaMA-3 demonstrate RILQ's consistent improvements in 2-bit quantized inference across various state-of-the-art quantizers and enhanced accuracy in task-specific fine-tuning. RILQ maintains computational efficiency comparable to existing LoRA methods, enabling adapter-merged weight-quantized LLM inference with significantly enhanced accuracy, making it a promising approach for boosting 2-bit LLM performance. Our code is available at https://github.com/aiha-lab/RILQ.
Figures
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
L.; Li, B.; Jaggi, M.; Alistarh, D.; Hoefler, T.; and Hensman, J
Ashkboos, S.; Mohtashami, A.; Croci, M. L.; Li, B.; Jaggi, M.; Alistarh, D.; Hoefler, T.; and Hensman, J. 2024 a . QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs . arXiv
work page 2024
-
[4]
L.; Li, B.; Jaggi, M.; Alistarh, D.; Hoefler, T.; and Hensman, J
Ashkboos, S.; Mohtashami, A.; Croci, M. L.; Li, B.; Jaggi, M.; Alistarh, D.; Hoefler, T.; and Hensman, J. 2024 b . QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs. arXiv preprint arXiv:2404.00456
arXiv 2024
-
[5]
Bisk, Y.; Zellers, R.; Bras, R. L.; Gao, J.; and Choi, Y. 2019. PIQA: Reasoning about Physical Commonsense in Natural Language. arXiv:1911.11641
arXiv 2019
-
[6]
Chai, Y.; Gkountouras, J.; Ko, G. G.; Brooks, D.; and Wei, G.-Y. 2023. INT2.1: Towards Fine-Tunable Quantized Large Language Models with Error Correction through Low-Rank Adaptation . arXiv
work page 2023
-
[7]
Chee, J.; Cai, Y.; Kuleshov, V.; and Sa, C. D. 2023. QuIP: 2-Bit Quantization of Large Language Models With Guarantees . arXiv
work page 2023
-
[8]
Chen, J.; Zhang, A.; Shi, X.; Li, M.; Smola, A.; and Yang, D. 2023 a . Parameter-Efficient Fine-Tuning Design Spaces . arXiv
work page 2023
Show all 63 references
-
[9]
Chen, T.; Ding, T.; Yadav, B.; Zharkov, I.; and Liang, L. 2023 b . LoRAShear: Efficient Large Language Model Structured Pruning and Knowledge Recovery . arXiv
2023
-
[10]
Chen, Y.; Qian, S.; Tang, H.; Lai, X.; Liu, Z.; Han, S.; and Jia, J. 2024. LongLo RA : Efficient Fine-tuning of Long-Context Large Language Models. In The Twelfth International Conference on Learning Representations
2024
-
[11]
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018. Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge. arXiv:1803.05457v1
2018 arXiv
-
[12]
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; Hesse, C.; and Schulman, J. 2021. Training Verifiers to Solve Math Word Problems. arXiv:2110.14168
2021 arXiv
-
[13]
Dettmers, T.; Pagnoni, A.; Holtzman, A.; and Zettlemoyer, L. 2023. QL o RA : Efficient Finetuning of Quantized LLM s. In Thirty-seventh Conference on Neural Information Processing Systems
2023
-
[14]
Ding, N.; Qin, Y.; Yang, G.; Wei, F.; Yang, Z.; Su, Y.; Hu, S.; Chen, Y.; Chan, C.-M.; Chen, W.; Yi, J.; Zhao, W.; Wang, X.; Liu, Z.; Zheng, H.-T.; Chen, J.; Liu, Y.; Tang, J.; Li, J.; and Sun, M. 2022. Delta Tuning: A Comprehensive Study of Parameter Efficient Methods for Pre...
2022
-
[15]
Egiazarian, V.; Panferov, A.; Kuznedelev, D.; Frantar, E.; Babenko, A.; and Alistarh, D. 2024. Extreme Compression of Large Language Models via Additive Quantization . arXiv
2024
-
[16]
Frantar, E.; Ashkboos, S.; Hoefler, T.; and Alistarh, D. 2023. OPTQ : Accurate Quantization for Generative Pre-trained Transformers. In The Eleventh International Conference on Learning Representations
2023
-
[17]
Gao, L.; Tow, J.; Biderman, S.; Black, S.; DiPofi, A.; Foster, C.; Golding, L.; Hsu, J.; McDonell, K.; Muennighoff, N.; Phang, J.; Reynolds, L.; Tang, E.; Thite, A.; Wang, B.; Wang, K.; and Zou, A. 2021. A framework for few-shot language model evaluation
2021
-
[18]
Guan, Z.; Huang, H.; Su, Y.; Huang, H.; Wong, N.; and Yu, H. 2024. APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language Models . arXiv
2024
-
[19]
Gunter, T.; Wang, Z.; Wang, C.; Pang, R.; Narayanan, A.; Zhang, A.; Zhang, B.; Chen, C.; Chiu, C.-C.; Qiu, D.; et al. 2024. Apple intelligence foundation language models. arXiv preprint arXiv:2407.21075
2024
-
[20]
Guo, C.; Tang, J.; Hu, W.; Leng, J.; Zhang, C.; Yang, F.; Liu, Y.; Guo, M.; and Zhu, Y. 2023. OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization . arXiv
2023
-
[21]
Guo, H.; Greengard, P.; Xing, E.; and Kim, Y. 2024. LQ -Lo RA : Low-rank plus Quantized Matrix Decomposition for Efficient Language Model Finetuning. In The Twelfth International Conference on Learning Representations
2024
-
[22]
Han, Z.; Gao, C.; Liu, J.; Zhang, J.; and Zhang, S. Q. 2024. Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey . arXiv
2024
-
[23]
H.; Kim, J.; Kwon, B.; Kim, B.; Kwon, S
Heo, J. H.; Kim, J.; Kwon, B.; Kim, B.; Kwon, S. J.; and Lee, D. 2023. Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language Models . arXiv
2023
-
[24]
J.; yelong shen; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W
Hu, E. J.; yelong shen; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022. Lo RA : Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations
2022
-
[25]
K.-W.; Bing, L.; and Poria, S
Hu, Z.; Lan, Y.; Wang, L.; Xu, W.; Lim, E.-P.; Lee, R. K.-W.; Bing, L.; and Poria, S. 2023. LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models . arXiv
2023
-
[26]
Y.; Pang, T.; Du, C.; and Lin, M
Huang, C.; Liu, Q.; Lin, B. Y.; Pang, T.; Du, C.; and Lin, M. 2023. LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition . arXiv
2023
-
[27]
Huang, W.; Ma, X.; Qin, H.; Zheng, X.; Lv, C.; Chen, H.; Luo, J.; Qi, X.; Liu, X.; and Magno, M. 2024 a . How Good Are Low-bit Quantized LLaMA3 Models? An Empirical Study. arXiv:2404.14047
2024 arXiv
-
[28]
Huang, X.; Liu, Z.; Liu, S.-Y.; and Cheng, K.-T. 2024 b . RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization. arXiv:2407.08044
2024 arXiv
-
[29]
Kamalloo, E.; Dziri, N.; Clarke, C.; and Rafiei, D. 2023. Evaluating Open-Domain Question Answering in the Era of Large Language Models. In Rogers, A.; Boyd-Graber, J.; and Okazaki, N., eds., Proceedings of the 61st Annual Meeting of the Association for Computational Linguisti...
2023
-
[30]
Kim, M.; Lee, S.; Lee, J.; Hong, S.; Chang, D.-S.; Sung, W.; and Choi, J. 2023 a . Token-Scaled Logit Distillation for Ternary Weight Generative Language Models. In Thirty-seventh Conference on Neural Information Processing Systems
2023
-
[31]
Kim, M.; Lee, S.; Sung, W.; and Choi, J. 2024. RA - L o RA : Rank-Adaptive Parameter-Efficient Fine-Tuning for Accurate 2-bit Quantized Large Language Models. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Findings of the Association for Computational Linguistics ACL 2024,...
2024
-
[32]
W.; and Keutzer, K
Kim, S.; Hooper, C.; Gholami, A.; Dong, Z.; Li, X.; Shen, S.; Mahoney, M. W.; and Keutzer, K. 2023 b . SqueezeLLM: Dense-and-Sparse Quantization . arXiv
2023
-
[33]
M.; and Moulines, E
Leconte, L.; Bedin, L.; Nguyen, V. M.; and Moulines, E. 2024. ReALLM: A general framework for LLM compression and fine-tuning. arXiv:2405.13155
2024 arXiv
-
[34]
Lee, C.; Jin, J.; Kim, T.; Kim, H.; and Park, E. 2024. OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models . Proceedings of the AAAI Conference on Artificial Intelligence, 38(12): 13355--13364
2024
-
[35]
H.; Kim, J.; Kwon, S
Lee, J. H.; Kim, J.; Kwon, S. J.; and Lee, D. 2023. FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization. In Krause, A.; Brunskill, E.; Cho, K.; Engelhardt, B.; Sabato, S.; and Scarlett, J., eds., International Conference on Machine Learn...
2023
-
[36]
Li, G.; Tang, Y.; and Zhang, W. 2024. LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language Models . arXiv
2024
-
[37]
Li, Y.; Gong, R.; Tan, X.; Yang, Y.; Hu, P.; Zhang, Q.; Yu, F.; Wang, W.; and Gu, S. 2021. Brecq: Pushing the limit of post-training quantization by block reconstruction. arXiv preprint arXiv:2102.05426
2021 arXiv
-
[38]
Li, Y.; Yu, Y.; Liang, C.; Karampatziakis, N.; He, P.; Chen, W.; and Zhao, T. 2024. LoftQ: Lo RA -Fine-Tuning-aware Quantization for Large Language Models. In The Twelfth International Conference on Learning Representations
2024
-
[39]
Liao, B.; and Monz, C. 2024. ApiQ: Finetuning of 2-Bit Quantized Large Language Model . arXiv
2024
-
[40]
Lin, J.; Tang, J.; Tang, H.; Yang, S.; Dang, X.; and Han, S. 2023. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration . arXiv
2023
-
[41]
Liu, J.; Gong, R.; Wei, X.; Dong, Z.; Cai, J.; and Zhuang, B. 2024. QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models. arXiv:2310.08041
2024 arXiv
-
[42]
Liu, Z.; Oguz, B.; Zhao, C.; Chang, E.; Stock, P.; Mehdad, Y.; Shi, Y.; Krishnamoorthi, R.; and Chandra, V. 2023. LLM-QAT: Data-Free Quantization Aware Training for Large Language Models . arXiv
2023
-
[43]
Meta. 2024. Llama 3 Model Card
2024
-
[44]
OpenAI. 2023. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774
2023 arXiv
-
[45]
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2019. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv e-prints
2019
-
[46]
E.; Adi, Y.; Liu, J.; Sauvestre, R.; Remez, T.; Rapin, J.; Kozhevnikov, A.; Evtimov, I.; Bitton, J.; Bhatt, M.; Ferrer, C
Rozière, B.; Gehring, J.; Gloeckle, F.; Sootla, S.; Gat, I.; Tan, X. E.; Adi, Y.; Liu, J.; Sauvestre, R.; Remez, T.; Rapin, J.; Kozhevnikov, A.; Evtimov, I.; Bitton, J.; Bhatt, M.; Ferrer, C. C.; Grattafiori, A.; Xiong, W.; Défossez, A.; Copet, J.; Azhar, F.; Touvron, H.; Mart...
2024 arXiv
-
[47]
L.; Bhagavatula, C.; and Choi, Y
Sakaguchi, K.; Bras, R. L.; Bhagavatula, C.; and Choi, Y. 2019. WinoGrande: An Adversarial Winograd Schema Challenge at Scale. arXiv preprint arXiv:1907.10641
2019 arXiv
-
[48]
Shao, W.; Chen, M.; Zhang, Z.; Xu, P.; Zhao, L.; Li, Z.; Zhang, K.; Gao, P.; Qiao, Y.; and Luo, P. 2024. OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models. In The Twelfth International Conference on Learning Representations
2024
-
[49]
E.; and Stoica, I
Sheng, Y.; Cao, S.; Li, D.; Hooper, C.; Lee, N.; Yang, S.; Chou, C.; Zhu, B.; Zheng, L.; Keutzer, K.; Gonzalez, J. E.; and Stoica, I. 2023. S-LoRA: Serving Thousands of Concurrent LoRA Adapters . arXiv
2023
-
[50]
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; Bikel, D.; Blecher, L.; Ferrer, C. C.; Chen, M.; Cucurull, G.; Esiobu, D.; Fernandes, J.; Fu, J.; Fu, W.; Fuller, B.; Gao, C.; Goswami, V.; Goyal, N....
2023 arXiv
-
[51]
Tseng, A.; Chee, J.; Sun, Q.; Kuleshov, V.; and Sa, C. D. 2024. QuIP\#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks . arXiv
2024
-
[52]
A.; Khashabi, D.; and Hajishirzi, H
Wang, Y.; Kordi, Y.; Mishra, S.; Liu, A.; Smith, N. A.; Khashabi, D.; and Hajishirzi, H. 2023. Self-Instruct: Aligning Language Models with Self-Generated Instructions. In Rogers, A.; Boyd-Graber, J.; and Okazaki, N., eds., Proceedings of the 61st Annual Meeting of the Associa...
2023
-
[53]
W.; Lester, B.; Du, N.; Dai, A
Wei, J.; Bosma, M.; Zhao, V.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2022 a . Finetuned Language Models are Zero-Shot Learners. In International Conference on Learning Representations
2022
-
[54]
Wei, X.; Zhang, Y.; Li, Y.; Zhang, X.; Gong, R.; Guo, J.; and Liu, X. 2023. Outlier Suppression+: Accurate quantization of large language models by equivalent and effective shifting and scaling . Proceedings of the 2023 Conference on Empirical Methods in Natural Language Proce...
2023
-
[55]
Wei, X.; Zhang, Y.; Zhang, X.; Gong, R.; Zhang, S.; Zhang, Q.; Yu, F.; and Liu, X. 2022 b . Outlier Suppression: Pushing the Limit of Low-bit Transformer Language Models . arXiv
2022
-
[56]
Xia, W.; Qin, C.; and Hazan, E. 2024. Chain of LoRA: Efficient Fine-tuning of Language Models via Residual Learning . arXiv
2024
-
[57]
Xu, Y.; Xie, L.; Gu, X.; Chen, X.; Chang, H.; Zhang, H.; Chen, Z.; ZHANG, X.; and Tian, Q. 2024. QA -Lo RA : Quantization-Aware Low-Rank Adaptation of Large Language Models. In The Twelfth International Conference on Learning Representations
2024
-
[58]
Y.; Zhang, M.; Wu, X.; Li, C.; and He, Y
Yao, Z.; Aminabadi, R. Y.; Zhang, M.; Wu, X.; Li, C.; and He, Y. 2022. ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers . arXiv
2022
-
[59]
Yao, Z.; Wu, X.; Li, C.; Youn, S.; and He, Y. 2023. ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation . arXiv
2023
-
[60]
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019. H ella S wag: Can a Machine Really Finish Your Sentence? In Korhonen, A.; Traum, D.; and M \`a rquez, L., eds., Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 4791--4...
2019
-
[61]
A.; and Zhao, Y
Zhang, C.; Cheng, J.; Constantinides, G. A.; and Zhao, Y. 2024 a . LQER: Low-Rank Quantization Error Reconstruction for LLMs. arXiv preprint arXiv:2402.02446
2024 arXiv
-
[62]
Zhang, M.; Chen, H.; Shen, C.; Yang, Z.; Ou, L.; Yu, X.; and Zhuang, B. 2023. LoRAPrune: Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning . arXiv
2023
-
[63]
Zhang, T.; Ladhak, F.; Durmus, E.; Liang, P.; McKeown, K.; and Hashimoto, T. B. 2024 b . Benchmarking Large Language Models for News Summarization . Transactions of the Association for Computational Linguistics, 12: 39--57
2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.