REVIEW 33 cited by
LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Quantization is an indispensable technique for serving Large Language Models (LLMs) and has recently found its way into LoRA fine-tuning. In this work we focus on the scenario where quantization and LoRA fine-tuning are applied together on a pre-trained model. In such cases it is common to observe a consistent gap in the performance on downstream tasks between full fine-tuning and quantization plus LoRA fine-tuning approach. In response, we propose LoftQ (LoRA-Fine-Tuning-aware Quantization), a novel quantization framework that simultaneously quantizes an LLM and finds a proper low-rank initialization for LoRA fine-tuning. Such an initialization alleviates the discrepancy between the quantized and full-precision model and significantly improves generalization in downstream tasks. We evaluate our method on natural language understanding, question answering, summarization, and natural language generation tasks. Experiments show that our method is highly effective and outperforms existing quantization methods, especially in the challenging 2-bit and 2/4-bit mixed precision regimes. The code is available on https://github.com/yxli2123/LoftQ.
Forward citations
Cited by 33 Pith papers
-
SoftWater: Class-Aware Rate Allocation for Softmax Quantization
SoftWater, a KL-divergence-based quantizer for LLM softmax heads, allocates bit rate by class frequency and variance and beats WaterSIC at matched head rates on 59 of 60 test points.
-
Small Foundation Models of Human Cognition and Behaviour
Tiny cognitively fine-tuned models match a 70B model on familiar experiments, and prompt ablations show they use stimulus and feedback content, not choice-history shortcuts alone.
-
ST-LoRA: Single Trajectory LoRA Ensemble for Uncertainty Aware Agricultural Segmentation
Combining LoRA with snapshot ensembling yields a parameter-efficient uncertainty-aware segmentation ensemble that matches snapshot full-rank baselines, with feed-forward layers identified as the critical LoRA target.
-
Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards
Non-vacuous PAC-Bayes generalization bounds for billion-parameter RLVR models, obtained by a Gumbel-max reparameterization and aggressive TinyLoRA distillation/quantization, are claimed for four tasks.
-
Dive Into the Implicit Biases of Low-rank Vision-language Alignment
Low-rank LLM adaptation during vision-language alignment outperforms full fine-tuning by preserving per-token visual structure and favoring flat, noise-robust subspaces.
-
AutoNeural: Co-Designing Vision-Language Models for NPU Inference
A NPU-native VLM combining a MobileNet-style encoder with a hybrid Transformer-SSM backbone claims 14x lower latency and 7x lower quantization error over ViT-Transformer baselines, though quantized accuracy is not reported.
-
ConsNoTrainLoRA: Data-driven Weight Initialization of Low-rank Adapters using Constraints
LoRA adapters can be initialized with a closed-form estimate derived from constraint sets linking source and target activations, improving fine-tuning speed and accuracy.
-
DeCAF: Decentralized Consensus-And-Factorization for Low-Rank Adaptation of Foundation Models
A truncated-SVD consensus step for decentralized LoRA is claimed to reach O(1/sqrt T) convergence, matching decentralized SGD, with supporting CLIP and LLAMA2-7B experiments.
-
Saliency-Aware Quantized Imitation Learning for Efficient Robotic Control
A saliency-weighted quantization-aware training method lets 4-bit quantized imitation-learning policies match full-precision success rates across robot manipulation, driving, and control benchmarks.
-
Communication-Efficient Hybrid Language Model via Uncertainty-Aware Opportunistic and Compressed Transmission
An uncertainty-aware hybrid language model skips and compresses uplink token transmissions, achieving up to 206 times higher token throughput with 97.4% accuracy in simulation.
-
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
LowRA enables LoRA fine-tuning with base weights at 1.15 to 4 bits per parameter, outperforming QLoRA and LoftQ at equal bit widths and matching their accuracy at lower bit widths.
-
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters
Sharing one MLP layer's weights across several layers plus low-rank adapters recovers most of a pretrained LLM's quality with a fraction of the storage and faster phone inference.
-
CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization
CLoQ initializes LoRA adapters on quantized LLMs with a closed-form calibration-aware low-rank solution, improving 2-bit fine-tuning accuracy.
-
Matryoshka Re-Ranker: A Flexible Re-Ranking Architecture With Configurable Depth and Width
A single LLM re-ranker can be configured at runtime to different depths and widths, with training tricks that keep compressed variants close to full-scale accuracy.
-
FBQuant: FeedBack Quantization for Large Language Models
FBQuant redefines sub-branch compensation as Q(W - Sigma) + Sigma, bounding per-weight reconstruction error by half the quantizer step and improving 3-bit LLM accuracy.
-
FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices
FlexQuant generates a family of shared-parameter quantized LLMs by gradually replacing modules with lower-bit versions, cutting storage and improving memory granularity.
-
AdvAnchor: Enhancing Diffusion Model Unlearning with Adversarial Anchors
AdvAnchor generates adversarial anchors, embeddings perturbed to be dissimilar from the target concept, and fine-tunes the model toward them, improving the erasure-preservation trade-off in diffusion model unlearning.
-
Uncertainty-Aware Hybrid Inference with On-Device Small and Remote Large Language Models
A small on-device language model can skip sending most tokens to a remote large language model when its temperature-perturbation uncertainty is low, cutting uplink load by 45.93% and speeding token throughput 2.54x.
-
Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You Need
A modified LLM that predicts the next pixel value in language space achieves state-of-the-art lossless image compression rates on several benchmarks.
-
Efficient Reasoning on the Edge
LoRA adapters, budget-forced GRPO, dynamic switching, parallel verification and FPTQuant enable practical chain-of-thought reasoning on quantized Qwen2.5-7B for edge devices.
-
Examining the Efficacy of Graph Neural Network Message-Passing in Regression Contexts
Across four NAS/DNN-predictor regression benchmarks, GEN (deep graph convolution) achieves the best average rank over 11 GNN message-passing layers, though attention GATv2 wins on the largest graphs.
-
Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs
The study introduces TruthfulnessEval and reports that 4-bit quantization preserves simple true/false accuracy, but explicit 'lie' prompts make quantized and full-precision LLMs output falsehoods even when internal pr...
-
Quaff: Quantized Parameter-Efficient Fine-Tuning under Outlier Spatial Stability Hypothesis
Quaff shows that activation outlier channels keep their spatial positions during LLM fine-tuning, and exploits this stability to cut fine-tuning memory and latency with INT8 quantization while matching or beating full...
-
Instance-dependent Early Stopping
IES removes already-mastered training examples from backpropagation using a threshold on the second-order difference of their loss, achieving comparable accuracy with 10-50% less backpropagation.
-
HASSLE-free: A unified Framework for Sparse plus Low-Rank Matrix Decomposition for LLMs
HASSLE-free gives a fuller-Hessian alternating-minimization recipe for sparse-plus-low-rank LLM compression and reports perplexity improvements over OATS on Llama-3 and Llama-3.2 models.
-
Large Language Model Enabled Multi-Task Physical Layer Network
A single fine-tuned LLM backbone with task-specific encoders, decoders, and text prompts performs three physical-layer wireless tasks with accuracy close to dedicated single-task networks.
-
Quantization-Aware Imitation-Learning for Resource-Efficient Robotic Control
Quantization-aware imitation learning with teacher-distillation (QAIL+QBC) recovers full-precision policy accuracy under 4-bit weight and activation quantization across robot manipulation, driving, and control benchmarks.
-
From Financial Sentiment Classification to Return Predictability: A QLoRA Benchmark of Large Language Models
QLoRA tuning improves financial sentiment classification, but no sentiment model shows robust next-session stock return predictability in the 2019 Benzinga sample.
-
Mixture of Low Rank Adaptation with Partial Parameter Sharing for Time Series Forecasting
MoLA adapts a pre-trained short-horizon forecaster to multiple forecast steps via segment-specific mixtures of shared low-rank adapters, reporting modest mean-squared-error gains over the base models on most of eight ...
-
A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms
A survey that organizes foundation-model recommender systems into feature-based, generative, and agentic paradigms and reviews tasks, empirical results, and open challenges.
-
Deploying Foundation Model Powered Agent Services: A Survey
This survey proposes a layered framework (execution, resource, model, agent, application) for deploying foundation-model-powered agent services across edge-cloud environments, and reviews optimization techniques at ea...
-
Progtuning: Progressive Fine-tuning Framework for Transformer-based Language Models
A progressive scheduling trick that updates only the last remaining blocks in later epochs reduces parameter-update counts by about 25% with roughly unchanged GLUE and SQuAD scores.
-
Optimizing Large Language Models with an Enhanced LoRA Fine-Tuning Algorithm for Efficiency and Robustness in NLP Tasks
A modified LoRA update with per-matrix learning rates and an object-detection-style density term is reported to slightly improve QQP accuracy over GPT-4 baselines.
Discussion (0). Continue with ORCID to comment.