Pith. sign in

REVIEW 2 cited by

Gradient dynamics for low-rank fine-tuning beyond kernels

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.15385 v1 pith:H5HSXINO submitted 2024-11-23 cs.LG math.STstat.MLstat.TH

classification cs.LGmath.STstat.MLstat.TH
keywords modellow-ranksettingweightsfine-tuninggivengradientteacher
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

LoRA has emerged as one of the de facto methods for fine-tuning foundation models with low computational cost and memory footprint. The idea is to only train a low-rank perturbation to the weights of a pre-trained model, given supervised data for a downstream task. Despite its empirical sucess, from a mathematical perspective it remains poorly understood what learning mechanisms ensure that gradient descent converges to useful low-rank perturbations. In this work we study low-rank fine-tuning in a student-teacher setting. We are given the weights of a two-layer base model $f$, as well as i.i.d. samples $(x,f^*(x))$ where $x$ is Gaussian and $f^*$ is the teacher model given by perturbing the weights of $f$ by a rank-1 matrix. This generalizes the setting of generalized linear model (GLM) regression where the weights of $f$ are zero. When the rank-1 perturbation is comparable in norm to the weight matrix of $f$, the training dynamics are nonlinear. Nevertheless, in this regime we prove under mild assumptions that a student model which is initialized at the base model and trained with online gradient descent will converge to the teacher in $dk^{O(1)}$ iterations, where $k$ is the number of neurons in $f$. Importantly, unlike in the GLM setting, the complexity does not depend on fine-grained properties of the activation's Hermite expansion. We also prove that in our setting, learning the teacher model "from scratch'' can require significantly more iterations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning

    cs.LG 2025-06 conditional novelty 7.0 of 10

    MoE provably learns clustered single-index regression with sample complexity Õ(d^{k*-1}), while a vanilla network provably cannot recover the shared global feature due to gradient cancellation.

  2. LoRA Training Provably Converges to a Low-Rank Global Minimum or It Fails Loudly (But it Probably Won't Fail)

    cs.LG 2025-02 conditional novelty 7.0 of 10

    Under restricted strong convexity and smoothness, every stable point of LoRA training is either a low-rank global minimum or a high-rank, large-magnitude spurious minimum, and practical initialization and weight decay...

Pith tools