REVIEW 7 cited by
Approaching Deep Learning through the Spectral Dynamics of Weights
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We propose an empirical approach centered on the spectral dynamics of weights -- the behavior of singular values and vectors during optimization -- to unify and clarify several phenomena in deep learning. We identify a consistent bias in optimization across various experiments, from small-scale ``grokking'' to large-scale tasks like image classification with ConvNets, image generation with UNets, speech recognition with LSTMs, and language modeling with Transformers. We also demonstrate that weight decay enhances this bias beyond its role as a norm regularizer, even in practical systems. Moreover, we show that these spectral dynamics distinguish memorizing networks from generalizing ones, offering a novel perspective on this longstanding conundrum. Additionally, we leverage spectral dynamics to explore the emergence of well-performing sparse subnetworks (lottery tickets) and the structure of the loss surface through linear mode connectivity. Our findings suggest that spectral dynamics provide a coherent framework to better understand the behavior of neural networks across diverse settings.
Forward citations
Cited by 7 Pith papers
-
Disentangling Geometry, Performance, and Training in Language Models
Effective rank of the unembedding matrix mainly reflects hyperparameters like batch size and weight decay and is not a reliable predictor of language-model performance.
-
At-Grok Is Not Converged:A Measurement-Validity Audit for Grokking Representation Metrics
Embedding effective rank at grokking is a transient that overstates the converged floor by 3–5× (MLP) / 1.3–1.5× (transformer), and compression lags generalization by order T_grok, modulated by LayerNorm.
-
Diffract: Spectral View of LLM Domain Adaptation
During continual pre-training of OLMo 2 models, singular value spectra remain largely fixed while singular vectors change; selectively rewinding low-importance attention heads improves math accuracy by up to 4%.
-
Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility
Logic pre-pretraining, training a small LM on next-step formal derivations before natural language, accelerates skill acquisition and improves pruning robustness at a 100B-token scale.
-
Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws
For a high-dimensional Gaussian sequence model, empirical risk minimization in single-head tied attention has exactly computable test error, interpolation and recovery thresholds, and a singular-value spectrum that be...
-
On the Local Complexity of Linear Regions in Deep ReLU Networks
The paper introduces a noise-regularized measure of linear-region density and derives inequalities relating it to representation rank, total variation, and representation cost, but key inequalities contain a false nor...
-
S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner
S3LoRA prunes LoRA layers with the sharpest spectral update concentration to improve safety in fine-tuned LLM agents without needing base models or extra data.
Discussion (0). Continue with ORCID to comment.