REVIEW 15 cited by
A Spectral Condition for Feature Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
The push to train ever larger neural networks has motivated the study of initialization and training at large network width. A key challenge is to scale training so that a network's internal representations evolve nontrivially at all widths, a process known as feature learning. Here, we show that feature learning is achieved by scaling the spectral norm of weight matrices and their updates like $\sqrt{\texttt{fan-out}/\texttt{fan-in}}$, in contrast to widely used but heuristic scalings based on Frobenius norm and entry size. Our spectral scaling analysis also leads to an elementary derivation of \emph{maximal update parametrization}. All in all, we aim to provide the reader with a solid conceptual understanding of feature learning in neural networks.
Forward citations
Cited by 15 Pith papers
-
Algorithmic Separation between Constant-Depth and Logarithmic-Depth Neural Networks
Logarithmic-depth networks trained by layerwise coordinate descent can learn hierarchical staircase Boolean functions that constant-depth networks with bounded spectral norms cannot approximate.
-
Conditional Optimal Bridge for Riemannian Activation Steering
Casting activation steering as a Schrödinger Bridge on the residual hypersphere derives the log-density-ratio objective and yields query-adaptive directions that beat fixed baselines without OOD collapse.
-
Training Transformers with Enforced Lipschitz Constants
Transformers can be trained with enforced spectral-norm constraints throughout training, but competitive accuracy requires an astronomical Lipschitz upper bound.
-
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism
Constraining transformer projection weights to a shared low-rank subspace reportedly enables near-lossless compression of pipeline-parallel communication, matching centralized convergence at 80Mbps bandwidth.
-
Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention
Spectral outliers in attention weight matrices, identified by the Marchenko-Pastur threshold, contain a dominant share of functionally important learned structure across 11 transformers.
-
LipSSD: Lipschitz-Constrained Single-Shot Detection for Adversarially Robust Object Detection
Lipschitz-constrained SSD variants improve white-box adversarial robustness in an attack-agnostic way and remain complementary to adversarial training on VOC, KITTI, and LARD.
-
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
Structured-matrix scoring functions, BTT and MLR, let attention escape the low-rank bottleneck and add a distance-dependent compute bias, improving accuracy for fixed compute on regression, language modeling, and forecasting.
-
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
NGN-M, a momentum variant of the NGN step-size, provably converges at O(1/sqrt(K)) under milder assumptions and shows wider step-size stability than Adam, Momo, and SGDM in vision and language tasks.
-
Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance
Falcon-H1 reports competitive benchmark scores for a 0.5B to 34B family of parallel hybrid attention/Mamba-2 models, claiming 2x to 4x parameter efficiency versus dense transformers.
-
Low-rank Momentum Factorization for Memory Efficient Training
MoFaSGD keeps a low-rank factored momentum and uses its singular vectors as the update direction, achieving LoRA-level memory with competitive fine-tuning performance, but its convergence proof is flawed.
-
Practical Efficiency of Muon for Pretraining
Muon trains large language models to the same loss as AdamW with 10 to 15 percent fewer tokens and keeps its advantage at large batch sizes, while muP hyperparameter transfer works with Muon.
-
$\mu$nit Scaling: Simple and Scalable FP8 LLM Training
µnit Scaling combines unit-variance initialization, rearranged LayerNorm, a square-root softmax analysis, and µ-Parametrization-style learning-rate rules to train 1B-13B LLMs in FP8 with no dynamic scaling and zero-sh...
-
Approximate Message Passing for Bayesian Neural Networks
A factor-graph message-passing method for Bayesian neural networks that handles CNNs, avoids double-counting, and shows competitive accuracy with improved calibration on CIFAR-10.
-
Transformed Low-rank Adaptation via Tensor Decomposition and Its Applications to Text-to-image Models
TLoRA combines a tensor-ring-matrix transform with a tensor-ring residual to fine-tune text-to-image models, achieving better or comparable performance than LoRA with far fewer parameters.
-
Scale Weight Decay and Train Better
Muon with weight decay scaled by η/η_max reaches the same MoE validation loss ~30% faster than constant-decay Muon while preserving asymptotic stationarity of the unregularized objective.
Discussion (0). Continue with ORCID to comment.