Pith. sign in

REVIEW 37 cited by

Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.03466 v2 pith:HIF4C7TW submitted 2022-03-07 cs.LG cond-mat.dis-nncs.NE

classification cs.LGcond-mat.dis-nncs.NE
keywords modeltuningparameterscostpretrainingbert-largehyperparametermutransfer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hyperparameter (HP) tuning in deep learning is an expensive process, prohibitively so for neural networks (NNs) with billions of parameters. We show that, in the recently discovered Maximal Update Parametrization (muP), many optimal HPs remain stable even as model size changes. This leads to a new HP tuning paradigm we call muTransfer: parametrize the target model in muP, tune the HP indirectly on a smaller model, and zero-shot transfer them to the full-sized model, i.e., without directly tuning the latter at all. We verify muTransfer on Transformer and ResNet. For example, 1) by transferring pretraining HPs from a model of 13M parameters, we outperform published numbers of BERT-large (350M parameters), with a total tuning cost equivalent to pretraining BERT-large once; 2) by transferring from 40M parameters, we outperform published numbers of the 6.7B GPT-3 model, with tuning cost only 7% of total pretraining cost. A Pytorch implementation of our technique can be found at github.com/microsoft/mup and installable via `pip install mup`.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 37 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error Feedback

    math.OC 2026-07 conditional novelty 7.0 of 10

    Sign-compressing Muon's update to one bit can make it ascend on linear objectives; error feedback works only on the gradient side, yet the divergent sign-after-the-LMO method wins in experiments.

  2. Bridging Compute- and Data-Optimal Pretraining

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Pretraining loss obeys a single law in which repeated or paraphrased tokens count as η(N, data-per-parameter, expansion-ratio) fresh tokens, with total effective data saturating as derived tokens grow.

  3. Geometric Dyson Brownian Motions and the Free Log-Normal Limit for a Non-Square Gaussian Matrix Product

    math.PR 2026-06 unverdicted novelty 7.0 of 10

    In double asymptotic limits, the squared singular value process of non-square matrix products obeys geometric Dyson Brownian motion whose T-transform solves a Burgers equation, producing the free log-normal law via fr...

  4. Deep Delta Learning

    cs.LG 2026-01 unverdicted novelty 7.0 of 10

    Replacing additive residual connections with a gated rank-1 delta update that interpolates identity, projection, and reflection slightly improves language modeling and downstream averages in reported 124M/353M runs.

  5. FM4NPP: A Scaling Foundation Model for Nuclear and Particle Physics

    cs.LG 2025-08 conditional novelty 7.0 of 10

    A 188M-parameter Mamba model pretrained on 11M+ simulated sPHENIX events with a new serialization and neighbor-prediction task beats task-specific baselines on three downstream detector tasks when frozen and paired wi...

  6. Geometry of Neural Reinforcement Learning in Continuous State and Action Spaces

    cs.LG 2025-07 conditional novelty 7.0 of 10

    For wide two-layer linearized neural policies in deterministic continuous RL, the locally attainable states concentrate on a manifold of dimension at most 2da+1, independent of the state dimension.

  7. Adaptive kernel predictors from feature-learning infinite limits of neural networks

    cs.LG 2025-02 conditional novelty 7.0 of 10

    Feature-learning infinite-width neural networks are kernel machines with data-dependent kernels, defined by a min-max saddle point (Bayesian/Langevin) or a DMFT fixed point (gradient flow with weight decay).

  8. Neural Scaling Laws Rooted in the Data Distribution

    cs.LG 2024-12 conditional novelty 7.0 of 10

    Percolation theory at criticality produces a Zipf distribution of subtasks, from which the paper derives neural scaling laws with alpha=1 quanta and a data-scaling exponent of 0.5.

  9. Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A hybrid NoPE visual diffusion backbone with HeteroP scaling yields ~7× pretraining compute efficiency versus matched full attention and near-stable zero-shot 5s→30s video extrapolation.

  10. Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Splitting weight matrices into a fixed-norm direction and learnable per-row/column magnitudes improves LLM training over AdamW/Muon, removes weight decay and warmup, and transfers the optimal LR across width.

  11. Spectral Reach: Understanding Neural Scaling as Progress into the Spectral Tail

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Neural scaling occurs because larger models maintain learning on weaker eigenmodes of the eNTK that smaller models cannot access.

  12. SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning

    cs.LG 2026-02 conditional novelty 6.0 of 10

    A width-progressive training method (RMS-preserving rescaling plus asymmetric optimizer-state reset and LR rewarmup) enables mid-training 2x width expansion with up to 35% compute savings over training from scratch.

  13. Scaling depth capacity via zero/one-layer model expansion

    cs.LG 2025-11 conditional novelty 6.0 of 10

    Training GPT2 from a zero/one-layer model and expanding depth at 80% of the schedule reaches fixed-size loss with approximately 5x less compute.

  14. Customizing the Inductive Biases of Softmax Attention using Structured Matrices

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Structured-matrix scoring functions, BTT and MLR, let attention escape the low-rank bottleneck and add a distance-dependent compute bias, improving accuracy for fixed compute on regression, language modeling, and forecasting.

  15. Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Falcon-H1 reports competitive benchmark scores for a 0.5B to 34B family of parallel hybrid attention/Mamba-2 models, claiming 2x to 4x parameter efficiency versus dense transformers.

  16. Language Models Improve When Pretraining Data Matches Target Tasks

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Ranking pretraining documents by similarity to benchmark training examples (BETR) yields consistent benchmark gains and a 2.1x compute multiplier over DCLM-Baseline.

  17. BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A ReLU-routed MoE with chunk-level sparsity training objectives and custom kernels combining activation sparsity with speculative decoding achieves over 70% 8-token chunk sparsity and up to 3.67x end-side speedup.

  18. Simple Convergence Proof of Adam From a Sign-like Descent Perspective

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Adam's O(1/T^1/4) convergence is proven from a sign-like descent perspective, but the dimension-free claim depends on restrictive coordinate-wise assumptions.

  19. Decoupled Relative Learning Rate Schedules

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Per-component relative learning-rate schedules speed up Transformer pretraining by up to 23% for MoE models, and the schedules tuned on a 34M model transfer to models 27x larger.

  20. Eigenspectrum Analysis of Neural Networks without Aspect Ratio Bias

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Aspect ratio of weight matrices biases heavy-tail spectral metrics; the new FARMS subsampling method removes this bias and improves downstream layer-wise tuning.

  21. Dense Local Dependencies Induce Attention-Logit Explosion and Training Instability During Long-Sequence Transformer Training

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Long-sequence transformers blow up because softmax attention cannot fit dense local dependency patterns with its low-rank parameterization, and separating local and global attention heads prevents the resulting logit ...

  22. Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression

    cs.RO 2025-02 conditional novelty 6.0 of 10

    HMA is a masked autoregressive transformer that predicts future video and actions across many robot embodiments, running up to 15x faster than prior diffusion-based video simulators while matching or improving visual ...

  23. Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Peri-LN, which normalizes both the input and output of each sublayer, reduces activation-variance growth and gradient spikes during LLM pretraining, outperforming Pre-LN and Post-LN at scales up to 3.2B parameters.

  24. Towards Precise Scaling Laws for Video Diffusion Transformers

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Video diffusion transformers follow scaling laws, and tuning learning rate and batch size per model and data size makes those laws precise enough to predict loss and optimal model size at larger scales.

  25. Scale Weight Decay and Train Better

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Muon with weight decay scaled by η/η_max reaches the same MoE validation loss ~30% faster than constant-decay Muon while preserving asymptotic stationarity of the unregularized objective.

  26. Index SLM Technical Report

    cs.CL 2026-07 accept novelty 5.0 of 10

    Index-1.9B-Base reaches 64.92 average benchmark score via WSD training with late curated data plus Norm-Head, with open Pure/Boost controls isolating instruction-data inflation.

  27. The Role of Rigor in Artificial Intelligence

    cs.AI 2026-05 conditional novelty 5.0 of 10

    Modern AI's distinctive trajectory is explained by the primacy of operational rigor over conceptual and epistemic rigor across successive paradigms.

  28. Tree-Structured Parzen Estimator Can Solve Black-Box Combinatorial Optimization More Efficiently

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A distance-based categorical kernel and two algorithmic modifications let TPE optimize combinatorial spaces more efficiently than the original TPE on synthetic benchmarks.

  29. SingLoRA: Low Rank Adaptation Using a Single Matrix

    cs.AI 2025-07 conditional novelty 5.0 of 10

    SingLoRA replaces LoRA's two matrices A and B with one matrix A and the symmetric update AA^T, cutting adapter parameters roughly in half while claiming more stable fine-tuning.

  30. MiniCPM4: Ultra-Efficient LLMs on End Devices

    cs.CL 2025-06 conditional novelty 5.0 of 10

    MiniCPM4-8B reportedly matches Qwen3-8B on standard benchmarks while using about 22% of the training tokens, and achieves large long-context speedups on edge devices.

  31. Xmodel-2 Technical Report

    cs.AI 2024-12 reject novelty 5.0 of 10

    A 1.2B language model trained with tensor-program hyperparameter transfer and WSD decay with SFT data mixing; the headline SOTA claim is not supported by the paper's own benchmark tables.

  32. Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs

    cs.LG 2024-12 conditional novelty 5.0 of 10

    For 3B-7B LLMs, larger batch sizes with lower learning rates improve instruction-tuning benchmarks, early gradient and loss signals predict final quality, and stacked training matches phased training with fewer samples.

  33. Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs

    cs.LG 2025-07 conditional novelty 4.0 of 10

    High data redundancy and over-training decelerate LLM performance gains, and the authors fit a sub-optimal scaling law with logistic correction terms to predict the slowdown.

  34. NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models

    cs.AI 2025-06 conditional novelty 4.0 of 10

    The authors propose a competition with new scoring metrics to find benchmarks that give clean early-training signals for small language models, and show MMLU-var outperforms MMLU as a baseline.

  35. YuLan-Mini: An Open Data-efficient Language Model

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A 2.42B-parameter base model trained on 1.08T tokens matches or beats several industry baselines trained on 7T to 18T tokens across math, code, and general benchmarks.

  36. Scaling Transformers for Time Series Forecasting: Do Pretrained Large Models Outperform Small-Scale Alternatives?

    cs.LG 2025-06 reject novelty 3.0 of 10

    LLM4TS_FS achieves the best MSE on four of seven long-term datasets, but the claimed broad advantage of pre-trained large models over small transformers is not consistent across all benchmarks.

  37. Quantum Machine Learning: A Hands-on Tutorial for Machine Learning Practitioners and Researchers

    quant-ph 2025-02 unverdicted novelty 2.0 of 10

    A structured tutorial that introduces quantum machine learning concepts, algorithms, theory, and PennyLane code to classical ML practitioners.

Pith tools