REVIEW 14 cited by
The Road Less Scheduled
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Existing learning rate schedules that do not require specification of the optimization stopping step T are greatly out-performed by learning rate schedules that depend on T. We propose an approach that avoids the need for this stopping time by eschewing the use of schedules entirely, while exhibiting state-of-the-art performance compared to schedules across a wide family of problems ranging from convex problems to large-scale deep learning problems. Our Schedule-Free approach introduces no additional hyper-parameters over standard optimizers with momentum. Our method is a direct consequence of a new theory we develop that unifies scheduling and iterate averaging. An open source implementation of our method is available at https://github.com/facebookresearch/schedule_free. Schedule-Free AdamW is the core algorithm behind our winning entry to the MLCommons 2024 AlgoPerf Algorithmic Efficiency Challenge Self-Tuning track.
Forward citations
Cited by 14 Pith papers
-
In-Context Multiple Instance Learning
A model pretrained on synthetic bag-structured data performs in-context learning for new MIL tasks from a handful of examples and outperforms task-specific supervised baselines on twelve benchmarks.
-
From Syntax to Semantics: Unveiling the Emergence of Chirality in SMILES Translation Models
Chirality emerges in SMILES translation models through an abrupt encoder-centered reorganization of representations after a long plateau, identified via checkpoint analysis and ablation.
-
Training Deep Learning Models with Norm-Constrained LMOs
Scion is a new stochastic LMO-based optimizer family that unifies existing methods, supports unconstrained problems, and delivers hyperparameter transferability plus speedups on nanoGPT training.
-
Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions
Under fixed innovation coupling, finite-horizon optimizers admit minimal pathwise realizations and incidence-identifiable Möbius effects, with a five-term readout transfer from hidden relaxation and a closed reduced-v...
-
Modeling Local, Global, and Cross-Modal Context in Multimodal 3D MRI
MICViT outperforms CNN and transformer baselines on brain age prediction from multimodal 3D MRI by combining modality-specific and cross-modal local/global attention across three heterogeneous datasets.
-
Predictable Scaling Laws of Optimal Hyperparameters for LLM Continued Pre-training
Optimal hyperparameters for LLM continued pre-training follow predictable scaling laws derived from proxy models, enabling a two-stage framework that predicts settings from compute budget and checkpoint state to reduc...
-
CoAction: Cross-task Correlation-aware Pareto Set Learning
CoAction introduces a task-aware transformer that simultaneously learns Pareto optimal solutions across multiple tasks by capturing inter-task correlations via self-attention and task embeddings.
-
CoAction: Cross-task Correlation-aware Pareto Set Learning
CoAction applies a transformer encoder with per-task embeddings to jointly solve multiple multi-objective optimization problems by capturing cross-task correlations.
-
Integrating Weather Foundation Model and Satellite to Enable Fine-Grained Solar Irradiance Forecasting
Baguan-solar integrates Baguan weather foundation model forecasts with geostationary satellite data via a decoupled two-stage multimodal framework to deliver kilometer-scale 24-hour solar irradiance predictions, cutti...
-
Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions
A modular calculus decomposes optimizer updates into geometric preconditioning plus structured nongeometric mechanisms, with a direction-expressivity theorem showing full SPD geometry captures exactly strict descent d...
-
Central limit theorem for the averaged Adam optimizer
Establishes a central limit theorem for averaged Adam with n^{-1/2} convergence rate to an attracting zero and covariance determined by the algorithm at the attractor.
-
Scalable Hyperparameter-Divergent Ensemble Training with Automatic Learning Rate Exploration for Large Models
HDET lets data-parallel replicas explore a spread of learning rates independently before averaging parameters, with an auto-LR controller driven by inter-replica loss differences to produce a self-adapting schedule wi...
-
Analysis of Schedule-Free Nonconvex Optimization
A Lyapunov framework yields O(1/log T) and O(log T/T) gradient-norm rates for Schedule-Free on smooth nonconvex objectives, with the faster rate depending on an unproven assumption.
-
Image-Based Malware Type Classification on MalNet-Image Tiny: Effects of Multi-Scale Fusion, Transfer Learning, Data Augmentation, and Schedule-Free Optimization
Pretraining plus Mixup/TrivialAugment and a feature pyramid network lift macro-F1 from 0.65 to 0.69 on 43-class malware image classification while cutting training epochs from 96 to 10.
Discussion (0). Sign in to comment.