Pith. sign in

REVIEW 14 cited by

The Road Less Scheduled

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.15682 v4 pith:FZZECVVP submitted 2024-05-24 cs.LG cs.AImath.OCstat.ML

classification cs.LGcs.AImath.OCstat.ML
keywords scheduleslearningproblemsapproachmethodrateschedule-freestopping
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing learning rate schedules that do not require specification of the optimization stopping step T are greatly out-performed by learning rate schedules that depend on T. We propose an approach that avoids the need for this stopping time by eschewing the use of schedules entirely, while exhibiting state-of-the-art performance compared to schedules across a wide family of problems ranging from convex problems to large-scale deep learning problems. Our Schedule-Free approach introduces no additional hyper-parameters over standard optimizers with momentum. Our method is a direct consequence of a new theory we develop that unifies scheduling and iterate averaging. An open source implementation of our method is available at https://github.com/facebookresearch/schedule_free. Schedule-Free AdamW is the core algorithm behind our winning entry to the MLCommons 2024 AlgoPerf Algorithmic Efficiency Challenge Self-Tuning track.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. In-Context Multiple Instance Learning

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    A model pretrained on synthetic bag-structured data performs in-context learning for new MIL tasks from a handful of examples and outperforms task-specific supervised baselines on twelve benchmarks.

  2. From Syntax to Semantics: Unveiling the Emergence of Chirality in SMILES Translation Models

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Chirality emerges in SMILES translation models through an abrupt encoder-centered reorganization of representations after a long plateau, identified via checkpoint analysis and ablation.

  3. Training Deep Learning Models with Norm-Constrained LMOs

    cs.LG 2025-02 unverdicted novelty 7.0 of 10

    Scion is a new stochastic LMO-based optimizer family that unifies existing methods, supports unconstrained problems, and delivers hyperparameter transferability plus speedups on nanoGPT training.

  4. Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Under fixed innovation coupling, finite-horizon optimizers admit minimal pathwise realizations and incidence-identifiable Möbius effects, with a five-term readout transfer from hidden relaxation and a closed reduced-v...

  5. Modeling Local, Global, and Cross-Modal Context in Multimodal 3D MRI

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    MICViT outperforms CNN and transformer baselines on brain age prediction from multimodal 3D MRI by combining modality-specific and cross-modal local/global attention across three heterogeneous datasets.

  6. Predictable Scaling Laws of Optimal Hyperparameters for LLM Continued Pre-training

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    Optimal hyperparameters for LLM continued pre-training follow predictable scaling laws derived from proxy models, enabling a two-stage framework that predicts settings from compute budget and checkpoint state to reduc...

  7. CoAction: Cross-task Correlation-aware Pareto Set Learning

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    CoAction introduces a task-aware transformer that simultaneously learns Pareto optimal solutions across multiple tasks by capturing inter-task correlations via self-attention and task embeddings.

  8. CoAction: Cross-task Correlation-aware Pareto Set Learning

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    CoAction applies a transformer encoder with per-task embeddings to jointly solve multiple multi-objective optimization problems by capturing cross-task correlations.

  9. Integrating Weather Foundation Model and Satellite to Enable Fine-Grained Solar Irradiance Forecasting

    cs.LG 2026-03 unverdicted novelty 6.0 of 10

    Baguan-solar integrates Baguan weather foundation model forecasts with geostationary satellite data via a decoupled two-stage multimodal framework to deliver kilometer-scale 24-hour solar irradiance predictions, cutti...

  10. Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A modular calculus decomposes optimizer updates into geometric preconditioning plus structured nongeometric mechanisms, with a direction-expressivity theorem showing full SPD geometry captures exactly strict descent d...

  11. Central limit theorem for the averaged Adam optimizer

    math.PR 2026-06 unverdicted novelty 5.0 of 10

    Establishes a central limit theorem for averaged Adam with n^{-1/2} convergence rate to an attracting zero and covariance determined by the algorithm at the attractor.

  12. Scalable Hyperparameter-Divergent Ensemble Training with Automatic Learning Rate Exploration for Large Models

    cs.LG 2026-04 unverdicted novelty 5.0 of 10

    HDET lets data-parallel replicas explore a spread of learning rates independently before averaging parameters, with an auto-LR controller driven by inter-replica loss differences to produce a self-adapting schedule wi...

  13. Analysis of Schedule-Free Nonconvex Optimization

    cs.LG 2025-08 conditional novelty 4.0 of 10

    A Lyapunov framework yields O(1/log T) and O(log T/T) gradient-norm rates for Schedule-Free on smooth nonconvex objectives, with the faster rate depending on an unproven assumption.

  14. Image-Based Malware Type Classification on MalNet-Image Tiny: Effects of Multi-Scale Fusion, Transfer Learning, Data Augmentation, and Schedule-Free Optimization

    cs.CR 2026-04 unverdicted novelty 2.0 of 10

    Pretraining plus Mixup/TrivialAugment and a feature pyramid network lift macro-F1 from 0.65 to 0.69 on 43-class malware image classification while cutting training epochs from 96 to 10.

Pith tools