Pith. sign in

REVIEW 3 major objections 5 minor 142 references

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers

T0 review · 3 major / 5 minor · reviewed 2026-07-11 · grok-4.5

Pith's one-line read OmniOpt maps more than a hundred modern optimizers onto a shared update pipeline and geometry, then shows that no single method dominates the quality–runtime–memory frontier.

desk verdict A usable survey-plus-benchmark package that couples a five-stage pipeline, LMO axes, dual taxonomy, and multi-objective LLM/vision results—worth engaging, with protocol caveats on primary labels and Stage-1 isolation. read the letter →

arxiv 2607.04033 v1 pith:D75ZLFEO submitted 2026-07-04 cs.LG cs.AI

classification cs.LGcs.AI
keywords optimizersLLMpretrainingmeta-pipelinelinearminimizationoracleoptimizertaxonomymulti-objectivebenchmarkAdamWMuon
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Optimizer choice for large-model training is no longer a formula preference; it is a joint decision about compute, memory, tuning budget, and task. OmniOpt argues that the fragmented landscape of over a hundred methods can be made operational by treating every update as a five-stage meta-pipeline—where most methods only change one or two stages—and by reading update directions as norm-constrained linear minimization oracles. Those two views ground a dual taxonomy: one axis groups methods by primary mechanism (adaptive moments, matrix structure, sign-like directions, state compression, geometric wrappers), and the other records which measurable training objectives each method targets. A unified cross-domain benchmark then compares representative methods across scales, architectures, contexts, and image classification, and reports systematic family trade-offs rather than a single winner. The paper’s practical claim is that this coordinate system lets practitioners select optimizers under explicit mechanism and objective assumptions instead of chasing unstable global rankings.

What carries the argument

The universal five-stage meta-pipeline (signal acquisition, scoping/routing, gradient transform, state evolution, reconstruction, finalization) plus an LMO-driven four-axis decomposition (update domain, state estimator, geometry/precondition, finalization wrapper) that jointly define the dual taxonomy and the benchmark axes.

What would settle it

Run the full optimizer set with weight decay and clipping always on, at long context and matched wall-clock budgets, and check whether family-level quality–cost–memory orderings reverse relative to the paper’s Stage-1/Stage-2 conclusions, or whether reassigning primary mechanism labels collapses the claimed family trade-offs.

Watch

Extended reading notes

Core claim

Most modern optimizers are sparse modifications of one shared update process, and their directions can be unified as norm-constrained linear minimization oracles along four axes; once methods are grouped that way and scored on multiple effect objectives, no single optimizer dominates, and family rankings cross with scale, context length, and domain.

Load-bearing premise

That each optimizer has one stable primary mechanism family and that the two-stage benchmark—first screening without weight decay or clipping, then transferring only stronger short-context methods—fairly isolates those mechanisms without warping real rankings.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. OmniOpt proposes a unified survey-and-benchmark framework for modern deep-learning optimizers, especially LLM training. It introduces a five-stage universal meta-pipeline (S0–S5) with an identity-mapping principle, an LMO-based four-axis decomposition of updates (domain, state estimator, geometry/precondition, finalization), and a dual taxonomy over 108 methods: mechanism families T1–T5 and effect objectives O1–O6. The core empirical contribution is a controlled multi-scale, multi-architecture pretraining benchmark (C4 short-context screening; FineWeb-Edu 32k transfer; vision) of 24 representative optimizers, plus a Muon mechanistic ablation, arguing that no single optimizer dominates the quality–runtime–memory frontier and that family rankings cross with scale, context, and domain.

Significance. If the framework holds as an operational coordinate system, it would be a high-value contribution: the field has many fragmented optimizer papers and protocol-sensitive benchmarks, but few mechanism-aligned maps that jointly organize theory, taxonomy, and multi-objective evaluation. Strengths include explicit alignment of pipeline stages with LMO axes, broad coverage of T1–T5, controlled-variable benchmarking across scales and architectures, multi-objective reporting (PPL, runtime, memory, stability, LR robustness, transfer), and a useful Muon ablation. The work is more synthesis-plus-benchmark than a new optimizer theorem, but that is appropriate for a survey/cookbook aimed at selection under explicit constraints.

major comments (3)
  1. §6.1 Stage-1 protocol disables weight decay and gradient clipping for all 24 optimizers, then Stage 2 transfers only stronger Stage-1 methods under a production-style recipe. This is load-bearing for family-level claims in §6.2.8–6.2.9 and the abstract’s “no single optimizer dominates / ranking crossings” message. Because many T2/T3/T4 methods interact strongly with S5 finalization (LR–WD–warmup coupling for Lion/Muon; memory-feasibility claims for T4), Stage-1 PPL orderings (e.g., APOLLO/Muon/MARS-Shampoo at 1B) may not isolate pure S2/S3 mechanisms. Please either (i) re-run a Stage-1 subset with matched WD/clip, or (ii) substantially qualify family rankings as protocol-conditional and report which Stage-1 losers would re-enter under regularized screening.
  2. §4.1–4.2 primary-mechanism rule (unique T1–T5 label by “incremental contribution” / dominant non-identity stage) is used to justify family-level O1–O6 summaries. Composite methods (Q-GaLore, MARS-*, Cautious wrappers, COSMOS, etc.) sit on multiple stages by the paper’s own composition notes (§3.1.2). The manuscript needs a clearer, falsifiable assignment protocol—e.g., a short appendix table of contested labels with secondary tags and a sensitivity check showing that reassigning a few boundary methods does not flip the family conclusions in §6.2.8.
  3. §6.2.1–6.2.2 quality claims rest heavily on final PPL under fixed step budgets and per-optimizer LR/knob tuning, while O2/O3 are isolated optimizer runtime/memory. For matrix methods with large per-step overhead (SOAP, Shampoo, Muon), token-normalized PPL alone can overstate practical advantage. The paper already discusses wall-clock trade-offs, but the main family summary should report at least one matched wall-clock or FLOP-normalized comparison at 350M/1B so that “competitive quality” is not conflated with “better under fixed steps.”
minor comments (5)
  1. Figure 1 / abstract claim “over one hundred methods” vs. explicit “108” in §4.2: keep a single count and state inclusion criteria (preprints, variants, wrappers).
  2. Table 5 is dense; a short legend for tags (+res, VR, matrix routing, factored/INT8) earlier in §3.2.3 would help non-specialists.
  3. §3.1 vs §3.2 occasionally switch between S1–S5 and P1–P4 labeling when relating pipeline stages to axes; unify terminology.
  4. Stage-2 Commonsense results are deferred to Appendix B; a compact main-text CS Avg. column or rank-stability summary would better support the O6 transfer claim.
  5. Typos/style: “wild range” in Fig. 1 caption; occasional missing spaces before citations; ensure arXiv IDs/venues for very recent methods are consistent.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: taxonomy is definitional organization; benchmark metrics are independent external measurements.

full rationale

OmniOpt is a survey-plus-benchmark paper. Its four coupled components (five-stage meta-pipeline, LMO four-axis view, dual taxonomy of 108 methods, multi-objective benchmark) organize existing optimizers and measure them on external quantities (PPL, step time, optimizer-state memory, GNormCV, LR perturbation, cross-scenario transfer). The identity-mapping principle and single primary T1–T5 labels are classification design choices, not a derivation that forces empirical rankings: Table 7 is explicitly labeled a “mechanism-informed prior, not a final empirical conclusion,” and family effect tables are “design priors for benchmark planning, not empirical conclusions.” LMO unification is attributed to external geometric work (Bernstein et al., Pethick et al., Sfyraki et al.), not a self-cited uniqueness theorem. Stage-1/2 protocol choices may affect ranking stability (a methodology concern), but they do not make reported trade-offs true by construction. No fitted parameter is renamed as a prediction; no central claim reduces to its inputs by definition. Score 0 is the honest finding.

Assumptions & free parameters 4 free parameters · 5 assumptions · 3 invented entities

The paper’s load-bearing contribution is organizational and empirical, not a closed-form theorem. It rests on standard optimization background plus three paper-specific structuring choices (pipeline stages, primary-mechanism uniqueness, effect objectives) and many benchmark protocol knobs. Invented entities are taxonomical coordinates, not physical objects; free parameters are mostly experimental hyperparameters that affect ranking stability.

free parameters (4)
  • Per-optimizer learning rates and method knobs (betas, eps, APOLLO rank/interval, etc.)
    Stage-1/2 allow optimizer-specific tuning of these while freezing architecture/data/schedule; rankings depend on this search budget.
  • Stage-1 disable of weight decay and gradient clipping
    Design choice to isolate S2/S3 mechanisms; changes absolute and relative performance versus production recipes.
  • Training budgets (steps/tokens) and model scales (60M–1B, 32k context)
    Protocol choices that the paper itself notes can reverse rankings; conclusions weight 350M/1B most heavily.
  • Primary-mechanism assignment rule for composite optimizers
    Hand-assigned unique T1–T5 labels for 108 methods; secondary tags exist but family-level analysis uses the primary label.
assumptions (5)
  • ad hoc to paper Most optimizers are sparse modifications of a shared five-stage update pipeline (identity-mapping principle).
    Stated as central premise in §1.2 and §3.1; used to justify non-overlapping family sites.
  • domain assumption Norm-constrained LMOs (and dual preconditioner readings) unify practical optimizer directions including Adam boxes, sign maps, and spectral polar maps.
    Builds on cited theory (Bernstein, Pethick, Sfyraki) and is extended to four axes in §3.2.
  • ad hoc to paper Each optimizer has a unique primary incremental mechanism for Dimension-A labeling.
    Taxonomy design principle §4.1; required for non-overlapping T1–T5 families.
  • domain assumption Classical stochastic optimization setup (ERM, mini-batch gradients, AdamW baseline memory model) for LLM training.
    §2 preliminaries; standard in the field.
  • ad hoc to paper Effect objectives O1–O6 are the right multi-objective evaluation axes for optimizer selection.
    Dimension B §4.3; drives benchmark design and family assessments.
invented entities (3)
  • Universal Meta-Pipeline (S0–S5)
    purpose: Locate where each optimizer intervenes in a single update and support identity-mapping classification.
    Paper-defined operational abstraction; not independently measured outside this organizational use.
  • Four-axis LMO decomposition (domain, state estimator, geometry/precondition, finalization)
    purpose: Give each optimizer a coordinate tuple unifying direction geometry and practical state choices.
    Extension of prior LMO ideas into a survey coordinate system specific to this paper.
  • Dual-dimension taxonomy T1–T5 / O1–O6 over 108 optimizers
    purpose: Non-overlapping mechanism families plus multi-label effect targets for survey and benchmark grouping.
    Core invented organizational structure; empirical usefulness is tested but the labels are definitional.

how reviews work

0 comments
Cite this review

Pith. "Pith review of OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers." pith.science (2026). https://pith.science/paper/D75ZLFEO

@misc{pith2026260704033,
  author       = {Pith},
  title        = {Pith review of: OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/D75ZLFEO}},
  note         = {Machine review of arXiv:2607.04033}
}
read the original abstract

Optimizer selection for large-scale model training has become a system-level design decision constrained jointly by compute, memory, tuning budget, and task diversity, yet the landscape of over one hundred methods remains fragmented. We therefore present OmniOpt, a unified survey and benchmark cookbook of optimizers for the research community. OmniOpt rests on four coupled components. First, we treat every optimizer update as a structured transformation through a five-stage meta-pipeline, and show that most methods engage only one or two of these stages. Second, we use norm-constrained linear minimization oracles (LMOs) to unify different optimizers. Third, these two views ground a dual-dimension taxonomy, one dimension assigning each method to a mechanism family and the other recording the measurable training objectives it aims to improve. Fourth, and at the core of this paper, we instantiate the full taxonomy in a unified cross-domain benchmark spanning representative optimizers, model scales, and training regimes from language model pretraining to image classification, systematically analyzing each method family across multiple effect objectives and laying out their trade-offs. OmniOpt thus supplies the research community with an operational coordinate system for selecting optimizers under explicit mechanism and objective assumptions, and charts a direction for the future development of the optimizer community.

Figures

Figures reproduced from arXiv: 2607.04033 by the authors.

Figure 1
Figure 1. Overview of the proposed survey and benchmark framework for a wild range of optimizers. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Evolution of optimizer design for deep learning and LLM training. The expanding coverage [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Universal meta-pipeline for one optimizer step. The training system provides a gradient or [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figures from the paper (17 more)
Figure 4
Figure 4. Figure 4: Mechanism overview of T1 element-wise adaptive-moment and scalar control methods. [PITH_FULL_IMAGE:figures/full_fig_p027_4.png]
Figure 5
Figure 5. Figure 5: Taxonomy of element-wise adaptive-moment and scalar-control optimizers. [PITH_FULL_IMAGE:figures/full_fig_p028_5.png]
Figure 6
Figure 6. Figure 6: Mechanism schematic for T2 matrix-level structural methods. The schematic summarizes the [PITH_FULL_IMAGE:figures/full_fig_p034_6.png]
Figure 7
Figure 7. Figure 7: Taxonomy of matrix-level structural optimizers. [PITH_FULL_IMAGE:figures/full_fig_p035_7.png]
Figure 9
Figure 9. Figure 9: Taxonomy of discretized and directionally quantized optimizers. [PITH_FULL_IMAGE:figures/full_fig_p039_9.png]
Figure 8
Figure 8. Figure 8: Mechanism schematic for T3 discretiza￾tion and directional quantization. This schematic provides a taxonomy guide rather than an empir￾ical ranking for the compact T3 family, grouping discretization and directional quantization mecha￾nisms by their sign-direction gener…
Figure 10
Figure 10. Figure 10: Mechanism schematic for T4 state compression and structural aggregation. The schematic organizes the family by the form of memory reduction: factored second-moment storage, low-bit state representation, shared adaptive statistics, and streaming gradient consumption. T…
Figure 11
Figure 11. Figure 11: Taxonomy of state-compressed and structurally aggregated optimizers. [PITH_FULL_IMAGE:figures/full_fig_p045_11.png]
Figure 12
Figure 12. Figure 12: Mechanism schematic for T5 curvature-aware and geometric regularization methods. The [PITH_FULL_IMAGE:figures/full_fig_p049_12.png]
Figure 13
Figure 13. Figure 13: Taxonomy of curvature-aware and geometry-regularized optimizers. The prefix C- denotes [PITH_FULL_IMAGE:figures/full_fig_p050_13.png]
Figure 14
Figure 14. Figure 14: Stage-1 Pareto frontiers (1B). PPL vs. per-step runtime (left) and vs. optimizer-state memory (right); lower-left is better. Colors denote families, stars mark frontier members. families therefore occupy different favorable regions in quality, runtime, and memory. Opt…
Figure 15
Figure 15. Figure 15: Optimizer-level heatmap of the three Stage-1 metrics (1B). Rows are grouped by family, and columns are C4 PPL, runtime, and optimizer-state memory. Green is favorable, red unfavorable. Optimizer-level heatmap. To complement the Pareto plots, [PITH_FULL_IMAGE:figures/…
Figure 16
Figure 16. Figure 16: Cross-scenario rank stability (FineWeb-Edu, 32k). Absolute ranks among all twelve optimizers per scenario. (a) AdamW vs. its T1 family; (b) AdamW vs. other families’ mean rank; (c) per-architecture mean rank per optimizer (color = family, hatch = architecture). aggreg…
Figure 17
Figure 17. Figure 17: Auxiliary O4 stability analysis from gradient-norm dynamics across architectures. Each cell shows the coefficient of variation of the gradient norm (GNormCV) for one optimizer in one architecture-scale scenario of the FineWeb-Edu long-context benchmark. Lower GNormCV …
Figure 18
Figure 18. Figure 18: Auxiliary learning-rate perturbation robustness. Each panel shows WikiText PPL under 0.2×, 1×, and 5× the tuned learning rate. The star marks the tuned learning rate. The sensitivity score sLR is shown as s in each panel title and measures the worst relative PPL degra…
Figure 19
Figure 19. Figure 19: Family-level objective summary. Per-family profile over O1–O6. O1–O3 are directly measured from Stage 1; O4 is measured from gradient-norm stability; O5 is probed by an auxiliary learning-rate perturbation test; O6 is measured through Stage 2 rank stability and sequen…
Figure 20
Figure 20. Figure 20: Mechanistic ablation of Muon (C4, 350M). Decomposition into core operations, gain operations, and operator-order constraints, with bars showing absolute PPL (lower is better). The 70.74 bar is truncated off-scale. Tiered-summary takeaway. Optimizer selection should be…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

142 extracted references · 56 linked inside Pith

  1. [1]

    Exadam: The power of adaptive cross-moments.arXiv preprint arXiv:2412.20302, 2024

    Ahmed M Adly. Exadam: The power of adaptive cross-moments.arXiv preprint arXiv:2412.20302, 2024

  2. [2]

    Dion: Distributed orthonormalized updates.arXiv preprint arXiv:2504.05295, 2025

    Kwangjun Ahn, Byron Xu, Natalie Abreu, Ying Fan, Gagik Magakyan, Pratyusha Sharma, Zheng Zhan, and John Langford. Dion: Distributed orthonormalized updates.arXiv preprint arXiv:2504.05295, 2025

  3. [3]

    Development of deep learning optimizers: Approaches, concepts, and update rules.arXiv preprint arXiv:2509.18396, 2025

    Do˘gay Altınel. Development of deep learning optimizers: Approaches, concepts, and update rules.arXiv preprint arXiv:2509.18396, 2025

  4. [4]

    Natural gradient works efficiently in learning.Neural computation, 10(2):251–276, 1998

    Shun-Ichi Amari. Natural gradient works efficiently in learning.Neural computation, 10(2):251–276, 1998

  5. [5]

    A pid controller approach for stochastic optimization of deep networks

    Wangpeng An, Haoqian Wang, Qingyun Sun, Jun Xu, Qionghai Dai, and Lei Zhang. A pid controller approach for stochastic optimization of deep networks. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 8522–8531, 2018

  6. [6]

    Memory efficient adaptive optimization

    Rohan Anil, Vineet Gupta, Tomer Koren, and Yoram Singer. Memory efficient adaptive optimization. Advances in Neural Information Processing Systems, 32, 2019

  7. [7]

    Affine-scaled attention: Towards flexible and stable transformer attention.arXiv preprint arXiv:2602.23057, 2026

    Jeongin Bae, Baeseong Park, Gunho Park, Minsub Kim, Joonhyung Lee, Junhee Yoo, Sunghyeon Woo, Jiwon Ryu, Se Jung Kwon, and Dongsoo Lee. Affine-scaled attention: Towards flexible and stable transformer attention.arXiv preprint arXiv:2602.23057, 2026

  8. [8]

    Gravity optimizer: a kinematic approach on optimization in deep learning.arXiv preprint arXiv:2101.09192, 2021

    Dariush Bahrami and Sadegh Pouriyan Zadeh. Gravity optimizer: a kinematic approach on optimization in deep learning.arXiv preprint arXiv:2101.09192, 2021

Show all 142 references
  1. [9]

    Dissecting adam: The sign, magnitude and variance of stochastic gradients

    Lukas Balles and Philipp Hennig. Dissecting adam: The sign, magnitude and variance of stochastic gradients. InInternational Conference on Machine Learning, pages 404–413. PMLR, 2018

  2. [10]

    Old optimizer, new norm: An anthology.arXiv preprint arXiv:2409.20325, 2024

    Jeremy Bernstein and Laker Newhouse. Old optimizer, new norm: An anthology.arXiv preprint arXiv:2409.20325, 2024

  3. [11]

    signsgd: Compressed optimisation for non-convex problems

    Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar. signsgd: Compressed optimisation for non-convex problems. InInternational conference on machine learning, pages 560–569. PMLR, 2018

  4. [12]

    Training neural networks for and by interpola- tion

    Leonard Berrada, Andrew Zisserman, and M Pawan Kumar. Training neural networks for and by interpola- tion. InInternational conference on machine learning, pages 799–809. PMLR, 2020

  5. [13]

    PIQA: Reasoning about physical commonsense in natural language

    Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. PIQA: Reasoning about physical commonsense in natural language. InProceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7432–7439, 2020. doi: 10.1609/aaai.v34i05.6239

  6. [14]

    High-performance large-scale image recognition without normalization

    Andy Brock, Soham De, Samuel L Smith, and Karen Simonyan. High-performance large-scale image recognition without normalization. InInternational conference on machine learning, pages 1059–1071. PMLR, 2021

  7. [15]

    Grams: Gradient descent with adaptive momentum scaling.arXiv preprint arXiv:2412.17107, 2024

    Yang Cao, Xiaoyu Li, and Zhao Song. Grams: Gradient descent with adaptive momentum scaling.arXiv preprint arXiv:2412.17107, 2024

  8. [16]

    Mgup: A momentum-gradient alignment update policy for stochastic optimization.Advances in Neural Information Processing Systems, 38:20488–20537, 2026

    Da Chang and Ganzhao Yuan. Mgup: A momentum-gradient alignment update policy for stochastic optimization.Advances in Neural Information Processing Systems, 38:20488–20537, 2026

  9. [17]

    Closing the generaliza- tion gap of adaptive gradient methods in training deep neural networks.arXiv preprint arXiv:1806.06763, 2018

    Jinghui Chen, Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu. Closing the generaliza- tion gap of adaptive gradient methods in training deep neural networks.arXiv preprint arXiv:1806.06763, 2018

  10. [18]

    Fira: Can we achieve full-rank training of llms under low-rank constraint?Advances in Neural Information Processing Systems, 38:120680–120712, 2026

    Xi Chen, Kaituo Feng, Changsheng Li, Xunhao Lai, Xiangyu Yue, Ye Yuan, and Guoren Wang. Fira: Can we achieve full-rank training of llms under low-rank constraint?Advances in Neural Information Processing Systems, 38:120680–120712, 2026

  11. [19]

    Symbolic discovery of optimization algorithms.Advances in neural information processing systems, 36:49205–49233, 2023

    Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, et al. Symbolic discovery of optimization algorithms.Advances in neural information processing systems, 36:49205–49233, 2023. 78 OmniOpt: Taxonomy, ...

  12. [20]

    BoolQ: Exploring the surprising difficulty of natural yes/no questions

    Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. BoolQ: Exploring the surprising difficulty of natural yes/no questions. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computatio...

  13. [21]

    Think you have solved question answering? try arc, the ai2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018

  14. [22]

    Why gradients rapidly increase near the end of training.arXiv preprint arXiv:2506.02285, 2025

    Aaron Defazio. Why gradients rapidly increase near the end of training.arXiv preprint arXiv:2506.02285, 2025

  15. [23]

    A momentumized, adaptive, dual averaged gradient method.Journal of Machine Learning Research, 23(144):1–34, 2022

    Aaron Defazio and Samy Jelassi. A momentumized, adaptive, dual averaged gradient method.Journal of Machine Learning Research, 23(144):1–34, 2022

  16. [24]

    Learning-rate-free learning by d-adaptation

    Aaron Defazio and Konstantin Mishchenko. Learning-rate-free learning by d-adaptation. InInternational conference on machine learning, pages 7449–7479. PMLR, 2023

  17. [25]

    The road less scheduled.Advances in Neural Information Processing Systems, 37:9974–10007, 2024

    Aaron Defazio, Xingyu Yang, Harsh Mehta, Konstantin Mishchenko, Ahmed Khaled, and Ashok Cutkosky. The road less scheduled.Advances in Neural Information Processing Systems, 37:9974–10007, 2024

  18. [26]

    Rmnp: Row-momentum normalized preconditioning for scalable matrix-based optimization.arXiv preprint arXiv:2603.20527, 2026

    Shenyang Deng, Zhuoli Ouyang, Tianyu Pang, Zihang Liu, Ruochen Jin, Shuhua Yu, and Yaoqing Yang. Rmnp: Row-momentum normalized preconditioning for scalable matrix-based optimization.arXiv preprint arXiv:2603.20527, 2026

  19. [27]

    8-bit optimizers via block-wise quantization

    Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer. 8-bit optimizers via block-wise quantization. arXiv preprint arXiv:2110.02861, 2021

  20. [28]

    An adaptive and momental bound method for stochastic learning.arXiv preprint arXiv:1910.12249, 2019

    Jianbang Ding, Xuancheng Ren, Ruixuan Luo, and Xu Sun. An adaptive and momental bound method for stochastic learning.arXiv preprint arXiv:1910.12249, 2019

  21. [29]

    Incorporating nesterov momentum into adam, 2016

    Timothy Dozat. Incorporating nesterov momentum into adam, 2016

  22. [30]

    diffgrad: an optimization method for convolutional neural networks.IEEE transactions on neural networks and learning systems, 31(11):4500–4511, 2019

    Shiv Ram Dubey, Soumendu Chakraborty, Swalpa Kumar Roy, Snehasis Mukherjee, Satish Kumar Singh, and Bidyut Baran Chaudhuri. diffgrad: an optimization method for convolutional neural networks.IEEE transactions on neural networks and learning systems, 31(11):4500–4511, 2019

  23. [31]

    Adaptive subgradient methods for online learning and stochastic optimization.Journal of machine learning research, 12(7), 2011

    John Duchi, Elad Hazan, and Yoram Singer. Adaptive subgradient methods for online learning and stochastic optimization.Journal of machine learning research, 12(7), 2011

  24. [32]

    Sharpness-aware minimization for efficiently improving generalization.arXiv preprint arXiv:2010.01412, 2020

    Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. Sharpness-aware minimization for efficiently improving generalization.arXiv preprint arXiv:2010.01412, 2020

  25. [33]

    A stable whitening optimizer for efficient neural network training.Advances in Neural Information Processing Systems, 38:174086–174110, 2026

    Kevin Frans, Sergey Levine, and Pieter Abbeel. A stable whitening optimizer for efficient neural network training.Advances in Neural Information Processing Systems, 38:174086–174110, 2026

  26. [34]

    The language model evaluation harness, 07 2024

    Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang S...

  27. [35]

    Stochastic gradient methods with layer-wise adaptive moments for training of deep networks.arXiv preprint arXiv:1905.11286, 2019

    Boris Ginsburg, Patrice Castonguay, Oleksii Hrinchuk, Oleksii Kuchaiev, Vitaly Lavrukhin, Ryan Leary, Jason Li, Huyen Nguyen, Yang Zhang, and Jonathan M Cohen. Stochastic gradient methods with layer-wise adaptive moments for training of deep networks.arXiv preprint arXiv:1905....

  28. [36]

    Scalable parameter and memory efficient pretraining for llm: Recent algorithmic advances and benchmarking.arXiv preprint arXiv:2505.22922, 2025

    Athanasios Glentis, Jiaxiang Li, Qiulin Shang, Andi Han, Ioannis Tsaknakis, Quan Wei, and Mingyi Hong. Scalable parameter and memory efficient pretraining for llm: Recent algorithmic advances and benchmarking.arXiv preprint arXiv:2505.22922, 2025

  29. [37]

    Towards efficient optimizer design for llm via structured fisher approximation with a low-rank extension.arXiv preprint arXiv:2502.07752, 2025

    Wenbo Gong, Meyer Scetbon, Chao Ma, and Edward Meeds. Towards efficient optimizer design for llm via structured fisher approximation with a low-rank extension.arXiv preprint arXiv:2502.07752, 2025

  30. [38]

    Shampoo: Preconditioned stochastic tensor optimization

    Vineet Gupta, Tomer Koren, and Yoram Singer. Shampoo: Preconditioned stochastic tensor optimization. InInternational Conference on Machine Learning, pages 1842–1850. PMLR, 2018. 79 OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers

  31. [39]

    Gcsam: Gradient centralized sharpness aware minimization.IEEE Access, 2025

    Mohamed Hassan, Aleksandar Vakanski, Boyu Zhang, and Min Xian. Gcsam: Gradient centralized sharpness aware minimization.IEEE Access, 2025

  32. [40]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016

  33. [41]

    Orthograd improves neural calibration.arXiv preprint arXiv:2506.04487, 2025

    C Evans Hedges. Orthograd improves neural calibration.arXiv preprint arXiv:2506.04487, 2025

  34. [42]

    Adamp: Slowing down the slowdown for momentum optimizers on scale-invariant weights.arXiv preprint arXiv:2006.08217, 2020

    Byeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han, Sangdoo Yun, Gyuwan Kim, Youngjung Uh, and Jung-Woo Ha. Adamp: Slowing down the slowdown for momentum optimizers on scale-invariant weights.arXiv preprint arXiv:2006.08217, 2020

  35. [43]

    Gradientstabilizer: Fix the norm, not the gradient

    Tianjin Huang, Zhangyang Wang, Haotian Hu, Zhenyu Zhang, Gaojie Jin, Xiang Li, Li Shen, Jiaxing Shang, Tianlong Chen, Ke Li, et al. Gradientstabilizer: Fix the norm, not the gradient

  36. [44]

    Spam: Spike-aware adam with momentum reset for stable llm training.arXiv preprint arXiv:2501.06842, 2025

    Tianjin Huang, Ziquan Zhu, Gaojie Jin, Lu Liu, Zhangyang Wang, and Shiwei Liu. Spam: Spike-aware adam with momentum reset for stable llm training.arXiv preprint arXiv:2501.06842, 2025

  37. [45]

    Dog is sgd’s best friend: A parameter-free dynamic step size schedule, 2023

    Maor Ivgi, Oliver Hinder, and Yair Carmon. Dog is sgd’s best friend: A parameter-free dynamic step size schedule, 2023

  38. [46]

    Adamd: Improved bias-correction in adam.arXiv preprint arXiv:2110.10828, 2021

    John St John. Adamd: Improved bias-correction in adam.arXiv preprint arXiv:2110.10828, 2021

  39. [47]

    On surprising effectiveness of masking updates in adaptive optimizers.arXiv preprint arXiv:2602.15322, 2026

    Taejong Joo, Wenhan Xia, Cheolmin Kim, Ming Zhang, and Eugene Ie. On surprising effectiveness of masking updates in adaptive optimizers.arXiv preprint arXiv:2602.15322, 2026

  40. [48]

    Muon: An optimizer for hidden layers in neural networks

    Keller Jordan, Yuchen Jin, Vlado Boza, Jiacheng You, Franz Cesista, Laker Newhouse, and Jeremy Bernstein. Muon: An optimizer for hidden layers in neural networks. https://kellerjordan.github.io/posts/muon/, 2024

  41. [49]

    Fineweb-edu-100b-shuffle

    Andrej Karpathy. Fineweb-edu-100b-shuffle. https://huggingface.co/datasets/karpathy/ fineweb-edu-100b-shuffle, 2024

  42. [50]

    Ano: Faster is better in noisy landscape.arXiv preprint arXiv:2508.18258, 2025

    Adrien Kegreisz. Ano: Faster is better in noisy landscape.arXiv preprint arXiv:2508.18258, 2025

  43. [51]

    Improving generalization performance by switching from adam to sgd.arXiv preprint arXiv:1712.07628, 2017

    Nitish Shirish Keskar and Richard Socher. Improving generalization performance by switching from adam to sgd.arXiv preprint arXiv:1712.07628, 2017

  44. [52]

    Dowg unleashed: An efficient universal parameter- free gradient descent method.Advances in Neural Information Processing Systems, 36:6748–6769, 2023

    Ahmed Khaled, Konstantin Mishchenko, and Chi Jin. Dowg unleashed: An efficient universal parameter- free gradient descent method.Advances in Neural Information Processing Systems, 36:6748–6769, 2023

  45. [53]

    On the insufficiency of existing momentum schemes for stochastic optimization, 2018

    Rahul Kidambi, Praneeth Netrapalli, Prateek Jain, and Sham Kakade. On the insufficiency of existing momentum schemes for stochastic optimization, 2018

  46. [54]

    Adam: A method for stochastic optimization.arXiv preprint arXiv:1412.6980, 2014

    Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization.arXiv preprint arXiv:1412.6980, 2014

  47. [55]

    Learning multiple layers of features from tiny images, 2009

    Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images, 2009

  48. [56]

    Zclip: Adaptive spike mitigation for llm pre-training.arXiv preprint arXiv:2504.02507, 2025

    Abhay Kumar, Louis Owen, Nilabhra Roy Chowdhury, and Fabian G¨ura. Zclip: Adaptive spike mitigation for llm pre-training.arXiv preprint arXiv:2504.02507, 2025

  49. [57]

    Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks

    Jungmin Kwon, Jeongseop Kim, Hyunseo Park, and In Kwon Choi. Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks. InInternational conference on machine learning, pages 5905–5914. PMLR, 2021

  50. [58]

    Unveiling the backbone-optimizer coupling bias in visual representation learning.arXiv preprint arXiv:2410.06373, 2024

    Siyuan Li, Juanxi Tian, Zedong Wang, Luyuan Zhang, Zicheng Liu, Weiyang Jin, Yang Liu, Baigui Sun, and Stan Z Li. Unveiling the backbone-optimizer coupling bias in visual representation learning.arXiv preprint arXiv:2410.06373, 2024

  51. [59]

    Taming llms by scaling learning rates with gradient grouping.arXiv preprint arXiv:2506.01049, 2025

    Siyuan Li, Juanxi Tian, Zedong Wang, Xin Jin, Zicheng Liu, Wentao Zhang, and Dan Xu. Taming llms by scaling learning rates with gradient grouping.arXiv preprint arXiv:2506.01049, 2025

  52. [60]

    SAC: Adaptive learning rate scaling with architectural constraints, 2026

    Siyuan Li, Juanxi Tian, Zedong Wang, Anna Wang, Xin Jin, Chang Yu, Ruoyu Sun, and Cheng Tan. SAC: Adaptive learning rate scaling with architectural constraints, 2026. URL https://openreview.net/forum? id=EB92tITeNq. 80 OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers

  53. [61]

    Friendly sharpness-aware minimiza- tion

    Tao Li, Pan Zhou, Zhengbao He, Xinwen Cheng, and Xiaolin Huang. Friendly sharpness-aware minimiza- tion. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5631–5640, 2024

  54. [62]

    Preconditioned stochastic gradient descent.IEEE transactions on neural networks and learning systems, 29(5):1454–1466, 2017

    Xi-Lin Li. Preconditioned stochastic gradient descent.IEEE transactions on neural networks and learning systems, 29(5):1454–1466, 2017

  55. [63]

    Black box lie group preconditioners for sgd.arXiv preprint arXiv:2211.04422, 2022

    Xilin Li. Black box lie group preconditioners for sgd.arXiv preprint arXiv:2211.04422, 2022

  56. [64]

    Cautious optimizers: Improving training with one line of code.arXiv preprint arXiv:2411.16085, 2024

    Kaizhao Liang, Lizhang Chen, Bo Liu, and Qiang Liu. Cautious optimizers: Improving training with one line of code.arXiv preprint arXiv:2411.16085, 2024

  57. [65]

    Sophia: A scalable stochastic second-order optimizer for language model pre-training

    Hong Liu, Zhiyuan Li, David Hall, Percy Liang, and Tengyu Ma. Sophia: A scalable stochastic second-order optimizer for language model pre-training. InInternational Conference on Learning Representations, volume 2024, pages 1621–1650, 2024

  58. [66]

    Cosmos: A hybrid adaptive optimizer for memory-efficient training of llms.arXiv preprint arXiv:2502.17410, 2025

    Liming Liu, Zhenghao Xu, Zixuan Zhang, Hao Kang, Zichong Li, Chen Liang, Weizhu Chen, and Tuo Zhao. Cosmos: A hybrid adaptive optimizer for memory-efficient training of llms.arXiv preprint arXiv:2502.17410, 2025

  59. [67]

    On the variance of the adaptive learning rate and beyond.arXiv preprint arXiv:1908.03265, 2019

    Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han. On the variance of the adaptive learning rate and beyond.arXiv preprint arXiv:1908.03265, 2019

  60. [68]

    Focus: First order concentrated updating scheme.arXiv preprint arXiv:2501.12243, 2025

    Yizhou Liu, Ziming Liu, and Jeff Gore. Focus: First order concentrated updating scheme.arXiv preprint arXiv:2501.12243, 2025

  61. [69]

    Towards efficient and scalable sharpness- aware minimization

    Yong Liu, Siqi Mai, Xiangning Chen, Cho-Jui Hsieh, and Yang You. Towards efficient and scalable sharpness- aware minimization. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12360–12370, 2022

  62. [70]

    Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017

  63. [71]

    Adasmooth: an adaptive learning rate method based on effective ratio

    Jun Lu. Adasmooth: an adaptive learning rate method based on effective ratio. InSentiment Analysis and Deep Learning: Proceedings of ICSADL 2022, pages 273–293. Springer, 2023

  64. [72]

    Aggregated momentum: Stability through passive damping.arXiv preprint arXiv:1804.00325, 2018

    James Lucas, Shengyang Sun, Richard Zemel, and Roger Grosse. Aggregated momentum: Stability through passive damping.arXiv preprint arXiv:1804.00325, 2018

  65. [73]

    Adaptive gradient methods with dynamic bound of learning rate.arXiv preprint arXiv:1902.09843, 2019

    Liangchen Luo, Yuanhao Xiong, Yan Liu, and Xu Sun. Adaptive gradient methods with dynamic bound of learning rate.arXiv preprint arXiv:1902.09843, 2019

  66. [74]

    Came: Confidence-guided adaptive memory efficient optimization

    Yang Luo, Xiaozhe Ren, Zangwei Zheng, Zhuo Jiang, Xin Jiang, and Yang You. Came: Confidence-guided adaptive memory efficient optimization. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4442–4453, 2023

  67. [75]

    Adalomo: Low-memory optimization with adaptive learning rate

    Kai Lv, Hang Yan, Qipeng Guo, Haijun Lv, and Xipeng Qiu. Adalomo: Low-memory optimization with adaptive learning rate. InFindings of the Association for Computational Linguistics: ACL 2024, pages 12486– 12502, 2024

  68. [76]

    Full parameter fine-tuning for large language models with limited resources

    Kai Lv, Yuqing Yang, Tengxiao Liu, Qipeng Guo, and Xipeng Qiu. Full parameter fine-tuning for large language models with limited resources. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8187–8198, 2024

  69. [77]

    Quasi-hyperbolic momentum and adam for deep learning.arXiv preprint arXiv:1810.06801, 2018

    Jerry Ma and Denis Yarats. Quasi-hyperbolic momentum and adam for deep learning.arXiv preprint arXiv:1810.06801, 2018

  70. [78]

    Torque-aware momentum.arXiv preprint arXiv:2412.18790, 2024

    Pranshu Malviya, Goncalo Mordido, Aristide Baratin, Reza Babanezhad Harikandeh, Gintare Karolina Dziugaite, Razvan Pascanu, and Sarath Chandar. Torque-aware momentum.arXiv preprint arXiv:2412.18790, 2024

  71. [79]

    Pointer sentinel mixture models

    Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843, 2016. 81 OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers

  72. [80]

    Can a suit of armor conduct electricity? a new dataset for open book question answering

    Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. InProceedings of the 2018 conference on empirical methods in natural language processing, pages 2381–2391, 2018

  73. [81]

    Prodigy: An expeditiously adaptive parameter-free learner

    Konstantin Mishchenko and Aaron Defazio. Prodigy: An expeditiously adaptive parameter-free learner. arXiv preprint arXiv:2306.06101, 2023

  74. [82]

    Sam as an optimal relaxation of bayes.arXiv preprint arXiv:2210.01620, 2022

    Thomas M¨ollenhoff and Mohammad Emtiyaz Khan. Sam as an optimal relaxation of bayes.arXiv preprint arXiv:2210.01620, 2022

  75. [83]

    Connections between schedule-free optimizers, ademamix, and accelerated sgd variants.arXiv preprint arXiv:2502.02431, 2025

    Depen Morwani, Nikhil Vyas, Hanlin Zhang, and Sham Kakade. Connections between schedule-free optimizers, ademamix, and accelerated sgd variants.arXiv preprint arXiv:2502.02431, 2025

  76. [84]

    The ademamix optimizer: Better, faster, older, 2025

    Matteo Pagliardini, Pierre Ablin, and David Grangier. The ademamix optimizer: Better, faster, older, 2025

  77. [85]

    The LAMBADA dataset: Word prediction requiring a broad discourse context

    Denis Paperno, Germ´an Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernandez. The LAMBADA dataset: Word prediction requiring a broad discourse context. InProceedings of the 54th Annual Meeting of t...

  78. [86]

    The fineweb datasets: Decanting the web for the finest text data at scale, 2024

    Guilherme Penedo, Hynek Kydl´ıˇcek, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. The fineweb datasets: Decanting the web for the finest text data at scale, 2024. URLhttps://arxiv.org/abs/2406.17557

  79. [87]

    Training deep learning models with norm-constrained lmos.arXiv preprint arXiv:2502.07529, 2025

    Thomas Pethick, Wanyun Xie, Kimon Antonakopoulos, Zhenyu Zhu, Antonio Silveti-Falls, and Volkan Cevher. Training deep learning models with norm-constrained lmos.arXiv preprint arXiv:2502.07529, 2025

  80. [88]

    Some methods of speeding up the convergence of iteration methods.Ussr computational mathematics and mathematical physics, 4(5):1–17, 1964

    Boris T Polyak. Some methods of speeding up the convergence of iteration methods.Ussr computational mathematics and mathematical physics, 4(5):1–17, 1964

  81. [89]

    Can muon fine-tune adam-pretrained models?arXiv preprint arXiv:2605.10468, 2026

    Xingyu Qu, Peigeng Huang, and Samuel Horvath. Can muon fine-tune adam-pretrained models?arXiv preprint arXiv:2605.10468, 2026

  82. [90]

    Exploring the limits of transfer learning with a unified text-to-text transformer

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67, 2020

  83. [91]

    Navigating llm valley: From adamw to memory-efficient and matrix-based optimizers

    Aditya Ranganath. Navigating llm valley: From adamw to memory-efficient and matrix-based optimizers. arXiv preprint arXiv:2605.09176, 2026

  84. [92]

    Choice of plausible alternatives: An evaluation of commonsense causal reasoning

    Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In2011 AAAI Spring Symposium Series, 2011. URL https: //people.ict.usc.edu/∼gordon/publications/AAAI-SPRING11A.PDF

  85. [93]

    Rlion: A refined lion optimizer for deep learning, 2024

    Jian Rong, ChenHao Ma, QingHui Zhang, and Yong Cao. Rlion: A refined lion optimizer for deep learning, 2024

  86. [94]

    Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021

    Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021

  87. [95]

    No more pesky learning rates

    Tom Schaul, Sixin Zhang, and Yann LeCun. No more pesky learning rates. InInternational conference on machine learning, pages 343–351. PMLR, 2013

  88. [96]

    Pre-training llms on a budget: A comparison of three optimizers.arXiv preprint arXiv:2507.08472, 2025

    Joel Schlotthauer, Christian Kroos, Chris Hinze, Viktor Hangya, Luzian Hahn, and Fabian K¨uch. Pre-training llms on a budget: A comparison of three optimizers.arXiv preprint arXiv:2507.08472, 2025

  89. [97]

    Benchmarking optimizers for large language model pretraining.arXiv preprint arXiv:2509.01440, 2025

    Andrei Semenov, Matteo Pagliardini, and Martin Jaggi. Benchmarking optimizers for large language model pretraining.arXiv preprint arXiv:2509.01440, 2025

  90. [98]

    Lions and muons: Optimization via stochastic frank-wolfe.arXiv preprint arXiv:2506.04192, 2025

    Maria-Eleni Sfyraki and Jun-Kun Wang. Lions and muons: Optimization via stochastic frank-wolfe.arXiv preprint arXiv:2506.04192, 2025

  91. [99]

    Adafactor: Adaptive learning rates with sublinear memory cost

    Noam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. In International conference on machine learning, pages 4596–4604. PMLR, 2018. 82 OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers

  92. [100]

    Muon is not that special: Random or inverted spectra work just as well.arXiv preprint arXiv:2605.11181, 2026

    Zakhar Shumaylov, Natha¨el Da Costa, Peter Zaika, B´alint Mucs´anyi, Alex Massucco, Yoav Gelberg, Carola- Bibiane Sch¨onlieb, Yarin Gal, and Philipp Hennig. Muon is not that special: Random or inverted spectra work just as well.arXiv preprint arXiv:2605.11181, 2026

  93. [101]

    Adamuon: Adaptive muon optimizer.arXiv preprint arXiv:2507.11005, 2025

    Chongjie Si, Debing Zhang, and Wei Shen. Adamuon: Adaptive muon optimizer.arXiv preprint arXiv:2507.11005, 2025

  94. [102]

    Through the river: Understanding the benefit of schedule-free methods for language model training.Advances in Neural Information Processing Systems, 38:127524–127555, 2026

    Minhak Song, Beomhan Baek, Kwangjun Ahn, and Chulhee Yun. Through the river: Understanding the benefit of schedule-free methods for language model training.Advances in Neural Information Processing Systems, 38:127524–127555, 2026

  95. [103]

    On the importance of initialization and momentum in deep learning

    Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton. On the importance of initialization and momentum in deep learning. InInternational conference on machine learning, pages 1139–1147. pmlr, 2013

  96. [104]

    ADOPT: Modified Adam can converge with anyβ 2 with the optimal rate.Advances in Neural Information Processing Systems, 37:72438–72474, 2024

    Shohei Taniguchi, Keno Harada, Gouki Minegishi, Yuta Oshima, Seong Cheol Jeong, Go Nagahara, Tomoshi Iiyama, Masahiro Suzuki, Yusuke Iwasawa, and Yutaka Matsuo. ADOPT: Modified Adam can converge with anyβ 2 with the optimal rate.Advances in Neural Information Processing System...

  97. [105]

    Divide the gradient by a running average of its recent magnitude

    Tijmen Tieleman and Geoffrey Hinton. Divide the gradient by a running average of its recent magnitude. coursera: Neural networks for machine learning. InTechnical report. University of Toronto, 2017

  98. [106]

    Training data-efficient image transformers & distillation through attention

    Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herv´e J´egou. Training data-efficient image transformers & distillation through attention. InInternational conference on machine learning, pages 10347–10357. PMLR, 2021

  99. [107]

    Attention is all you need.Advances in neural information processing systems, 30, 2017

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez,Łukasz Kaiser, and Illia Polosukhin. Attention is all you need.Advances in neural information processing systems, 30, 2017

  100. [108]

    Soap: Improving and stabilizing shampoo using adam for language modeling

    Nikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira, David Brandfonbrener, Lucas Janson, and Sham Kakade. Soap: Improving and stabilizing shampoo using adam for language modeling. InInternational Conference on Learning Representations, volume 2025, pages 93423–93444, 2025

  101. [109]

    Adagc: Improving training stability for large language model pretraining.arXiv preprint arXiv:2502.11034, 2025

    Guoxia Wang, Shuai Li, Congliang Chen, Jinle Zeng, Jiabin Yang, Dianhai Yu, Yanjun Ma, and Li Shen. Adagc: Improving training stability for large language model pretraining.arXiv preprint arXiv:2502.11034, 2025

  102. [110]

    The sharpness disparity principle in transformers for accelerating language model pre-training

    Jinbo Wang, Mingze Wang, Zhanpeng Zhou, Junchi Yan, Lei Wu, et al. The sharpness disparity principle in transformers for accelerating language model pre-training. InInternational Conference on Machine Learning, pages 64859–64879. PMLR, 2025

  103. [111]

    Conda: Column-normalized adam for training large language models faster.arXiv preprint arXiv:2509.24218, 2025

    Junjie Wang, Pan Zhou, Yiming Dong, Huan Li, Jia Li, Xun Zhou, Qicheng Lao, Cong Fang, and Zhouchen Lin. Conda: Column-normalized adam for training large language models faster.arXiv preprint arXiv:2509.24218, 2025

  104. [112]

    Crowdsourcing multiple choice science questions

    Johannes Welbl, Nelson F Liu, and Matt Gardner. Crowdsourcing multiple choice science questions. In Proceedings of the 3rd Workshop on Noisy User-generated Text, pages 94–106, 2017

  105. [113]

    Fantastic pretraining optimizers and where to find them.arXiv preprint arXiv:2509.02046, 2025

    Kaiyue Wen, David Hall, Tengyu Ma, and Percy Liang. Fantastic pretraining optimizers and where to find them.arXiv preprint arXiv:2509.02046, 2025

  106. [114]

    Stable and low-precision training for large-scale vision-language models.Advances in Neural Information Processing Systems, 36:10271–10298, 2023

    Mitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, Ari Morcos, Ali Farhadi, and Ludwig Schmidt. Stable and low-precision training for large-scale vision-language models.Advances in Neural Information Processing Systems, 36:10271–10298, 2023

  107. [115]

    Ranger21: a synergistic deep learning optimizer.arXiv preprint arXiv:2106.13731, 2021

    Less Wright and Nestor Demeure. Ranger21: a synergistic deep learning optimizer.arXiv preprint arXiv:2106.13731, 2021

  108. [116]

    Controlled llm training on spectral sphere.arXiv preprint arXiv:2601.08393, 2026

    Tian Xie, Haoming Luo, Haoyu Tang, Yiwen Hu, Jason Klein Liu, Qingnan Ren, Yang Wang, Wayne Xin Zhao, Rui Yan, Bing Su, et al. Controlled llm training on spectral sphere.arXiv preprint arXiv:2601.08393, 2026

  109. [117]

    Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12):9508–9520, 2024

    Xingyu Xie, Pan Zhou, Huan Li, Zhouchen Lin, and Shuicheng Yan. Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12):9508–9520, 2024. 83 OmniOpt: Taxonomy, Geometry, and Benchmarking...

  110. [118]

    No more adam: Learning rate scaling at initializa- tion is all you need.arXiv preprint arXiv:2412.11768, 2024

    Minghao Xu, Lichuan Xiang, Xu Cai, and Hongkai Wen. No more adam: Learning rate scaling at initializa- tion is all you need.arXiv preprint arXiv:2412.11768, 2024

  111. [119]

    On the width scaling of neural optimizers under matrix operator norms i: Row/column normalization and hyperparameter transfer.arXiv preprint arXiv:2603.09952, 2026

    Ruihan Xu, Jiajin Li, and Yiping Lu. On the width scaling of neural optimizers under matrix operator norms i: Row/column normalization and hyperparameter transfer.arXiv preprint arXiv:2603.09952, 2026

  112. [120]

    Gated linear attention trans- formers with hardware-efficient training.arXiv preprint arXiv:2312.06635, 2023

    Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim. Gated linear attention trans- formers with hardware-efficient training.arXiv preprint arXiv:2312.06635, 2023

  113. [121]

    Parallelizing linear transformers with the delta rule over sequence length.Advances in neural information processing systems, 37:115491–115522, 2024

    Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim. Parallelizing linear transformers with the delta rule over sequence length.Advances in neural information processing systems, 37:115491–115522, 2024

  114. [122]

    Gated delta networks: Improving mamba2 with delta rule

    Songlin Yang, Jan Kautz, and Ali Hatamizadeh. Gated delta networks: Improving mamba2 with delta rule. InInternational Conference on Learning Representations, volume 2025, pages 29687–29707, 2025

  115. [123]

    Adahes- sian: An adaptive second order optimizer for machine learning

    Zhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa, Kurt Keutzer, and Michael Mahoney. Adahes- sian: An adaptive second order optimizer for machine learning. Inproceedings of the AAAI conference on artificial intelligence, volume 35, pages 10665–10673, 2021

  116. [124]

    Gradient centralization: A new opti- mization technique for deep neural networks

    Hongwei Yong, Jianqiang Huang, Xiansheng Hua, and Lei Zhang. Gradient centralization: A new opti- mization technique for deep neural networks. InEuropean Conference on Computer Vision, pages 635–652. Springer, 2020

  117. [125]

    Large batch optimization for deep learning: Training bert in 76 minutes.arXiv preprint arXiv:1904.00962, 2019

    Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. Large batch optimization for deep learning: Training bert in 76 minutes.arXiv preprint arXiv:1904.00962, 2019

  118. [126]

    Metaformer is actually what you need for vision

    Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan. Metaformer is actually what you need for vision. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10819–10829, 2022

  119. [127]

    Metaformer baselines for vision.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(2):896–912, 2023

    Weihao Yu, Chenyang Si, Pan Zhou, Mi Luo, Yichen Zhou, Jiashi Feng, Shuicheng Yan, and Xinchao Wang. Metaformer baselines for vision.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(2):896–912, 2023

  120. [128]

    Mars: Unleashing the power of variance reduction for training large models.arXiv preprint arXiv:2411.10438, 2024

    Huizhuo Yuan, Yifeng Liu, Shuang Wu, Xun Zhou, and Quanquan Gu. Mars: Unleashing the power of variance reduction for training large models.arXiv preprint arXiv:2411.10438, 2024

  121. [129]

    Sharpness-aware minimization revisited: Weighted sharpness as a regularization term

    Yun Yue, Jiadi Jiang, Zhiling Ye, Ning Gao, Yongchao Liu, and Ke Zhang. Sharpness-aware minimization revisited: Weighted sharpness as a regularization term. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 3185–3194, 2023

  122. [130]

    Adadelta: an adaptive learning rate method.arXiv preprint arXiv:1212.5701, 2012

    Matthew D Zeiler. Adadelta: an adaptive learning rate method.arXiv preprint arXiv:1212.5701, 2012

  123. [131]

    HellaSwag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. HellaSwag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800. Association for Computational Linguistics, ...

  124. [132]

    Lookahead optimizer: k steps forward, 1 step back.Advances in neural information processing systems, 32, 2019

    Michael Zhang, James Lucas, Jimmy Ba, and Geoffrey E Hinton. Lookahead optimizer: k steps forward, 1 step back.Advances in neural information processing systems, 32, 2019

  125. [133]

    Adagrad meets muon: Adaptive stepsizes for orthogonal updates.arXiv preprint arXiv:2509.02981, 2025

    Minxin Zhang, Yuxuan Liu, and Hayden Schaeffer. Adagrad meets muon: Adaptive stepsizes for orthogonal updates.arXiv preprint arXiv:2509.02981, 2025

  126. [134]

    Evolution of optimization methods: Algorithms, scenarios, and evaluations

    Tong Zhang, Jiangning Zhang, Zhucun Xue, Juntao Jiang, Yicheng Xu, Chengming Xu, Teng Hu, Xingyu Xie, Xiaobin Hu, Yabiao Wang, et al. Evolution of optimization methods: Algorithms, scenarios, and evaluations. arXiv preprint arXiv:2604.12968, 2026

  127. [135]

    Revisiting zeroth-order optimization for memory-efficient llm fine-tuning: A benchmark.arXiv preprint arXiv:2402.11592, 2024

    Yihua Zhang, Pingzhi Li, Junyuan Hong, Jiaxiang Li, Yimeng Zhang, Wenqing Zheng, Pin-Yu Chen, Jason D Lee, Wotao Yin, Mingyi Hong, et al. Revisiting zeroth-order optimization for memory-efficient llm fine-tuning: A benchmark.arXiv preprint arXiv:2402.11592, 2024. 84 OmniOpt: T...

  128. [136]

    Adam-mini: Use fewer learning rates to gain more

    Yushun Zhang, Congliang Chen, Ziniu Li, Tian Ding, Chenwei Wu, Diederik Durk Kingma, Yinyu Ye, Zhi-Quan Luo, and Ruoyu Sun. Adam-mini: Use fewer learning rates to gain more. InInternational Conference on Learning Representations, volume 2025, pages 28033–28063, 2025

  129. [137]

    Q-galore: Quantized galore with int4 projection and layer-adaptive low-rank gradients.arXiv preprint arXiv:2407.08296, 2024

    Zhenyu Zhang, Ajay Jaiswal, Lu Yin, Shiwei Liu, Jiawei Zhao, Yuandong Tian, and Zhangyang Wang. Q-galore: Quantized galore with int4 projection and layer-adaptive low-rank gradients.arXiv preprint arXiv:2407.08296, 2024

  130. [138]

    Galore: Memory-efficient llm training by gradient low-rank projection.arXiv preprint arXiv:2403.03507, 2024

    Jiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang, Anima Anandkumar, and Yuandong Tian. Galore: Memory-efficient llm training by gradient low-rank projection.arXiv preprint arXiv:2403.03507, 2024

  131. [139]

    Deconstructing what makes a good optimizer for language models, 2025.URL https://arxiv

    Rosie Zhao, Depen Morwani, David Brandfonbrener, Nikhil Vyas, and Sham Kakade. Deconstructing what makes a good optimizer for language models, 2025.URL https://arxiv. org/abs/2407.07972

  132. [140]

    Apollo: Sgd-like memory, adamw-level performance.Proceedings of Machine Learning and Systems, 7, 2025

    Hanqing Zhu, Zhenyu Zhang, Wenyan Cong, Xi Liu, Sem Park, Vikas Chandra, Bo Long, David Z Pan, Zhangyang Wang, and Jinwon Lee. Apollo: Sgd-like memory, adamw-level performance.Proceedings of Machine Learning and Systems, 7, 2025

  133. [141]

    Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.Advances in neural information processing systems, 33:18795–18806, 2020

    Juntang Zhuang, Tommy Tang, Yifan Ding, Sekhar C Tatikonda, Nicha Dvornek, Xenophon Papademetris, and James Duncan. Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.Advances in neural information processing systems, 33:18795–18806, 2020

  134. [142]

    Surrogate gap minimization improves sharpness-aware training.arXiv preprint arXiv:2203.08065, 2022

    Juntang Zhuang, Boqing Gong, Liangzhe Yuan, Yin Cui, Hartwig Adam, Nicha Dvornek, Sekhar Tatikonda, James Duncan, and Ting Liu. Surrogate gap minimization improves sharpness-aware training.arXiv preprint arXiv:2203.08065, 2022. 85 OmniOpt: Taxonomy, Geometry, and Benchmarking ...

Pith tools

Reviewed July 11, 2026 · model on record in the stance chip above.