Pith. sign in

Paper Citation Record · LEDGER

Deconstructing What Makes a Good Optimizer for Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2407.07972.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.07972 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:57:29.263146Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:27:30.162768Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d37440de-1ffc-48ba-8fb9-af58f132c7e8 · inbound

Old Optimizer, New Norm: An Anthology cites this paper.

Old Optimizer, New Norm: An Anthology Deconstructing What Makes a Good Optimizer for Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:27:52.955453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T07:27:52.883335Z digest=sha256:056435a66b802d56d01f0a7721a820d058d5c1486dcda47d14448e3b53c8e56f

Observation b19eaa3e-d997-4475-af1e-df2e6a124e0c · inbound

Taming LLMs by Scaling Learning Rates with Gradient Grouping cites this paper.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Deconstructing What Makes a Good Optimizer for Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.263146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.263146Z digest=sha256:210e386bf8eb5285117528637bab5eaf03d49deb5d577deb7848b993a062992c

Observation d6d36155-b1fa-4519-8f90-785c1d93fa03 · inbound

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling cites this paper.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Deconstructing What Makes a Good Optimizer for Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.636520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.636520Z digest=sha256:6533819269301d5e81d75ef28404e46d61d2d566afc80d0c917fdc46e1e76a4b

Observation d12a81c7-2229-430d-9787-4b99a6c15a39 · inbound

On Design Principles for Private Adaptive Optimizers cites this paper.

On Design Principles for Private Adaptive Optimizers Deconstructing What Makes a Good Optimizer for Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:09:33.200101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:09:33.200101Z digest=sha256:7304d928074d1b33ee81d6959eba392753c49ebe60cb700f167afb46a1efc920

Observation 7ea7beba-0650-4b10-aaf8-17de7a19d88a · inbound

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling cites this paper.

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling Deconstructing What Makes a Good Optimizer for Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T09:38:51.648559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:38:51.648559Z digest=sha256:1300b86cbffa41089bfab76cd5f56fe76f0d4f3e24207978e7bde6f202c1f704

Observation e2572600-d266-46aa-9b73-c18d7d38e2fe · inbound

Prototype Transformer: Towards Language Model Architectures Interpretable by Design cites this paper.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Deconstructing What Makes a Good Optimizer for Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:46.444321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:46.444321Z digest=sha256:b1bbc55fec22d8bba298d9eede03d74ff31025fea9aa4239b7b44c5f06de918e

Observation d38d8f07-0ca0-46f3-863d-c89996a5d9e8 · inbound

Why Muon Outperforms Adam: A Curvature Perspective cites this paper.

Why Muon Outperforms Adam: A Curvature Perspective Deconstructing What Makes a Good Optimizer for Language Models

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:06:44.884155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T07:04:21.012269Z digest=sha256:d17c02159fc7e6e65247ca0d208f6acf47a959815f8ab5aa6fe9b142b6514f11

Observation d6158db1-28c9-472b-918e-6470b164ddee · inbound

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss cites this paper.

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss Deconstructing What Makes a Good Optimizer for Language Models

Reference 283

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:56:56.098113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T02:35:39.845487Z digest=sha256:bb9a36f7b3792558b067e7a4820a9d24d6b5c305a3b709872808250c64a24b05

Observation ac0f7cd7-60a2-45d1-a928-e2242878a79c · inbound

Muon Learns More Robust and Transferable Features than Adam cites this paper.

Muon Learns More Robust and Transferable Features than Adam Deconstructing What Makes a Good Optimizer for Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:30.164103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T17:08:30.717799Z digest=sha256:4a14b29210611231212d85ec9d58c9b2358d21be467d553420eba2603f1f22ca

Observation c531ab18-8baf-4ef5-a237-ab573d950f87 · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Deconstructing What Makes a Good Optimizer for Language Models

Reference 139

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:cc198677c8e5a8f29096db2b8e7e3b2a9e322724f0c0dc2d580305d6b2ef64d5

Observation 4e96dacf-3e2f-44a5-9d04-d6acdee852fa · inbound

Muon Meets Mamba: Spectral Optimization for State Space Models cites this paper.

Muon Meets Mamba: Spectral Optimization for State Space Models Deconstructing What Makes a Good Optimizer for Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T05:27:16.990675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:27:16.990675Z digest=sha256:4959f9d78700cecc7ea3c9888f93cae42dd4026eed62d06641138fab126a6864