Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:30.649181Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 6 inbound Pith citation observations for arXiv:2506.12543.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:30.649181Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T09:38:48.446865Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
30 of 30 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation a21bd626-5831-4ece-acb5-d4840f36f7b9 · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Disentangling Adaptive Gradient Methods from Learning Rates
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49d4b4c9-e468-4c4f-b591-c01e47235889 · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling As before, the gap decreases the longer we train, and SGD can eventually outperform Adam
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42de01fd-13d3-4930-a9bc-9a2017e833c3 · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bbf646a2-602d-4105-bf75-187a47e1043a · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faba8271-5cc9-4dd9-a404-e1ec57e7b9b5 · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling How Does Adaptive Optimization Impact Local Neural Network Geometry?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ed2fc01-02dd-4f7e-831e-dbf28261e37e · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Adam: A Method for Stochastic Optimization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93b8ace5-b1e8-41e7-976f-b4ee1f3a29da · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f2a746-919f-4dfd-9dad-d458551d844e · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Muon is Scalable for LLM Training
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf808663-47b5-4c1a-9afc-7e27f355ba34 · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c264e59-9a2c-40dd-b301-bfb40bb9ea61 · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Orvieto and R
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c90daf0d-ba2c-4c9c-8d1c-cfba25a472cb · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Toward Understanding Why Adam Converges Faster Than SGD for Transformers
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69e7fee2-9ad5-4733-b8d8-8f7515dc23f4 · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling arXiv:2502.00213 [cs]
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff2174f1-3d25-4dac-8c73-a60ed35c049d · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Adam Exploits $\ell_\infty$-geometry of Loss Landscape via Coordinate-wise Adaptivity
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37a5bffc-f69f-4701-ab14-eb9cf9c9450d · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling How Does Critical Batch Size Scale in Pre-training?
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7dcaf01-f68f-4aad-a8d0-86ec4ad73dc9 · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Why Transformers Need Adam: A Hessian Perspective
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6d36155-b1fa-4519-8f90-785c1d93fa03 · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Deconstructing What Makes a Good Optimizer for Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a976dd82-e9e3-4506-b999-0f76593a6436 · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08b0145f-91e0-48b2-ae3e-3a5593e047d9 · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling [2024], and uses the codebase of Orvieto and Gower [2025]
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f34f1c0d-9218-4ca8-aec6-2a8131c3d5ec · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Unresolved cited work
Reference 1024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f176b608-2526-4295-95c2-1538cb7575ee · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Transformers without Tears: Improving the Normalization of Self-Attention
Reference 1986
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6799d912-962d-42ed-bda7-de130b44a4a0 · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling How to Fine-Tune Vision Models with SGD
Reference 2014
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c597f601-3722-465a-851c-2e2ab7b14dcb · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling The Llama 3 Herd of Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4d8fd51-e618-4b44-9367-f940273b414c · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling signSGD: Compressed Optimisation for Non-Convex Problems
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef191c64-b70d-4e05-a88f-27d993caa9c9 · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Practical Efficiency of Muon for Pretraining
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f04113b-7173-4087-a8fc-853e94d5a18b · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72b5ec18-fdd9-4bf8-b615-731d1aa0231f · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling SGDR: Stochastic Gradient Descent with Warm Restarts
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a959df74-d31d-472b-b5f2-150c36d047b8 · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cec26bc-75cf-40f6-8c17-d216703fe868 · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c56be0f-8610-4504-8510-4e9111b6f761 · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Linear attention is (maybe) all you need (to understand transformer optimization)
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0da4f519-d61b-47a1-8033-3430a73bbbff · outbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling GLU Variants Improve Transformer
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34bb4621-ccf2-4792-bf5c-1e3ccfebbc71 · inbound
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d414ffe-531d-4c19-a57a-892ac77a246f · inbound
How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01bd3cbc-d651-4a8c-bf5a-ad2e192dce3d · inbound
Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53449721-9569-4766-a9a1-050f989768bf · inbound
Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
Reference 235
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3aa9feae-5033-46d6-8a2e-f130184609b1 · inbound
A Theoretical and Experimental Study of a Novel Adaptive Learning Algorithm Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9e7170e-6db5-4851-b0e2-92c553ad4910 · inbound
Quantifying the Agreement Between Data-Influence and Data-Similarity to Understand LLM Behavior Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.