Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:22:23.931757Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 3 inbound Pith citation observations for arXiv:2507.08472.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:22:23.931757Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T11:19:35.239826Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-10T12:15:22.206452Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7a2ea393-f05a-43d5-a9c0-434ecaa518c0 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers sign SGD with majority vote is communication efficient and fault tolerant
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4fb7e08d-9d50-4d3c-869b-4e0a89bbbd10 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd89f21a-c756-47b7-b3a5-5b92a4e5978f · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Language models are few-shot learners
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab00a737-6491-4258-8b39-6a8453ca908d · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Symbolic discovery of optimization algorithms
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b138bef4-ea8b-4b2e-b0ba-148afe2a5154 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 318d535f-1e5d-4bdd-9b46-383f247151c7 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers BTLM-3B-8K: 7B Parameter Performance in a 3B Parameter Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 589650ec-6d73-4964-9628-2395376204ee · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers A framework for few-shot language model evaluation, 07 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f589883-c4bb-4b03-a752-766fe9ab3b97 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers OLMES: A Standard for Language Model Evaluations
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dab9aca4-5267-47b8-9945-a66bd0f6c6c1 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Measuring massive multitask language understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 718cde59-c1d3-4d8c-9802-b0ba220e4bd2 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Rae, and Laurent Sifre
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b1c014c-1ea5-4000-ab3c-92f0fae0f339 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers On the Parameterization of Second-Order Optimization Effective Towards the Infinite Width
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97449433-a998-45a4-86f9-76aa52746a28 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers No train no gain: Revisiting efficient training algorithms for transformer-based language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 297ef640-4e68-4fd3-87d3-48d6b58f84c6 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers A method for stochastic optimization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c50b93d0-4f3b-4381-92cc-9206bc9fab8d · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 405688fe-4eae-4972-8b32-0ae8ffec189b · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers ROPE : Reading order equivariant positional encoding for graph-based document information extraction
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f702a233-b0e3-4ee1-9389-32653e13a488 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96cc9c6d-e052-43a9-a3e7-940e0e07b170 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers An Empirical Study of $\mu$P Learning Rate Transfer
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec7119fa-76ab-4581-be0e-0493269ea6ca · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Sophia: A scalable stochastic second-order optimizer for language model pre-training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c3f85f4-b2ee-4d14-a27e-94a65b1cbd47 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Decoupled Weight Decay Regularization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 616636a0-c779-4e6e-be13-096e23376066 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Scaling data-constrained language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdd2d457-dc2a-4937-8cad-2bd12a46ee15 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Language models are unsupervised multitask learners
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3cfcba2-07bc-456a-a1dd-b56b60f81b68 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers A modified A dam algorithm for deep neural network optimization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f782298b-0f15-41d1-8d96-eaaeea223815 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers An overview of gradient descent optimization algorithms
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d23e4df-9c67-47f1-b7d5-d39e12080c38 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Adafactor: Adaptive learning rates with sublinear memory cost
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1de602ec-89cb-49b7-bdc8-a716687951ec · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers SlimPajama: A 627B token cleaned and deduplicated version of RedPajama
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1df782e-021a-4b02-bf3f-58d33eddde44 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Spike no more: Stabilizing the pre-training of large language models, 2025
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f3481f1d-0cd6-4369-a2bf-aa8047fea6d1 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers LLaMA: Open and Efficient Foundation Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9ed3627-872a-4194-bb55-ecae015a44b2 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Evolution and role of optimizers in training deep learning models, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e4f294a-5d8d-4cab-a6d6-faeb9a94718c · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Ranger21: a synergistic deep learning optimizer
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f250c12-6b80-4fca-9c58-2be7d5e6a03e · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c83071cc-073c-4e99-b631-ab1a1292e651 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Tuning large neural networks via zero-shot hyperparameter transfer
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 38cdee5c-1ba7-41cd-a7cb-2e6166473774 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71c1903a-373d-4d1c-b18a-7bf22963b90f · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Adam-mini: Use Fewer Learning Rates To Gain More
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7156432-f7cf-4697-b702-72cfd2c05fc4 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Improved adam optimizer for deep neural networks
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6f168c4e-d217-4171-9806-b2a140876725 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b5ec9fb-1624-45c9-a126-83da3fa1f79f · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Adabelief optimizer: Adapting stepsizes by the belief in observed gradients
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe6525a8-af09-40f1-99a7-af0421227267 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers write newline
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6062aa24-20ae-455b-85c0-f3a9a36eddc4 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers @esa (Ref
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b64a7ad-d192-4238-aff7-d1f3aaf472e4 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8b538c4-8e24-424d-9cfa-7ef785a12a42 · outbound
Pre-Training LLMs on a budget: A comparison of three optimizers Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89f747dd-bff9-48d0-b28f-70f65461bca6 · inbound
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models Pre-Training LLMs on a budget: A comparison of three optimizers
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2391eeff-d65c-459e-99ec-1f5d71f89933 · inbound
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Pre-Training LLMs on a budget: A comparison of three optimizers
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6feb931c-ffda-4945-bc51-44caea724920 · inbound
From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference Pre-Training LLMs on a budget: A comparison of three optimizers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.