Pith. sign in

Paper Citation Record · LEDGER

Pre-Training LLMs on a budget: A comparison of three optimizers

As of 11 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 3 inbound Pith citation observations for arXiv:2507.08472.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08472 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:22:23.931757Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T11:19:35.239826Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T12:15:22.206452Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7a2ea393-f05a-43d5-a9c0-434ecaa518c0 · outbound

This paper cites sign SGD with majority vote is communication efficient and fault tolerant.

Pre-Training LLMs on a budget: A comparison of three optimizers sign SGD with majority vote is communication efficient and fault tolerant

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:26.282195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.483030Z digest=sha256:b8351d72c4835b9c2fcec921a92ce4ad2eed62054a097e8327a9706e4c897663

Observation 4fb7e08d-9d50-4d3c-869b-4e0a89bbbd10 · outbound

This paper cites Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling.

Pre-Training LLMs on a budget: A comparison of three optimizers Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.496384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.496384Z digest=sha256:34f2ad982d1bb3e0dcfcf4d24f06f8ff1690098fb3695d1550c7d713602dafbc

Observation cd89f21a-c756-47b7-b3a5-5b92a4e5978f · outbound

This paper cites Language models are few-shot learners.

Pre-Training LLMs on a budget: A comparison of three optimizers Language models are few-shot learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.506585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.506585Z digest=sha256:a06e4dea5fc393acb3096f21f3db2e3d20799bc7fcf01c1c2fef90bdb8928b60

Observation ab00a737-6491-4258-8b39-6a8453ca908d · outbound

This paper cites Symbolic discovery of optimization algorithms.

Pre-Training LLMs on a budget: A comparison of three optimizers Symbolic discovery of optimization algorithms

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:26.188949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.516425Z digest=sha256:4c901f5982b96ff3d171c1100a364cddeda03950d45a2eb6595a06e1cbb3086e

Observation b138bef4-ea8b-4b2e-b0ba-148afe2a5154 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Pre-Training LLMs on a budget: A comparison of three optimizers Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.526060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.526060Z digest=sha256:aa98c93c7a5490f95446d0ddd358d4112bb23949092d553e60553c78feeb508b

Observation 318d535f-1e5d-4bdd-9b46-383f247151c7 · outbound

This paper cites BTLM-3B-8K: 7B Parameter Performance in a 3B Parameter Model.

Pre-Training LLMs on a budget: A comparison of three optimizers BTLM-3B-8K: 7B Parameter Performance in a 3B Parameter Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.535485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.535485Z digest=sha256:c085e7ef6ff7d047de8758305063c5959646f47e943321263c7e4124190c8081

Observation 589650ec-6d73-4964-9628-2395376204ee · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Pre-Training LLMs on a budget: A comparison of three optimizers A framework for few-shot language model evaluation, 07 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.545685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.545685Z digest=sha256:b1b91d1855f4ae92a0d8a7e9557a6d892245524e1b701b19a5f6af0dc60e3f3e

Observation 2f589883-c4bb-4b03-a752-766fe9ab3b97 · outbound

This paper cites OLMES: A Standard for Language Model Evaluations.

Pre-Training LLMs on a budget: A comparison of three optimizers OLMES: A Standard for Language Model Evaluations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.555619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.555619Z digest=sha256:e880149d1a1484197e05e046c4b757e40634b206481d25d0a2c7ca0c9e614a83

Observation dab9aca4-5267-47b8-9945-a66bd0f6c6c1 · outbound

This paper cites Measuring massive multitask language understanding.

Pre-Training LLMs on a budget: A comparison of three optimizers Measuring massive multitask language understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.580428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.580428Z digest=sha256:b3eb7883e1b43b93ff8aef53d2bc7490d737c679bb2dc80eaac1463f7b395920

Observation 718cde59-c1d3-4d8c-9802-b0ba220e4bd2 · outbound

This paper cites Rae, and Laurent Sifre.

Pre-Training LLMs on a budget: A comparison of three optimizers Rae, and Laurent Sifre

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.590334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.590334Z digest=sha256:a074da0e9ebf0cd55331e74e1cf49dc2edadce007a7e363ccc9522b099918e7d

Observation 4b1c014c-1ea5-4000-ab3c-92f0fae0f339 · outbound

This paper cites On the Parameterization of Second-Order Optimization Effective Towards the Infinite Width.

Pre-Training LLMs on a budget: A comparison of three optimizers On the Parameterization of Second-Order Optimization Effective Towards the Infinite Width

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.602995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.602995Z digest=sha256:68f6fe4cca99cc38e07e7533f7f3ebcb0d6656e8bf6cf14abd7c61e4458c1c6c

Observation 97449433-a998-45a4-86f9-76aa52746a28 · outbound

This paper cites No train no gain: Revisiting efficient training algorithms for transformer-based language models.

Pre-Training LLMs on a budget: A comparison of three optimizers No train no gain: Revisiting efficient training algorithms for transformer-based language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:26.059920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.712776Z digest=sha256:12a0d73604ac1428b68814abdc2c37ee5b19e9955406bbcee76b874ed2a512a4

Observation 297ef640-4e68-4fd3-87d3-48d6b58f84c6 · outbound

This paper cites A method for stochastic optimization.

Pre-Training LLMs on a budget: A comparison of three optimizers A method for stochastic optimization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.918161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.768846Z digest=sha256:8322a5da0b9be3078625805e3ad9f7d50e4bfd11a4be4515d00d19b7ee7166cf

Observation c50b93d0-4f3b-4381-92cc-9206bc9fab8d · outbound

This paper cites Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks.

Pre-Training LLMs on a budget: A comparison of three optimizers Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.776470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.835815Z digest=sha256:69905be7d08a6200506c9a72b53db5afd4782861875b6536cbe85f2f4eb1660b

Observation 405688fe-4eae-4972-8b32-0ae8ffec189b · outbound

This paper cites ROPE : Reading order equivariant positional encoding for graph-based document information extraction.

Pre-Training LLMs on a budget: A comparison of three optimizers ROPE : Reading order equivariant positional encoding for graph-based document information extraction

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.616982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.935665Z digest=sha256:a33dee1d2e7bae4c075be30455a895a6479ea87877ce3ae1e4c8932b9bbf5e12

Observation f702a233-b0e3-4ee1-9389-32653e13a488 · outbound

This paper cites an unresolved cited work.

Pre-Training LLMs on a budget: A comparison of three optimizers Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:22:25.497357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:22.013963Z digest=sha256:1050a489a5eb9f674203c0e8783e86a67a117b2258fd90d5c537196b976e0da1

Observation 96cc9c6d-e052-43a9-a3e7-940e0e07b170 · outbound

This paper cites An Empirical Study of $\mu$P Learning Rate Transfer.

Pre-Training LLMs on a budget: A comparison of three optimizers An Empirical Study of $\mu$P Learning Rate Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.099491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.099491Z digest=sha256:f4c91496b2c0b4bdd3008bc7e93d4bd91f86dc2773c774781e6d96acc01c5722

Observation ec7119fa-76ab-4581-be0e-0493269ea6ca · outbound

This paper cites Sophia: A scalable stochastic second-order optimizer for language model pre-training.

Pre-Training LLMs on a budget: A comparison of three optimizers Sophia: A scalable stochastic second-order optimizer for language model pre-training

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.349057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:22.233634Z digest=sha256:e35d9c99389eff98b0297bb9008cb26c2bda331ad14a0d46e8ad65e0f4fc2cbd

Observation 1c3f85f4-b2ee-4d14-a27e-94a65b1cbd47 · outbound

This paper cites Decoupled Weight Decay Regularization.

Pre-Training LLMs on a budget: A comparison of three optimizers Decoupled Weight Decay Regularization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.322826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.322826Z digest=sha256:0e840f2bd777f0883ccac0bea04cc427a3e3b79cc1bfebded623c7fdaade04a4

Observation 616636a0-c779-4e6e-be13-096e23376066 · outbound

This paper cites Scaling data-constrained language models.

Pre-Training LLMs on a budget: A comparison of three optimizers Scaling data-constrained language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.396965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.396965Z digest=sha256:02a3a3545d7dd589b85b54940b60d2f087df40954d4b580b4da12c10adb777a2

Observation cdd2d457-dc2a-4937-8cad-2bd12a46ee15 · outbound

This paper cites Language models are unsupervised multitask learners.

Pre-Training LLMs on a budget: A comparison of three optimizers Language models are unsupervised multitask learners

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.479151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.479151Z digest=sha256:ef17de46c0e4076fa5521851700f8dc062cb2aa2647cf110a698035bbf05d4e7

Observation f3cfcba2-07bc-456a-a1dd-b56b60f81b68 · outbound

This paper cites A modified A dam algorithm for deep neural network optimization.

Pre-Training LLMs on a budget: A comparison of three optimizers A modified A dam algorithm for deep neural network optimization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.218513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:22.557255Z digest=sha256:03f338995215ef47b0e4ddca9efe0f6df703a399f25e025c3f0b763ccb8a3b5d

Observation f782298b-0f15-41d1-8d96-eaaeea223815 · outbound

This paper cites An overview of gradient descent optimization algorithms.

Pre-Training LLMs on a budget: A comparison of three optimizers An overview of gradient descent optimization algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.619341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.619341Z digest=sha256:9125098688cfea8b75951ad9d6455635fe9243fbdcc0cbe8d5c6cf2f1343bb6a

Observation 4d23e4df-9c67-47f1-b7d5-d39e12080c38 · outbound

This paper cites Adafactor: Adaptive learning rates with sublinear memory cost.

Pre-Training LLMs on a budget: A comparison of three optimizers Adafactor: Adaptive learning rates with sublinear memory cost

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.139815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:22.702979Z digest=sha256:1750fe88e7d6acb986f471e7390dfe4d3b53c3f72514c4c1670d2ef87d9fb7bb

Observation 1de602ec-89cb-49b7-bdc8-a716687951ec · outbound

This paper cites SlimPajama: A 627B token cleaned and deduplicated version of RedPajama.

Pre-Training LLMs on a budget: A comparison of three optimizers SlimPajama: A 627B token cleaned and deduplicated version of RedPajama

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.820809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.820809Z digest=sha256:008bbf8fc6b204066065d30e23ff688a86c6376b1525a8e869fae36462864802

Observation c1df782e-021a-4b02-bf3f-58d33eddde44 · outbound

This paper cites Spike no more: Stabilizing the pre-training of large language models, 2025.

Pre-Training LLMs on a budget: A comparison of three optimizers Spike no more: Stabilizing the pre-training of large language models, 2025

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.996150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:22.890439Z digest=sha256:b66c7cebc3ba0d9a79c772989c878d5d9f17e7af342a4864b24418f4ae602bed

Observation f3481f1d-0cd6-4369-a2bf-aa8047fea6d1 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Pre-Training LLMs on a budget: A comparison of three optimizers LLaMA: Open and Efficient Foundation Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.020109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.020109Z digest=sha256:3dd695e450c01fcf1e33c60a767c41ce2111f99e925ce858977a6b34878bff27

Observation b9ed3627-872a-4194-bb55-ecae015a44b2 · outbound

This paper cites Evolution and role of optimizers in training deep learning models, 2024.

Pre-Training LLMs on a budget: A comparison of three optimizers Evolution and role of optimizers in training deep learning models, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.868709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.098112Z digest=sha256:560caf4cd72681f3c8ca56207668cb0798ed19f64f81ed06897167dbb4fc2120

Observation 9e4f294a-5d8d-4cab-a6d6-faeb9a94718c · outbound

This paper cites Ranger21: a synergistic deep learning optimizer.

Pre-Training LLMs on a budget: A comparison of three optimizers Ranger21: a synergistic deep learning optimizer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.166969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.166969Z digest=sha256:d35f327a5362211384fc554f648ea01a87138ce0efe4c2248445fe18d11e3c65

Observation 6f250c12-6b80-4fca-9c58-2be7d5e6a03e · outbound

This paper cites Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models.

Pre-Training LLMs on a budget: A comparison of three optimizers Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.767914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.243728Z digest=sha256:26cdf73e6fe5299178fdc7f4cfe9a1c3bce46b32b844920cf3ce2cf41dc23d5f

Observation c83071cc-073c-4e99-b631-ab1a1292e651 · outbound

This paper cites Tuning large neural networks via zero-shot hyperparameter transfer.

Pre-Training LLMs on a budget: A comparison of three optimizers Tuning large neural networks via zero-shot hyperparameter transfer

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.654484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.319784Z digest=sha256:7ef3befcb611c292620e508c73d057f6c400383311fe359e58753a7d08a12e8c

Observation 38cdee5c-1ba7-41cd-a7cb-2e6166473774 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.

Pre-Training LLMs on a budget: A comparison of three optimizers Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.370940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.370940Z digest=sha256:9b51fd70920e85096b7ab3e7b6a767ea5af58a40f78f9d763b737fb964e40db3

Observation 71c1903a-373d-4d1c-b18a-7bf22963b90f · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

Pre-Training LLMs on a budget: A comparison of three optimizers Adam-mini: Use Fewer Learning Rates To Gain More

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.426489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.426489Z digest=sha256:56a47bc9ef6304c25826bdc221c674e495bc305a5829435f650a459a31880a01

Observation e7156432-f7cf-4697-b702-72cfd2c05fc4 · outbound

This paper cites Improved adam optimizer for deep neural networks.

Pre-Training LLMs on a budget: A comparison of three optimizers Improved adam optimizer for deep neural networks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.506253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.522210Z digest=sha256:ff65151cd7e46af70b586027f85441a4a50c1ab4b05bdaae18727e8e52831a70

Observation 6f168c4e-d217-4171-9806-b2a140876725 · outbound

This paper cites an unresolved cited work.

Pre-Training LLMs on a budget: A comparison of three optimizers Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:22:24.359342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.554627Z digest=sha256:3686654dc053eb2435c67a91d7cd4e1a1e60635709c051b6d50295fc239c2159

Observation 6b5ec9fb-1624-45c9-a126-83da3fa1f79f · outbound

This paper cites Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.

Pre-Training LLMs on a budget: A comparison of three optimizers Adabelief optimizer: Adapting stepsizes by the belief in observed gradients

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.258442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.592781Z digest=sha256:dbf88733862646fc2b7274c2594232a745167a6c7738bf9b2383b4e57c230b74

Observation fe6525a8-af09-40f1-99a7-af0421227267 · outbound

This paper cites write newline.

Pre-Training LLMs on a budget: A comparison of three optimizers write newline

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.691536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.691536Z digest=sha256:d7fb49634ac5c4abf0f3ed80979448aeecab0fbe893ef5cfec34e8a7eb953c70

Observation 6062aa24-20ae-455b-85c0-f3a9a36eddc4 · outbound

This paper cites @esa (Ref.

Pre-Training LLMs on a budget: A comparison of three optimizers @esa (Ref

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.785237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.785237Z digest=sha256:800e31b0c9a2f4f114dee9d7201323a4a9a840af66c7c67df69d2e0f0039e9de

Observation 6b64a7ad-d192-4238-aff7-d1f3aaf472e4 · outbound

This paper cites an unresolved cited work.

Pre-Training LLMs on a budget: A comparison of three optimizers Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.885318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.885318Z digest=sha256:151c602aa6934208442a67a29c33445031a9f96d2023c1380e68f8ca50fa84b6

Observation b8b538c4-8e24-424d-9cfa-7ef785a12a42 · outbound

This paper cites an unresolved cited work.

Pre-Training LLMs on a budget: A comparison of three optimizers Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.931757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.931757Z digest=sha256:f120de3928ded411c3deba2aa3fa1256bdb4e066ef1e4ed46f512ea59648e6c3

Pith citing papers

Observation 89f747dd-bff9-48d0-b28f-70f65461bca6 · inbound

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models cites this paper.

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models Pre-Training LLMs on a budget: A comparison of three optimizers

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:15:22.208322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T12:10:44.802059Z digest=sha256:e04883ea06e139ccfe4e234a341ce4f28a36a5f8ad14be368f3550e2a978aac3

Observation 2391eeff-d65c-459e-99ec-1f5d71f89933 · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Pre-Training LLMs on a budget: A comparison of three optimizers

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:f381f6a34752d14f1e0796c075c05fd9a818b719dbb0272766129e4315a32c34

Observation 6feb931c-ffda-4945-bc51-44caea724920 · inbound

From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference cites this paper.

From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference Pre-Training LLMs on a budget: A comparison of three optimizers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T11:19:35.239826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T11:19:35.239826Z digest=sha256:218f22360ce86e09965ed2e5985795355cdf4f8300094a9a481acd638bf86d0e