Pith. sign in

Paper Citation Record · LEDGER

Pre-Training LLMs on a budget: A comparison of three optimizers

As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 3 inbound Pith citation observations for arXiv:2507.08472.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08472 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:22:23.931757Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T11:19:35.239826Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T12:15:22.206452Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7a2ea393-f05a-43d5-a9c0-434ecaa518c0 · outbound

This paper cites sign SGD with majority vote is communication efficient and fault tolerant.

Pre-Training LLMs on a budget: A comparison of three optimizers sign SGD with majority vote is communication efficient and fault tolerant

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:26.282195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.483030Z digest=sha256:d06b3fa937c0e56aeef01de04d709679dff601aa3419fd99b99e4578a6abf550

Observation 4fb7e08d-9d50-4d3c-869b-4e0a89bbbd10 · outbound

This paper cites Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling.

Pre-Training LLMs on a budget: A comparison of three optimizers Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.496384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.496384Z digest=sha256:49c7d8b28494479b734bc7486cc3e1dc0206a24abd28596cb95d9e5ac6ddc225

Observation cd89f21a-c756-47b7-b3a5-5b92a4e5978f · outbound

This paper cites Language models are few-shot learners.

Pre-Training LLMs on a budget: A comparison of three optimizers Language models are few-shot learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.506585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.506585Z digest=sha256:a64310b815d8423cb3547a88d298a9e174ef53ab080dadba00c1cbb5be7a9c8e

Observation ab00a737-6491-4258-8b39-6a8453ca908d · outbound

This paper cites Symbolic discovery of optimization algorithms.

Pre-Training LLMs on a budget: A comparison of three optimizers Symbolic discovery of optimization algorithms

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:26.188949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.516425Z digest=sha256:42d945a0304aca64d7026d952c3505af6aeb29c83a4a334cade2fd2405dde17b

Observation b138bef4-ea8b-4b2e-b0ba-148afe2a5154 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Pre-Training LLMs on a budget: A comparison of three optimizers Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.526060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.526060Z digest=sha256:4de7a8bb281347486f77e5e78c9115b0d14176e5c042245180a62b5fa10a9120

Observation 318d535f-1e5d-4bdd-9b46-383f247151c7 · outbound

This paper cites BTLM-3B-8K: 7B Parameter Performance in a 3B Parameter Model.

Pre-Training LLMs on a budget: A comparison of three optimizers BTLM-3B-8K: 7B Parameter Performance in a 3B Parameter Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.535485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.535485Z digest=sha256:e8c47a1ee276f9e256f8f9ca378a08b4430126a9b34cb7e3cf2e79aa32b8aa83

Observation 589650ec-6d73-4964-9628-2395376204ee · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Pre-Training LLMs on a budget: A comparison of three optimizers A framework for few-shot language model evaluation, 07 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.545685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.545685Z digest=sha256:198edbf02c0666c48f003c533da84a2e70271d2047911759ea67d13921483dfb

Observation 2f589883-c4bb-4b03-a752-766fe9ab3b97 · outbound

This paper cites OLMES: A Standard for Language Model Evaluations.

Pre-Training LLMs on a budget: A comparison of three optimizers OLMES: A Standard for Language Model Evaluations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.555619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.555619Z digest=sha256:f8f73a104883592e7fd28650d8edb2569798fbeaa907974c2de016ad74c0e069

Observation dab9aca4-5267-47b8-9945-a66bd0f6c6c1 · outbound

This paper cites Measuring massive multitask language understanding.

Pre-Training LLMs on a budget: A comparison of three optimizers Measuring massive multitask language understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.580428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.580428Z digest=sha256:e819d0d35acafcf042005fb17e5380babed78c9f45b35eaed5834ce2db79a73b

Observation 718cde59-c1d3-4d8c-9802-b0ba220e4bd2 · outbound

This paper cites Rae, and Laurent Sifre.

Pre-Training LLMs on a budget: A comparison of three optimizers Rae, and Laurent Sifre

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.590334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.590334Z digest=sha256:e4cedcc213f386a9cb0cf154a95375da9a336437fa9e7e4eda19c8a1118bca08

Observation 4b1c014c-1ea5-4000-ab3c-92f0fae0f339 · outbound

This paper cites On the Parameterization of Second-Order Optimization Effective Towards the Infinite Width.

Pre-Training LLMs on a budget: A comparison of three optimizers On the Parameterization of Second-Order Optimization Effective Towards the Infinite Width

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.602995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.602995Z digest=sha256:b5fa3350b588d66f16473de2324cd4bf43122b75331bd6e88f12b07e1a98ccfc

Observation 97449433-a998-45a4-86f9-76aa52746a28 · outbound

This paper cites No train no gain: Revisiting efficient training algorithms for transformer-based language models.

Pre-Training LLMs on a budget: A comparison of three optimizers No train no gain: Revisiting efficient training algorithms for transformer-based language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:26.059920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.712776Z digest=sha256:9f231789e24c801d10d3899d951d87a4ac240bbc79b766aa906a3b14a6ed6003

Observation 297ef640-4e68-4fd3-87d3-48d6b58f84c6 · outbound

This paper cites A method for stochastic optimization.

Pre-Training LLMs on a budget: A comparison of three optimizers A method for stochastic optimization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.918161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.768846Z digest=sha256:9c1e4f1c97c6c90be7ea20364a5b209015ba39f210d0227760aa69b17a21de6c

Observation c50b93d0-4f3b-4381-92cc-9206bc9fab8d · outbound

This paper cites Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks.

Pre-Training LLMs on a budget: A comparison of three optimizers Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.776470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.835815Z digest=sha256:58db5210fe0e3212e2b7703fa25bf580a119fd5225e9de82abd0a9bbf1775a5a

Observation 405688fe-4eae-4972-8b32-0ae8ffec189b · outbound

This paper cites ROPE : Reading order equivariant positional encoding for graph-based document information extraction.

Pre-Training LLMs on a budget: A comparison of three optimizers ROPE : Reading order equivariant positional encoding for graph-based document information extraction

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.616982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.935665Z digest=sha256:bd26f4febc42d033207cc9c152abe657c8b1b985085ced32869fce0166a316b4

Observation f702a233-b0e3-4ee1-9389-32653e13a488 · outbound

This paper cites an unresolved cited work.

Pre-Training LLMs on a budget: A comparison of three optimizers Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:22:25.497357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:22.013963Z digest=sha256:7242695a44bff3a75bcb0f3a23869329cd69693b2ddf1d54e804c6eb9fb7c439

Observation 96cc9c6d-e052-43a9-a3e7-940e0e07b170 · outbound

This paper cites An Empirical Study of $\mu$P Learning Rate Transfer.

Pre-Training LLMs on a budget: A comparison of three optimizers An Empirical Study of $\mu$P Learning Rate Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.099491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.099491Z digest=sha256:67fc441f0fb5aa6ec8aed8a0aaf178ad490da7707a351b854182e1cb967df6a3

Observation ec7119fa-76ab-4581-be0e-0493269ea6ca · outbound

This paper cites Sophia: A scalable stochastic second-order optimizer for language model pre-training.

Pre-Training LLMs on a budget: A comparison of three optimizers Sophia: A scalable stochastic second-order optimizer for language model pre-training

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.349057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:22.233634Z digest=sha256:d6a9b06cadf37eae1e9acd88ad4e9b2bdf5d05b90dc61bbd9701c954ab67a732

Observation 1c3f85f4-b2ee-4d14-a27e-94a65b1cbd47 · outbound

This paper cites Decoupled Weight Decay Regularization.

Pre-Training LLMs on a budget: A comparison of three optimizers Decoupled Weight Decay Regularization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.322826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.322826Z digest=sha256:739c2f8606ba3278a0005ce19392ff343aa9410a0ab9c8c63fe706adde205aa3

Observation 616636a0-c779-4e6e-be13-096e23376066 · outbound

This paper cites Scaling data-constrained language models.

Pre-Training LLMs on a budget: A comparison of three optimizers Scaling data-constrained language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.396965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.396965Z digest=sha256:7bd60f78f212765f6ff75c4476011b2eb6dd06755a9ad32d7d08559833de6e0f

Observation cdd2d457-dc2a-4937-8cad-2bd12a46ee15 · outbound

This paper cites Language models are unsupervised multitask learners.

Pre-Training LLMs on a budget: A comparison of three optimizers Language models are unsupervised multitask learners

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.479151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.479151Z digest=sha256:dad974f1d950ec9984fe50124ceb9b5efb01bb485d2eee5fc595fec486c783dd

Observation f3cfcba2-07bc-456a-a1dd-b56b60f81b68 · outbound

This paper cites A modified A dam algorithm for deep neural network optimization.

Pre-Training LLMs on a budget: A comparison of three optimizers A modified A dam algorithm for deep neural network optimization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.218513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:22.557255Z digest=sha256:1803a8640a1cec8bdf418511740229154ca88b4dc4ba1500ad0e4eb209d2cd88

Observation f782298b-0f15-41d1-8d96-eaaeea223815 · outbound

This paper cites An overview of gradient descent optimization algorithms.

Pre-Training LLMs on a budget: A comparison of three optimizers An overview of gradient descent optimization algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.619341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.619341Z digest=sha256:7fd12377d501c663477f5d4ff49ea8043589173e4124a98917ca4598a347376f

Observation 4d23e4df-9c67-47f1-b7d5-d39e12080c38 · outbound

This paper cites Adafactor: Adaptive learning rates with sublinear memory cost.

Pre-Training LLMs on a budget: A comparison of three optimizers Adafactor: Adaptive learning rates with sublinear memory cost

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.139815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:22.702979Z digest=sha256:ffd7bb5f9fb5f0d45354d5aa23e51b38dd51975246356c253faa67b578c80c41

Observation 1de602ec-89cb-49b7-bdc8-a716687951ec · outbound

This paper cites SlimPajama: A 627B token cleaned and deduplicated version of RedPajama.

Pre-Training LLMs on a budget: A comparison of three optimizers SlimPajama: A 627B token cleaned and deduplicated version of RedPajama

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.820809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.820809Z digest=sha256:fbf93320c11cf4faf2be8d7e731bcc19934db4577f749095e4b815290e3129dc

Observation c1df782e-021a-4b02-bf3f-58d33eddde44 · outbound

This paper cites Spike no more: Stabilizing the pre-training of large language models, 2025.

Pre-Training LLMs on a budget: A comparison of three optimizers Spike no more: Stabilizing the pre-training of large language models, 2025

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.996150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:22.890439Z digest=sha256:1d50e3c60c172467b0d9bea5c900f301ba700baa9fef6c8894661ad5fa60921b

Observation f3481f1d-0cd6-4369-a2bf-aa8047fea6d1 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Pre-Training LLMs on a budget: A comparison of three optimizers LLaMA: Open and Efficient Foundation Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.020109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.020109Z digest=sha256:157b7efa028919a73a2ea71206ea9d6b765431790ed6d3eaa8d4e7b28bc8982e

Observation b9ed3627-872a-4194-bb55-ecae015a44b2 · outbound

This paper cites Evolution and role of optimizers in training deep learning models, 2024.

Pre-Training LLMs on a budget: A comparison of three optimizers Evolution and role of optimizers in training deep learning models, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.868709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.098112Z digest=sha256:ab7801a6c4990a5ee5ab1d3dbfe4f077e673c51d315e9abfcfda5a836540ab53

Observation 9e4f294a-5d8d-4cab-a6d6-faeb9a94718c · outbound

This paper cites Ranger21: a synergistic deep learning optimizer.

Pre-Training LLMs on a budget: A comparison of three optimizers Ranger21: a synergistic deep learning optimizer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.166969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.166969Z digest=sha256:c6319bed823de4f9c175ac9a5fec580d48ecbeb3257de1648e73ec24bd88aa29

Observation 6f250c12-6b80-4fca-9c58-2be7d5e6a03e · outbound

This paper cites Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models.

Pre-Training LLMs on a budget: A comparison of three optimizers Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.767914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.243728Z digest=sha256:4ad2186c250965fa5c393f2a2f90c0936eb51548388a4a5c53a3ff040a80ac0a

Observation c83071cc-073c-4e99-b631-ab1a1292e651 · outbound

This paper cites Tuning large neural networks via zero-shot hyperparameter transfer.

Pre-Training LLMs on a budget: A comparison of three optimizers Tuning large neural networks via zero-shot hyperparameter transfer

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.654484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.319784Z digest=sha256:50d10817d2482f6b984a9ee9ac498add4fbec2b032b453280c9680d7ae09778a

Observation 38cdee5c-1ba7-41cd-a7cb-2e6166473774 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.

Pre-Training LLMs on a budget: A comparison of three optimizers Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.370940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.370940Z digest=sha256:96bc6f8260bfdad95bf80526b6778960f582894ec41fdd03549662a5be6d170c

Observation 71c1903a-373d-4d1c-b18a-7bf22963b90f · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

Pre-Training LLMs on a budget: A comparison of three optimizers Adam-mini: Use Fewer Learning Rates To Gain More

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.426489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.426489Z digest=sha256:f113a893a3e8eb700b9b2b2feded14d4026554fb1099cfc7c7f68f025fb4dcdb

Observation e7156432-f7cf-4697-b702-72cfd2c05fc4 · outbound

This paper cites Improved adam optimizer for deep neural networks.

Pre-Training LLMs on a budget: A comparison of three optimizers Improved adam optimizer for deep neural networks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.506253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.522210Z digest=sha256:3591fc7d1d6d61b026e2dae683f510a14064a4775446f9747f5359808b725edf

Observation 6f168c4e-d217-4171-9806-b2a140876725 · outbound

This paper cites an unresolved cited work.

Pre-Training LLMs on a budget: A comparison of three optimizers Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:22:24.359342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.554627Z digest=sha256:99b8aa2a51de2f7ddf3e1338e6bc1ec66652e863c73a90f890117764cdde9dbe

Observation 6b5ec9fb-1624-45c9-a126-83da3fa1f79f · outbound

This paper cites Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.

Pre-Training LLMs on a budget: A comparison of three optimizers Adabelief optimizer: Adapting stepsizes by the belief in observed gradients

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.258442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.592781Z digest=sha256:bad05d475dbdb0652aaf71f17fdc1a454449bdc16ceeb41bf4bdf6aca58a8aa8

Observation fe6525a8-af09-40f1-99a7-af0421227267 · outbound

This paper cites write newline.

Pre-Training LLMs on a budget: A comparison of three optimizers write newline

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.691536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.691536Z digest=sha256:8d713113c55b1c26b3bcb4d2ff38b03fedeb70458c9b5289ace7de7062aba8c6

Observation 6062aa24-20ae-455b-85c0-f3a9a36eddc4 · outbound

This paper cites @esa (Ref.

Pre-Training LLMs on a budget: A comparison of three optimizers @esa (Ref

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.785237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.785237Z digest=sha256:3bf2bb15e78876ce8895dd8c0a50d8b53fc1d96863be3a9669b09995ef0dcc83

Observation 6b64a7ad-d192-4238-aff7-d1f3aaf472e4 · outbound

This paper cites an unresolved cited work.

Pre-Training LLMs on a budget: A comparison of three optimizers Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.885318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.885318Z digest=sha256:fc4fcdd8ab80a9630517c6ef26001dbb7127073c99bf701da67a5989e96aa37e

Observation b8b538c4-8e24-424d-9cfa-7ef785a12a42 · outbound

This paper cites an unresolved cited work.

Pre-Training LLMs on a budget: A comparison of three optimizers Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.931757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.931757Z digest=sha256:6f1e71a946717c75db91edc61538458f7103ce8aa5c4d77c2642bc7f110a108b

Pith citing papers

Observation 89f747dd-bff9-48d0-b28f-70f65461bca6 · inbound

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models cites this paper.

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models Pre-Training LLMs on a budget: A comparison of three optimizers

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:15:22.208322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T12:10:44.802059Z digest=sha256:b4cba25ea31c3670d58d790187b17c12af4cd9b9110751e3bf58e4327da8c40c

Observation 2391eeff-d65c-459e-99ec-1f5d71f89933 · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Pre-Training LLMs on a budget: A comparison of three optimizers

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:3822fc9e9e5daa2b3d6be01ce736464532b6cf6992b6820aae43d38d83e8c1ff

Observation 6feb931c-ffda-4945-bc51-44caea724920 · inbound

From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference cites this paper.

From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference Pre-Training LLMs on a budget: A comparison of three optimizers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T11:19:35.239826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T11:19:35.239826Z digest=sha256:9eee288fb6452bb7d7fc8dbc53b6479528f0a3ce8019abaa60e7605e758f96ce