Pith. sign in

Paper Citation Record · LEDGER

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning

As of 21 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2506.05447.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05447 v2

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:31:20.550988Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:02:05.516512Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy2
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a381240c-4449-458e-9476-d0579b6f0a85 · outbound

This paper cites online" 'onlinestring :=.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.306883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.306883Z digest=sha256:73e8afb71f70e35df86d402966ead1f97c2a9e088bbf8b37684546bd74afdd28

Observation 673f7902-411a-4e2f-bc9f-3480e903d649 · outbound

This paper cites write newline.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.314438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.314438Z digest=sha256:a6d6bca109050a0d88177bdb4da798dc5fc405b1495d0bb00a92b1e214b2f4a1

Observation 187ee80c-5b8b-4d2e-bde7-363c5159e6eb · outbound

This paper cites Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.328935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.328935Z digest=sha256:7216c3ca61b93150e1105abcce5fc03fbeb80a5deb313713895069c38ee7f220

Observation 94b2fce2-c6d0-47c3-86a8-d3e2020a40b2 · outbound

This paper cites Atanasov, Jacob A.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Atanasov, Jacob A

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:31:21.353615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.335963Z digest=sha256:701ebe90a3ae478357f536be7e5b94e3de097cbff51a0b9dafe2b9778d7a76f1

Observation 28757c54-f468-4906-b66d-be7fa4b406f8 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.341972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.341972Z digest=sha256:b7721ca0d5f1d0e25b1ac2f1ea894e5fd69b388f452aa0f89b6faa74566dc548

Observation 03619f5c-1633-405e-a256-803fcd2d4e81 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.334057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.347469Z digest=sha256:127a4acb8be8e1d84bbbae4af44d1563efd473031f61aab37d983329fca7e3aa

Observation bfa26d7c-1cd2-4395-9cd4-0e91a68d634d · outbound

This paper cites Broken Neural Scaling Laws.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Broken Neural Scaling Laws

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.353468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.353468Z digest=sha256:a38584fcddc1f797dbc84bded82ec92c25840bf7454100f618eff38ecfb42152

Observation 6c24d1c1-a387-45a4-ab02-9f5023e092bc · outbound

This paper cites Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.359142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.359142Z digest=sha256:5d4dc085b2f4e2c94b8e2bed9a2c767d3b74f79c7522eb3db7eebabe8b5d7f7a

Observation f225e47c-d49a-40d2-8c6e-9d2916046295 · outbound

This paper cites Scaling Exponents Across Parameterizations and Optimizers.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Scaling Exponents Across Parameterizations and Optimizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.364789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.364789Z digest=sha256:fd8440c2eae0d36e08534eaf4b2985f774fe2a63a685cb50ce0a805fee3c707c

Observation ab140c9e-642f-4dae-939a-209b848441b4 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.370181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.370181Z digest=sha256:b7cfb74880d15f91e1ec8865c90dc1b7e689cb2dded478d523325b157714cdf3

Observation 3bb8d72f-b197-4dc5-883b-01621b12d273 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 11

Resolution
verified exact
doi, observed 2026-08-07T10:31:20.907800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.375809Z digest=sha256:ad49d87604697fa06748c05d81de7bfe58468cc9ef7bf45d802460d754699899

Observation 2988fb8e-629f-48d4-ace9-b76f3659ff9c · outbound

This paper cites Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.382188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.382188Z digest=sha256:094f35c7d7ff8c0bd8e6390bf55e6f3146291ed11b834de26746107d76cdfe30

Observation 1027a3ea-4290-441e-9d5e-0e74972add43 · outbound

This paper cites OLMo: Accelerating the Science of Language Models.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning OLMo: Accelerating the Science of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.387856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.387856Z digest=sha256:55c2cfaa426d6f6627a1f12a2fa599273ee9d98161201401873b9a6944ee3bee

Observation c9652438-c0eb-450e-aaa4-e0fe3a553bd3 · outbound

This paper cites Rae, and Laurent Sifre.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Rae, and Laurent Sifre

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:31:21.308776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.394487Z digest=sha256:78d30acb9624f98c09e228c53cb6723ff00e5a34e90bcd2c0f3efef4fa3acf56

Observation 7fbdc2cd-c03f-4aed-8ad7-76703561d66d · outbound

This paper cites Learning Curve Theory.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Learning Curve Theory

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.400174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.400174Z digest=sha256:0627ba199994c8447e3f09f226228146f1872e7c3ff667b1320048d7dd9fd123

Observation 0e645430-c0ac-4d66-8e53-a2ae77ccffed · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.288053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.406165Z digest=sha256:27e92a629ded4f5d236dac7d5b8845ce3dbd60ec310dd8ca5cb63608e2c11895

Observation 4de56d0a-b8f0-4b67-ba50-8e5e6a1e5294 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.271940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.411836Z digest=sha256:9069a3277dcc08f00e1965ee671442af2fba9492326c1504c3180e2e07cda6d8

Observation 190826fb-ace0-4549-bb0a-34db1785d24c · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.417724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.417724Z digest=sha256:1bf2c7522f1946aaf3f08900430edfb5b7afacd89f8158ec1d81fa377b0903f4

Observation ffb6496d-3839-474e-ad3d-49ff68b8cb74 · outbound

This paper cites Scaling Laws for Neural Language Models.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Scaling Laws for Neural Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.423534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.423534Z digest=sha256:3aa8fd69db43dba804947cd65548e53f49c8b86c2a205ae463a4f51f75512ae5

Observation 811dc1f4-9f45-4c92-b30e-d18d83c128aa · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.252089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.428496Z digest=sha256:4985e7b94818efb126768d97f2dda268cef7aa8635e5b3c7dfb9108f428f32c7

Observation cb665d93-206a-4ac8-b17e-ec2bffb28845 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.433609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.433609Z digest=sha256:8f99b97da88e483135bc4b68a0581b7b53b7cc1ef4fbe810c99b229774d47a4c

Observation 2e8a43a8-a092-47cd-8fa6-5eecde1d9bf7 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.235256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.441055Z digest=sha256:e14dde482480ce0994e18345edf53bbacf126b7c463e286c7b6e09d172eb5ec5

Observation 8e99f216-4cb2-4585-a3ea-18661cbc1fa2 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.446261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.446261Z digest=sha256:7c518d4911cd2bd07356edfa5043f86162c107fa4b47bb6e687195c7e7a942c5

Observation b5d9cf26-1b55-47ec-b814-6e278a268e92 · outbound

This paper cites Paloma: A Benchmark for Evaluating Language Model Fit.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Paloma: A Benchmark for Evaluating Language Model Fit

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.450742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.450742Z digest=sha256:6b532da51e341de62c829f7e9624ee44a2af057717d9fe16e7ed23b1df94f78e

Observation 64bb2aec-abbc-4d16-aafe-645de5f60d18 · outbound

This paper cites When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.455874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.455874Z digest=sha256:113d02bbd570fb6c132ab76151af0a7709b5fa67eab71e55c7cae0d3578e7654

Observation 0afb2e2b-3515-4df5-8662-a91f2621106d · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.202276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.460735Z digest=sha256:ee1ff6365b6962ee05177363ca1055edcc1f550a3e3d0895966d8ae57f782867

Observation 736de624-1040-4b06-a353-90424ceda43e · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.182023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.465412Z digest=sha256:7b2b21c931381648b243173291d29257f6e84c8e83764deb91a83bbec51d32d6

Observation f894a601-7ebb-4e98-9e4d-ad84dd5733de · outbound

This paper cites In-context Learning and Induction Heads.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning In-context Learning and Induction Heads

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.470207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.470207Z digest=sha256:efe1f4861713607e82f6d62bda2e5de06dcb26131d68eb9a03e9dc02c5b1b5b4

Observation fb5891ae-b06e-403b-9648-dfb55cf572dd · outbound

This paper cites The AdEMAMix Optimizer: Better, Faster, Older.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning The AdEMAMix Optimizer: Better, Faster, Older

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.474636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.474636Z digest=sha256:aba038b9d1f790f0a8049ee5256d4811eeff88f36532b64b756dedf8e6d84a62

Observation 9da32be0-03f4-4c1a-b0d1-cb583aa553db · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.166392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.479519Z digest=sha256:99fb0fb9d86394d5d84d1b42b7a59c0d335ec5913858761a3756604c084ecd94

Observation 77df1296-b4a4-4f56-962b-3c7e4e1842d9 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.149829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.484527Z digest=sha256:99978f3b8c2c67b66a7c5a70702c3eb8be8a39d0446764be9daaa044d7d80ea6

Observation 07729239-587b-4a1a-ae67-0d987caa8ced · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.490228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.490228Z digest=sha256:dfebb970b6d0509927c4df8f572795cb0fa019e63c4b77ada2388e7674babf2f

Observation 3d2cf27f-cac3-4065-9290-b35030a91c69 · outbound

This paper cites Outliers with Opposing Signals Have an Outsized Effect on Neural Network Optimization.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Outliers with Opposing Signals Have an Outsized Effect on Neural Network Optimization

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:31:20.681375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.496089Z digest=sha256:22cbc0f1174560144d7e109efeed4488b2b422f66b0b1c26abb3d743b04be1f7

Observation 6162da8a-0c80-4923-bf02-e979cd7c508b · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.133494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.501651Z digest=sha256:7c1fdee62b48cea3e6c0e06f63e1a7eba6fb54fe5d0606d6e8449d1dbbfb855c

Observation 0acf9fcd-435d-40de-8a1c-bb12889ed4e6 · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:21.117029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.506125Z digest=sha256:3f0591bf6a1feab11a010430f0eac56b72d49418a6a805de4f98fefbce3f3eef

Observation e74392e0-56ae-4ef1-9ee2-b5fde6cddedc · outbound

This paper cites an unresolved cited work.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.511694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.511694Z digest=sha256:26fa8fcad3dfdcb0fe4f8461a4b5926c9891e062ed389071f4b4c81f4f00f298

Observation 2b8ec33d-d619-4193-a0c7-3afc4d71caad · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Gemma 2: Improving Open Language Models at a Practical Size

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.516813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.516813Z digest=sha256:83bca8ff15d581c814c13b6af71e7b90219102acfed5e985c53226cd0ba8fe83

Observation 8b1f55d0-2325-4943-aadd-c42eb35e4330 · outbound

This paper cites Scaling Law with Learning Rate Annealing.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Scaling Law with Learning Rate Annealing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.522740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.522740Z digest=sha256:6ac1c07181ec789f2c2702c2c0b264f1f93b346dfc50999461fa95f959e3fe60

Observation 0ef2f770-b496-4cd6-ae99-715b7f8e8c68 · outbound

This paper cites The Shape of Learning Curves: a Review.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning The Shape of Learning Curves: a Review

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.528055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.528055Z digest=sha256:2141af78d918fc5a1abc4bb21471d6ce0d35b2b206bc03edfc059557eb605d45

Observation 22a06aaa-38ec-4f17-98a5-881b1a0aa397 · outbound

This paper cites Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.534009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.534009Z digest=sha256:1d7ed566d17c913a55ce222edd8a70d9352b97039199025ab2feaabe52b7eaaa

Observation a13dbdef-d570-432a-b52a-5c9b2282293a · outbound

This paper cites Data-Dependence of Plateau Phenomenon in Learning with Neural Network --- Statistical Mechanical Analysis.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Data-Dependence of Plateau Phenomenon in Learning with Neural Network --- Statistical Mechanical Analysis

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:31:21.027574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T10:31:20.539480Z digest=sha256:4065a592bf84e89088f637fb55f47e3ab1d9faccbe641eceaf02171540de660c

Observation 380618aa-20e5-4ba1-9f0c-1ad070b497d9 · outbound

This paper cites Opacus: User-Friendly Differential Privacy Library in PyTorch.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Opacus: User-Friendly Differential Privacy Library in PyTorch

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.544798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.544798Z digest=sha256:0a0084fc8c31b8f33febb4dad2ad9d328bae7f88a14228f1887731f05dc1c214

Observation 62f9b0f3-c2ac-4530-842f-51ca8ffb82cf · outbound

This paper cites Gradient Surgery for Multi-Task Learning.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Gradient Surgery for Multi-Task Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.550988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.550988Z digest=sha256:f617596000108b7fe9df4371f2cf042c0a3824177f6f4fe837ab288ba52edc3a

Pith citing papers

Observation 21d0767b-bbd7-44ae-b12e-a8508b3d6055 · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-01T03:08:36.328547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-01T03:02:05.516512Z digest=sha256:cabb94649c55df34ea18cf25f3d5ea38a002561f4393ff9e17b78750a9a93bfe