Pith. sign in

Paper Citation Record · LEDGER

How to set AdamW's weight decay as you scale model and dataset size

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2405.13698.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.13698 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:30.277440Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:30:07.608050Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2a802f31-d8a0-4a08-a5ba-7884a5a91ec3 · inbound

MuLoCo: Muon is a practical inner optimizer for DiLoCo cites this paper.

MuLoCo: Muon is a practical inner optimizer for DiLoCo How to set AdamW's weight decay as you scale model and dataset size

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:30.277440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:30.277440Z digest=sha256:1d3451fb8bc3bc27f3bdc93cfa9f2a4d969e9d9842d1fd9afa53cd249c72fb83

Observation 16c13459-0ce0-4da9-8d2f-dab0399daa4c · inbound

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks cites this paper.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks How to set AdamW's weight decay as you scale model and dataset size

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:54.538611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:54.538611Z digest=sha256:3ff0aa94d30074ca7052c6333c4ed0409d14da9ca52ece351dbfdb40addd92e4

Observation b442b740-5ea5-493a-b93e-605a2a8ff4ac · inbound

Weight Decay Improves Language Model Plasticity cites this paper.

Weight Decay Improves Language Model Plasticity How to set AdamW's weight decay as you scale model and dataset size

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T00:15:37.343946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:15:37.343946Z digest=sha256:5d4dfdb06901f71ab13c5bd760b6a5dc421b4815796f952f5d041ba7c3076734

Observation 62b52619-7bc6-46fd-90cd-b5cbe7b9bb18 · inbound

Rethinking Language Model Scaling under Transferable Hypersphere Optimization cites this paper.

Rethinking Language Model Scaling under Transferable Hypersphere Optimization How to set AdamW's weight decay as you scale model and dataset size

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:53:02.441467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:51:14.678941Z digest=sha256:8dc2e3739e4077eff78945449900a28706cf94ccb7bbe7eae7f05f1f4b3e1a94

Observation b57ebadf-4f86-465b-9aff-eab0e2766a85 · inbound

GQA-{\mu}P: The maximal parameterization update for grouped query attention cites this paper.

GQA-{\mu}P: The maximal parameterization update for grouped query attention How to set AdamW's weight decay as you scale model and dataset size

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.822671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T16:35:40.231293Z digest=sha256:ca013185479c70fdd5a725b3fc479e4353082b9cda23777f78241a0eb8b90ff0

Observation 347648e8-8a1f-4e37-9355-d8eec62b6283 · inbound

A Two-Parameter Weibull Framework for Diagnosing Transformer Weight Distributions cites this paper.

A Two-Parameter Weibull Framework for Diagnosing Transformer Weight Distributions How to set AdamW's weight decay as you scale model and dataset size

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:23:21.546428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T14:20:15.499338Z digest=sha256:4d37100c73bed120e13c9421027c4c7f7c2e92556aff1349a678b6050b286b4e

Observation f297a98d-c26e-4bc0-a89f-45ca90684f91 · inbound

Muon Learns More Robust and Transferable Features than Adam cites this paper.

Muon Learns More Robust and Transferable Features than Adam How to set AdamW's weight decay as you scale model and dataset size

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:30.106422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T17:08:30.717799Z digest=sha256:d34f327a71a82aca5e5742e567b471042d92d6b51d1f1b94145a0bfb33d172d4

Observation 0a73887b-e141-43ab-ad54-9fd14da19f9b · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors How to set AdamW's weight decay as you scale model and dataset size

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:30:07.609833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:2449dd6583ae4b1a36de359dc93500e9a10d2b469023c374c83bace679b0cd75

Observation 30de1928-ab05-4104-8e23-21758f905ff7 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors How to set AdamW's weight decay as you scale model and dataset size

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:12.180596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:12.180596Z digest=sha256:b8810ad36ffea947690260e077d004ec3ea147eca960c950513c65650b533925