Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2405.13698.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:30.277440Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T20:30:07.608050Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 2a802f31-d8a0-4a08-a5ba-7884a5a91ec3 · inbound
MuLoCo: Muon is a practical inner optimizer for DiLoCo How to set AdamW's weight decay as you scale model and dataset size
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16c13459-0ce0-4da9-8d2f-dab0399daa4c · inbound
Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks How to set AdamW's weight decay as you scale model and dataset size
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b442b740-5ea5-493a-b93e-605a2a8ff4ac · inbound
Weight Decay Improves Language Model Plasticity How to set AdamW's weight decay as you scale model and dataset size
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62b52619-7bc6-46fd-90cd-b5cbe7b9bb18 · inbound
Rethinking Language Model Scaling under Transferable Hypersphere Optimization How to set AdamW's weight decay as you scale model and dataset size
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b57ebadf-4f86-465b-9aff-eab0e2766a85 · inbound
GQA-{\mu}P: The maximal parameterization update for grouped query attention How to set AdamW's weight decay as you scale model and dataset size
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 347648e8-8a1f-4e37-9355-d8eec62b6283 · inbound
A Two-Parameter Weibull Framework for Diagnosing Transformer Weight Distributions How to set AdamW's weight decay as you scale model and dataset size
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f297a98d-c26e-4bc0-a89f-45ca90684f91 · inbound
Muon Learns More Robust and Transferable Features than Adam How to set AdamW's weight decay as you scale model and dataset size
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a73887b-e141-43ab-ad54-9fd14da19f9b · inbound
Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors How to set AdamW's weight decay as you scale model and dataset size
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 30de1928-ab05-4104-8e23-21758f905ff7 · inbound
Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors How to set AdamW's weight decay as you scale model and dataset size
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.