Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T08:15:20.191222Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2606.29158.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T08:15:20.191222Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1ed80a9e-47b6-467a-a23c-9d2ae7b9bb60 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Tune My Adam, Please!
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b476bcbc-729f-4330-94de-9e18dd9d44db · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Layer Normalization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2b49e810-a187-4c01-be95-494c36a627de · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Power lines: Scaling laws for weight decay and batch size in llm pre-training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d04d66e2-b1fe-4124-bc94-a89b80481c78 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Scaling optimal lr across token horizons
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2942bb31-e203-45f3-807e-d8b56bbdb2f0 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Y., Deiseroth, B., Cruz-Salinas, A
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 85bfd103-e0ed-4bf5-8edf-3ae048ad1712 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Depthwise hyperparameter transfer in residual networks: Dynamics and scaling limit
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 326afa66-a69d-41c6-b494-96c0f9d9240f · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation acdbe677-c97d-4750-a1ca-d887f791c183 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Don’t be lazy: Completep enables compute-efficient deep transformers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fad2912c-901a-4aa5-ac76-44710a4a7a07 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training E., Xiao, L., Wortsman, M., Alemi, A
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8cfee3b9-eef0-4808-9210-cc8a42fb94df · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Robust layerwise scaling rules by proper weight decay tuning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d5ba71d2-ef7d-4300-a530-e599984dba5b · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Nemotron-flash: Towards latency-optimal hybrid small language models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e5b1cde-973b-4d68-9a8f-4bef60ff7942 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Norm matters: efficient and accurate normalization schemes in deep networks.Advances in Neural Information Processing Systems, 31, 2018
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b6e36f17-a54f-4e53-a522-bcbc3966accc · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Minicpm: Unveiling the potential of small language models with scalable training strategies
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b9fc8db1-5b25-45c6-b8e5-5912fbd4607b · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training H., and Leyton-Brown, K
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 216b8889-4b67-4bab-a95b-9f06bbd1caa0 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training and Szegedy, C
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f1ed5e63-8183-4b39-9fb3-943fbbad2508 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Three Factors Influencing Minima in SGD
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b2d893dc-ba19-47b2-9599-720bda639dba · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Scaling Laws for Neural Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 93482c5a-9c43-43d1-8828-1011787364a5 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training nanoGPT.https://github.com/karpathy/nanoGPT, 2022
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 70994301-6fd4-488d-9b4b-cb94d4fd16c2 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Rotational equilibrium: How weight decay balances learning across neural networks
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1d47b165-16e1-4c2f-b048-95fcb2473e01 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Weight decay may matter more than mup for learning rate transfer in practice
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9b8a732c-2eb5-45f3-af9d-0924c8286b4f · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Efficient hyperparameter tuning via trajectory invariance principle.arXiv preprint arXiv:2509.25049, 2025
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 23542f01-bb32-401a-9b90-90b1e45102bb · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 910d5cd0-ae41-426c-805e-22c39419f4f0 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Reconciling modern deep learning with traditional optimization analyses: The intrinsic learning rate.Advances in Neural Information Processing Systems, 33: 14544–14555, 2020
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fc444885-f740-4f4d-bc3f-7df9c1d70f0a · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Muon is Scalable for LLM Training
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d7dd65cb-4a23-4d67-83b9-227a7d2daab7 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training and Hutter, F
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f04a8424-4645-408b-a2fd-edd22af09e6b · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training ngpt: Normalized transformer with representation learning on the hypersphere
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fb782221-1cd1-4f28-9fa5-ed58f30af273 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training A multi-power law for loss curve prediction across learning rate schedules
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7336a97c-5837-4d9d-832f-2fce3e899626 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training On the sdes and scaling rules for adaptive gradient algorithms.Advances in Neural Information Processing Systems, 35:7697–7711, 2022
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 70b8bbfd-6f20-459a-b87b-080c0bc99d32 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training G., and Goldblum, M
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6c3812b3-5e75-4cf6-8c25-6dd533c7b0c2 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training An Empirical Model of Large-Batch Training
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8d1dc6d3-028a-4b07-aa18-9ce07405935b · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Completed hyperparameter transfer across modules, width, depth, batch and duration
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f0dc6fca-e4f4-49b3-be1a-e3b4f7e87669 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training The fineweb datasets: Decanting the web for the finest text data at scale.Advances in Neural Information Processing Systems, 37:30811–30849, 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 70f0f246-979c-4d3a-9506-8d9675bd2dab · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Resolving Discrepancies in Compute-Optimal Scaling of Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a7ae3b72-9161-42ae-9408-24e534c7285f · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Language models are un- supervised multitask learners
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 87995147-01b3-4849-ad29-b17493b54b62 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 068cd734-0431-4fcc-8a9b-977367f48ca5 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Scaling Law with Learning Rate Annealing
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b0bb89b-2852-45b5-8e8c-2359a173e3d5 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training L2 Regularization versus Batch and Weight Normalization
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ec48e0a1-bd5e-4a9c-b52e-9034a6388842 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training N., Kaiser, L
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0b715f71-b032-424c-bf81-edaac5d03c5b · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training The sharpness disparity principle in transformers for accelerating language model pre-training
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 81f308be-1414-4f09-97d9-38b1fe3beed7 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Scaling laws across model architectures: A comparative analysis of dense and M o E models in large language models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 14de7d54-e1a2-439f-a978-41d89baa9892 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training and Aitchison, L
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d5b9adb0-4a9f-418d-bc2e-5ad7c3c40901 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Fantastic pretraining optimizers and where to find them 2.1: Hyperball optimization, 12 2025
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ed70c51b-7dd4-42f4-9e0f-e43c287b5ac2 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Controlled llm training on spectral sphere
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ac30bcd1-4fd1-4bdd-b1f5-2cddd2272537 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training and Hu, E
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 31917e82-803b-4c49-8f4a-7a98ccdc03d2 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Tuning large neural networks via zero-shot hyperparameter transfer.Advances in Neural Information Processing Systems, 34:17084–17097, 2021
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 46be5cf6-44e2-4b78-b138-60d16ba37af6 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fbcbe078-632d-4549-a4b2-31dae2eaa7de · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training arXiv , author =:2602.10300 , file =
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1e8e661a-dcd3-42cf-8e8d-9ce6fdbf9989 · outbound
On the Nonlinearity of Learning Rate Scaling for LLM Training How to set the learning rate for large-scale pre-training?arXiv preprint arXiv:2601.05049, 2026
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
No inbound Pith citation observations are available.