Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T20:45:31.427677Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2409.04777.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T20:45:31.427677Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T22:38:09.654884Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-04T22:38:10.015046Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 71b857a4-ad49-4b43-bbad-b05c7db2613b · outbound
Optimization Hyper-parameter Laws for Large Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 899e10d2-4277-4c5c-8efb-a159e65415de · outbound
Optimization Hyper-parameter Laws for Large Language Models Qwen Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9d989cec-7771-41d1-a130-01b1618bc8a9 · outbound
Optimization Hyper-parameter Laws for Large Language Models Chinchilla Scaling: A replication attempt
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f25b8a45-79ba-4bf4-89b6-dcda6f86d320 · outbound
Optimization Hyper-parameter Laws for Large Language Models DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 29fc9b2b-188c-4ba0-b630-2c8f205a7ad6 · outbound
Optimization Hyper-parameter Laws for Large Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation acffcd64-5776-47f1-b1ba-7b4b07a12ea9 · outbound
Optimization Hyper-parameter Laws for Large Language Models Stochastic Bregman Subgradient Methods for Nonsmooth Nonconvex Optimization Problems
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 99dd7265-15b3-4db5-bc31-bd10b8b39e89 · outbound
Optimization Hyper-parameter Laws for Large Language Models Adam-family Methods with Decoupled Weight Decay in Deep Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 30383b20-e117-4d2a-a365-e88e3066d929 · outbound
Optimization Hyper-parameter Laws for Large Language Models The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0d247a0a-72ab-4130-8a6f-053fc7e1a02d · outbound
Optimization Hyper-parameter Laws for Large Language Models Sharpness-Aware Minimization for Efficiently Improving Generalization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f97c4029-899f-445f-bcd8-352a29f3a66b · outbound
Optimization Hyper-parameter Laws for Large Language Models Exponential convergence rates for momentum stochastic gradient descent in the overparametrized setting
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 536f8344-1450-4939-bc2f-fd99b41393a1 · outbound
Optimization Hyper-parameter Laws for Large Language Models Accelerated Objective Gap and Gradient Norm Convergence for Gradient Descent via Long Steps
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dec96b09-c71c-48d9-a0f1-9a0f712cd44f · outbound
Optimization Hyper-parameter Laws for Large Language Models Efficient Continual Pre-training by Mitigating the Stability Gap
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3feba15a-7bca-4398-a3a3-68a93ee0a0ec · outbound
Optimization Hyper-parameter Laws for Large Language Models A Novel Convergence Analysis for Algorithms of the Adam Family
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5c6e8e3e-a02d-4d8b-959e-cbcb79d340f2 · outbound
Optimization Hyper-parameter Laws for Large Language Models Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 58294b13-000a-4652-8853-27cc490a9b5d · outbound
Optimization Hyper-parameter Laws for Large Language Models Scaling Laws for Transfer
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6672f655-71d9-4a21-9daa-bfff810d6d1f · outbound
Optimization Hyper-parameter Laws for Large Language Models Training Compute-Optimal Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 778dbfa0-7a15-4bc5-93da-7da1272a3d8d · outbound
Optimization Hyper-parameter Laws for Large Language Models MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9d1a3c41-6262-40ba-8d3a-7e2dcbd5aa77 · outbound
Optimization Hyper-parameter Laws for Large Language Models Simple and Scalable Strategies to Continually Pre-train Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2920b524-ea34-4253-a30c-c14534a6a0e9 · outbound
Optimization Hyper-parameter Laws for Large Language Models Scaling laws for downstream task performance of large language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2a59b6c9-4c31-43e5-81db-eb207d1d272d · outbound
Optimization Hyper-parameter Laws for Large Language Models Three Factors Influencing Minima in SGD
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2821de96-b257-4f92-8b91-6275d9357eb6 · outbound
Optimization Hyper-parameter Laws for Large Language Models Rethinking Learning Rate Tuning in the Era of Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cee8d0b5-bf8e-47eb-af8b-71f0d9956ee6 · outbound
Optimization Hyper-parameter Laws for Large Language Models Scaling Laws for Neural Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 921063f1-6e6e-4ce5-8f3d-19c5d5d8f569 · outbound
Optimization Hyper-parameter Laws for Large Language Models Continual Pre-training of Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 66668037-1846-4b0c-b002-81c88b430151 · outbound
Optimization Hyper-parameter Laws for Large Language Models On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7cb45d94-fae8-4d42-9c33-cacdb13bfa97 · outbound
Optimization Hyper-parameter Laws for Large Language Models Adam: A Method for Stochastic Optimization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 68927d72-0c6c-41a1-bb77-4ef54908852b · outbound
Optimization Hyper-parameter Laws for Large Language Models Decoupled Weight Decay Regularization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 671738de-d351-4a27-b967-585eccb6ef81 · outbound
Optimization Hyper-parameter Laws for Large Language Models Full Parameter Fine-tuning for Large Language Models with Limited Resources
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 56eabdd4-9e51-49a5-9325-d982698991b5 · outbound
Optimization Hyper-parameter Laws for Large Language Models An SDE Perspective on Stochastic Inertial Gradient Dynamics with Time-Dependent Viscosity and Geometric Damping
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0ba8aaa7-da0c-4ab4-81cd-3341c05d465c · outbound
Optimization Hyper-parameter Laws for Large Language Models Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f3b64c03-6053-4dbe-a6dc-f44a35ed181a · outbound
Optimization Hyper-parameter Laws for Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3b89fc53-5457-4ee3-8787-2eabdaa4ee53 · outbound
Optimization Hyper-parameter Laws for Large Language Models Rotaru, F
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2e2cb812-1518-49a7-bfff-a1f2760f19ec · outbound
Optimization Hyper-parameter Laws for Large Language Models RepoFusion: Training Code Models to Understand Your Repository
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b262fb87-6456-4f4c-a566-632f6612805f · outbound
Optimization Hyper-parameter Laws for Large Language Models An SDE perspective on stochastic convex optimization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9f78d267-bcdb-418d-99d7-86ab56a043d3 · outbound
Optimization Hyper-parameter Laws for Large Language Models An elementary proof of anti-concentration for degree two non-negative Gaussian polynomials
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 733bf6a5-6cd0-414c-a0a7-c47f13c2bb89 · outbound
Optimization Hyper-parameter Laws for Large Language Models Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0f862a80-8a16-4d01-82d7-a0d843490679 · outbound
Optimization Hyper-parameter Laws for Large Language Models Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 770af79a-1280-4f19-9b18-16f11a657b15 · outbound
Optimization Hyper-parameter Laws for Large Language Models LoCo: Low-Bit Communication Adaptor for Large-scale Model Training
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7c9da6ef-cfaa-4ab7-baab-3039495d370a · outbound
Optimization Hyper-parameter Laws for Large Language Models Qwen2 Technical Report
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6ea7a178-6fbc-4f44-8eed-c145524dd004 · outbound
Optimization Hyper-parameter Laws for Large Language Models LongSkywork: A Training Recipe for Efficiently Extending Context Length in Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 238e251c-24b1-4d59-82da-23d5d43a8f14 · inbound
Systematic Optimization of Open Source Large Language Models for Mathematical Reasoning Optimization Hyper-parameter Laws for Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.