Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-30T12:53:41.205912Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 0 inbound Pith citation observations for arXiv:2607.23777.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-30T12:53:41.205912Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
74 of 74 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 779190fd-96e4-45cc-b818-7c5e72291d75 · outbound
Scale Weight Decay and Train Better Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 276f9969-64b7-42fd-9f5b-83ec6d6a3308 · outbound
Scale Weight Decay and Train Better Scaling Laws for Neural Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ec2ef78-4e4b-4ff2-8dd9-ac917bbf00f5 · outbound
Scale Weight Decay and Train Better Training Compute-Optimal Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a77a1adf-f488-4b0f-83a5-4a20f16a1d13 · outbound
Scale Weight Decay and Train Better Explaining neural scaling laws
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cf1cc50-7468-421f-aad3-e21fb22704e6 · outbound
Scale Weight Decay and Train Better An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55b13f6c-2c76-4e87-b39a-92a8567d0e1e · outbound
Scale Weight Decay and Train Better Scaling Vision Transformers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fd4085d-e984-4019-ab2c-1d6c49ae3d4c · outbound
Scale Weight Decay and Train Better Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e874d25-fdcd-44f7-8324-a1de5c0199b4 · outbound
Scale Weight Decay and Train Better Neural Scaling Laws in Robotics
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11f0e4a3-670c-453c-8ddd-786047876460 · outbound
Scale Weight Decay and Train Better Kimi K2: Open Agentic Intelligence
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccc70d7e-a03d-487c-8525-aa737e8f1e96 · outbound
Scale Weight Decay and Train Better GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d353d982-f150-4baa-9ad5-6b6b537de4e2 · outbound
Scale Weight Decay and Train Better Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0669252d-c402-42c7-bd53-32a758d5c53c · outbound
Scale Weight Decay and Train Better A Simple Weight Decay Can Improve Generalization
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b48b4df0-84dc-45e0-8e09-424743d8b37b · outbound
Scale Weight Decay and Train Better Why Do We Need Weight Decay in Modern Deep Learning?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cfdcf2d-ac1d-4ba0-9c3e-9e5461e0649c · outbound
Scale Weight Decay and Train Better Decoupled Weight Decay Regularization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9737d18c-75bd-4ee0-bc67-7445f196d449 · outbound
Scale Weight Decay and Train Better Why Warmup the Learning Rate? Underlying Mechanisms and Improvements
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 278735a2-2063-4ab2-9f53-c192fd02a7ac · outbound
Scale Weight Decay and Train Better SGDR: Stochastic Gradient Descent with Warm Restarts
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03e11547-df7e-45e9-961d-a968c1cc15e5 · outbound
Scale Weight Decay and Train Better 2024.url:https : //kellerjordan.github.io/posts/muon/
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a26d2d7-3548-4363-826a-20840195459d · outbound
Scale Weight Decay and Train Better Muon is Scalable for LLM Training
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cd153e1-72dc-4c18-94d5-1872b22e2f1e · outbound
Scale Weight Decay and Train Better A Stochastic Approximation Method
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e669c20-a196-457f-8b62-74c44e0c2476 · outbound
Scale Weight Decay and Train Better Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18742985-3239-465c-a100-fa1b65c49831 · outbound
Scale Weight Decay and Train Better Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf743017-35ec-4b6e-a066-38f82f0bb173 · outbound
Scale Weight Decay and Train Better Why Gradients Rapidly Increase Near the End of Training
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 875cf34d-83ba-4843-8eb2-c3b3de5b968c · outbound
Scale Weight Decay and Train Better Shampoo: Preconditioned Stochastic Tensor Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 142487d7-9078-4488-8392-681cf80f7dc5 · outbound
Scale Weight Decay and Train Better Scalable Second Order Optimization for Deep Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb89b922-ff20-4ccf-be67-3ca20a3dccef · outbound
Scale Weight Decay and Train Better SOAP: Improving and Stabilizing Shampoo using Adam
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98613942-ad1f-44a9-9718-ed96c96bd207 · outbound
Scale Weight Decay and Train Better Aurora: A Leverage-Aware Spectral Optimizer
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffc2807d-03c3-4f46-9f86-c9d5212623e6 · outbound
Scale Weight Decay and Train Better Dion: Distributed Orthonormalized Updates
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 483bd6a9-0653-4ad7-8ffd-2b40df48d89f · outbound
Scale Weight Decay and Train Better Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9973faa-5435-47d3-a70a-ed5cc56f6bac · outbound
Scale Weight Decay and Train Better The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 911aa333-4fbd-400d-93a8-4a6c81f74004 · outbound
Scale Weight Decay and Train Better Anytime Training with Schedule-Free Spectral Optimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c7fb9f6-0bcf-42ec-be46-d5395a69548d · outbound
Scale Weight Decay and Train Better Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43073ace-70a4-47b1-8997-4d7465993ec2 · outbound
Scale Weight Decay and Train Better Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a23e4123-d953-48da-9821-295eba574a0c · outbound
Scale Weight Decay and Train Better Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92e69bab-9d3e-4f8c-b25c-0c5f98a6dd57 · outbound
Scale Weight Decay and Train Better Optimization Methods for Large-Scale Machine Learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e04770ed-0bf2-400d-aa57-03816929c724 · outbound
Scale Weight Decay and Train Better Muown: Row-Norm Control for Muon Optimization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc72e5ca-0af6-449e-8e6a-a644cf546291 · outbound
Scale Weight Decay and Train Better Convergence Bound and Critical Batch Size of Muon Optimizer
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9df2c4f6-02af-4a7c-bc0a-45ede472414b · outbound
Scale Weight Decay and Train Better On the Convergence Analysis of Muon
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5627a78-50d9-4101-929b-57ece23b4f20 · outbound
Scale Weight Decay and Train Better On the Convergence of Muon and Beyond
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3155f39f-83fc-4fd8-9db8-fd674e0bf77f · outbound
Scale Weight Decay and Train Better Language Models are Unsupervised Multitask Learners
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd28eaa8-3601-43c0-8170-5269eafe5a5a · outbound
Scale Weight Decay and Train Better Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38cbf6cf-6290-4515-aa64-80b98712a8be · outbound
Scale Weight Decay and Train Better Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a31867d-ee51-4a9c-a149-e080d835209c · outbound
Scale Weight Decay and Train Better Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db4d0e64-30a4-4959-871d-4543df27f422 · outbound
Scale Weight Decay and Train Better Scaling Optimal LR Across Token Horizons
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4c0e474-cbe7-43f0-8789-3ee316ab8a79 · outbound
Scale Weight Decay and Train Better L2 Regularization versus Batch and Weight Normalization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08e3f009-f43f-49fc-aaa1-048424f4be36 · outbound
Scale Weight Decay and Train Better Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98bd111d-c63b-49d3-bf56-722ec3e78ec2 · outbound
Scale Weight Decay and Train Better Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a7f9e62-8d0b-4e5d-b464-af15068857ed · outbound
Scale Weight Decay and Train Better 2022.url: https://openreview.net/forum?id=J7V_4aauV6B
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b2cd47a-2e82-4f08-8730-f15f62c23f53 · outbound
Scale Weight Decay and Train Better Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ead716dd-08e9-45f8-82e0-348633c9b71a · outbound
Scale Weight Decay and Train Better Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e0c32fe-8577-4f5d-97ec-90c1cca4c89c · outbound
Scale Weight Decay and Train Better Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3337b5f0-3d51-4517-a330-ccb02fa01e20 · outbound
Scale Weight Decay and Train Better Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130db0ba-753b-4621-8371-0c81a93b997d · outbound
Scale Weight Decay and Train Better Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4be47bb7-04b2-48d5-b48f-175cc1858efc · outbound
Scale Weight Decay and Train Better Gated Delta Networks: Improving Mamba2 with Delta Rule
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6381bee1-28d0-4edd-926d-8d7797d0d13b · outbound
Scale Weight Decay and Train Better Parallax: Parameterized Local Linear Attention for Language Modeling
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75dac705-d999-4a3f-bec5-0e797ecd4b76 · outbound
Scale Weight Decay and Train Better Training language models to follow instructions with human feedback
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7ce3610-a63b-4f44-94d1-14a6decf3732 · outbound
Scale Weight Decay and Train Better Deep reinforcement learning from human preferences
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87455bc6-5cb8-485e-929a-4c26d58a777c · outbound
Scale Weight Decay and Train Better Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad191bf0-4ca6-4f5c-bb49-cf36bf670141 · outbound
Scale Weight Decay and Train Better DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 620e6b72-6b4d-4a11-92a8-b313380f6b19 · outbound
Scale Weight Decay and Train Better DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80569cfd-2968-4e13-864e-cd06e2a97b27 · outbound
Scale Weight Decay and Train Better Scaling Laws for Fine-Grained Mixture of Experts
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fea8ca95-1122-4e78-ab85-9d8c13c619ce · outbound
Scale Weight Decay and Train Better Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8dc26db-bf8f-4bb6-b57e-7126535a1c92 · outbound
Scale Weight Decay and Train Better Inference economics of language models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b2f1fff-d1a5-4f20-9eda-f16568a08cfb · outbound
Scale Weight Decay and Train Better Root Mean Square Layer Normalization
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed33a81e-32e0-4c9c-8d04-753f9a6eba21 · outbound
Scale Weight Decay and Train Better RoFormer: Enhanced Transformer with Rotary Position Embedding
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36607c74-c48e-4e93-9420-5b0dbde4f094 · outbound
Scale Weight Decay and Train Better Unresolved cited work
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 970c7596-d80c-40ac-8060-6449687b52f1 · outbound
Scale Weight Decay and Train Better Benchmarking Optimizers for Large Language Model Pretraining
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c8d2a59-dea0-4030-bb23-dd38c5a17c02 · outbound
Scale Weight Decay and Train Better A Spectral Condition for Feature Learning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e09b8ad-c1b1-4069-adca-33ab7def3d1a · outbound
Scale Weight Decay and Train Better Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c428c06-d042-42ca-8493-4167a46d629a · outbound
Scale Weight Decay and Train Better GLU Variants Improve Transformer
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5b7f52a-9f64-4d35-80c5-a5db93db43e4 · outbound
Scale Weight Decay and Train Better PyTorch: An Imperative Style, High-Performance Deep Learning Library
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93edebdd-6373-46d8-a16b-c672ba690870 · outbound
Scale Weight Decay and Train Better Unresolved cited work
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad9c477a-93a7-45b9-a38d-34e26d5816b1 · outbound
Scale Weight Decay and Train Better Rethinking Language Model Scaling under Transferable Hypersphere Optimization
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca591d46-0694-44a9-8a69-f1879540b5c5 · outbound
Scale Weight Decay and Train Better The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64874ec6-a8ef-4481-9818-fec04fc460fd · outbound
Scale Weight Decay and Train Better Unresolved cited work
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.