Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T19:56:54.047126Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2605.26895.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T19:56:54.047126Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T10:14:12.034776Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T20:30:07.719081Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f6d6abf6-c617-43bd-99f5-e0a0db965c59 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models On the optimization of deep networks: Implicit acceleration by overparameterization
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be1ed8c9-061f-47e1-8c2b-2fa5736ea1c1 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Layer Normalization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7740af1e-5abc-4fa9-ab43-027344c3c110 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Optimization methods for large-scale machine learning.SIAM review, 60(2):223–311, 2018
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c0ad59-05ac-41fc-aef6-a8cf0d736575 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Seednorm: Self-rescaled dynamic normalization.arXiv preprint arXiv:2510.22777, 2025
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c24cae3c-4c71-49f5-bb39-cd4dbb4baa3b · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Post-layernorm is back: Stable, expressive, and deep.arXiv preprint arXiv:2601.19895, 2026
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a0381982-7521-4446-a942-06f7f078b0d3 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Label noise SGD provably prefers flat global minimizers.Advances in Neural Information Processing Systems, 34:27449–27461, 2021
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 563167d2-5db2-4da5-97c5-274a4b6bcf0c · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Scaling vision transformers to 22 billion parameters
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aad33d69-f219-4309-9bc4-20206e40352d · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models The inverse variance–flatness relation in stochastic gradient descent is critical for finding flat minima.Proceedings of the National Academy of Sciences, 118(9):e2015617118, 2021
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa54f2e4-338c-4884-a997-1a193ffbe27a · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Shampoo: Preconditioned stochastic tensor optimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe55fa24-f041-4dd1-9ed0-1e0cf2177635 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Shape matters: Understanding the implicit bias of the noise covariance
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be4249ca-42b4-40b0-bd1c-8f5e91ffe824 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Introduction to online convex optimization.Foundationsand Trends®in Optimization, 2(3-4): 157–325, 2016
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c3e6461-6baa-48d5-bbc6-9586e4540a15 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Training Compute-Optimal Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8086825e-b5f4-4692-b21d-3ce67c91b14d · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4f71305d-73b7-4d70-882e-fca01f58e58b · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e5b62692-d2c3-4f98-9aad-1a19e2bd022a · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fa44733-cacf-498b-899e-a32e52f58720 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Muon optimizer.URL https://github.com/KellerJordan/Muon?tab=readme-ov-file, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0795519-9065-46c4-9819-74de45bd95e2 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Adam: A Method for Stochastic Optimization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e762af73-3c2d-4a56-aa4e-20f81ae699a5 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Stochastic modified equations and adaptive stochastic gradient algorithms
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3802c18e-7dd9-4db0-a8fd-24a9e2427693 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models What happens after SGD reaches zero loss?–a mathematical framework
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d697157-3b3d-4a0d-84ea-dae74da3b441 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Muon is Scalable for LLM Training
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ccd897ff-7d7d-4b49-916b-c6a55070a280 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Noise and fluctuation of finite learning rate stochastic gradient descent
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eee105d-aa22-4d74-a3f1-9c07e298930e · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Decoupled Weight Decay Regularization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8bdfbc38-ee2b-4c5d-a588-e30c4a66f3f1 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Optimizing neural networks with kronecker-factored approximate curvature
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 229fa5f6-b91b-4260-b47b-b7cfe68eb540 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Power-law escape rate of SGD
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 43dc79a1-f30a-4837-8e97-59f363d273e7 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Power-law escape rate of sgd
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a20dea62-5991-4a16-9724-79f52a353b8b · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Transformers without tears: Improving the normalization of self-attention
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86de0c30-2a11-40e6-8cd7-179624a610be · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models 2 OLMo 2 Furious
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6d02f29f-e9e8-4d49-926c-a25aca2057fd · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models A unified view of attention and residual sinks: Outlier-driven rescaling is essential for transformer training
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bc063411-3758-4471-bda5-d7e09e89a366 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 781a9a10-2291-4cab-94c9-b55e86dea0ba · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Exact solutions to the nonlinear dynamics of learning in deep linear neural networks.International Conference on Learning Representations, 2014
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa631fa3-04d9-462f-8e79-a12d46d152ca · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8152584-9a2d-4647-9b51-e1493635dfdb · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Gemma 2: Improving Open Language Models at a Practical Size
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ac12bede-7af8-429f-a75e-bdeab9238f83 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Gemma 3 technical report, 2025
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9400ecb-df4b-4eb8-9543-5a1a02d9ed55 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Tieleman and G
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ef7b968-45f8-458c-b16f-d722427c8cc0 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9f20bf04-4300-44b0-8ac3-b13f92c56f58 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Attention is all you need.Advancesin neural information processing systems, 30, 2017
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5524de1-3089-4f4f-89b7-d32b035335e4 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Soap: Improving and stabilizing shampoo using adam for language modeling
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f927861a-c287-4f22-8414-a935da506baf · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Deepnet: Scaling transformers to 1,000 layers.IEEE Transactionson Pattern Analysis and Machine Intelligence, 46(10):6761–6774, 2024
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6ce5507-9cde-4779-b69e-f99be7ff6771 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models The sharpness disparity principle in transformers for accelerating language model pre-training.International Conference on Machine Learning, pages 64859–64879, 2025
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6704bdd-adb1-4559-a6db-bb215eefb6bf · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Gradpower: Powering gradients for faster language model pre-training.International Conference on Machine Learning, 2026
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21d359c4-b7af-48de-9e93-613595c7e38c · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models A Theoretical Analysis of Noise Geometry in Stochastic Gradient Descent
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 492f1ed6-efa2-4daf-823c-e2410eff2bad · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Improving generalization and convergence by enhancing implicit regularization.Advances in Neural Information Processing Systems, 2024
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf9f3de2-31a2-420c-bd45-26326b1cff4a · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Bayesian learning via stochastic gradient langevin dynamics
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 832087ea-8f8e-47c5-bc71-c2ec5f8c04b1 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Stochastic gradient descent with noise of machine learning type. Part II: Continuous time analysis
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a5ca6b9c-f9f9-4c65-b800-16c329c1264c · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models The alignment property of sgd noise and how it helps select flat minima: A stability analysis
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0d7d435-572a-49e8-9dec-ddf14da7dabd · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models On layer normalization in the transformer architecture
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a64532f0-4ff0-4ae7-b228-00a36c377464 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Qwen2 Technical Report
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2224e8a4-c0cb-4f8f-b368-d9ab2f785c05 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Qwen3 Technical Report
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6445e9cd-4947-42aa-ad06-3609600da68a · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Scaling vision transformers
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea413e91-d95c-498c-95b7-d2863b9f0c20 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Root mean square layer normalization.Advancesin Neural Information Processing Systems, 32, 2019
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3b3fa41-7013-45d3-b3c8-b746c5dfed62 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Transformers without normalization
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8333a4a-3b2d-479b-9b98-5206ef6a6822 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models arXiv preprint arXiv:2602.22681 , year=
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b67ff9bd-9274-440b-812b-594310558c26 · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3a4233b1-3f6a-4b7e-8454-87f75ec667cc · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Parameter symmetry and noise equilibrium of stochastic gradient descent.Advancesin Neural Information Processing Systems, 37:93874–93906, 2024
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92c2c495-d03c-4e03-b568-cdf8b149eb2e · outbound
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Sincew=0π-a.s., we also have a=γ⊙w=0π-a.s
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ccad840-ac3a-4815-864f-33ebe42ec849 · inbound
Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 85f586ff-f2a5-42b6-89f4-830783f91b0d · inbound
Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.