Pith. sign in

Paper Citation Record · LEDGER

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

As of 6 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2605.26895.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.26895 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T19:56:54.047126Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T10:14:12.034776Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T20:30:07.719081Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact19
  • verified fuzzy0
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f6d6abf6-c617-43bd-99f5-e0a0db965c59 · outbound

This paper cites On the optimization of deep networks: Implicit acceleration by overparameterization.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models On the optimization of deep networks: Implicit acceleration by overparameterization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:037fd0b0e030da8fc88de7f4570ac9aa52582955d3be1a07fdf9204d0cb9a389

Observation be1ed8c9-061f-47e1-8c2b-2fa5736ea1c1 · outbound

This paper cites Layer Normalization.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Layer Normalization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.377843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:960a4e6fca3c0ec4c256fc5c696b246d79e4781f9c5f82aa6598a2f2983e4b02

Observation 7740af1e-5abc-4fa9-ab43-027344c3c110 · outbound

This paper cites Optimization methods for large-scale machine learning.SIAM review, 60(2):223–311, 2018.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Optimization methods for large-scale machine learning.SIAM review, 60(2):223–311, 2018

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:cfa775f7ee9caed139da6b0b739c24f111b354aa62ef3b525b77433e8b6668a9

Observation 73c0ad59-05ac-41fc-aef6-a8cf0d736575 · outbound

This paper cites Seednorm: Self-rescaled dynamic normalization.arXiv preprint arXiv:2510.22777, 2025.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Seednorm: Self-rescaled dynamic normalization.arXiv preprint arXiv:2510.22777, 2025

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:03:56.392947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:46044cad5dbc24270a5970a421713e4f32a84a975f60db5aefa93b52251b74a8

Observation c24cae3c-4c71-49f5-bb39-cd4dbb4baa3b · outbound

This paper cites Post-layernorm is back: Stable, expressive, and deep.arXiv preprint arXiv:2601.19895, 2026.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Post-layernorm is back: Stable, expressive, and deep.arXiv preprint arXiv:2601.19895, 2026

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:03:56.354792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:e9abd84082a9f5a3ffc427183f7e0d98084aded43a36fade50b2b35ec4932e3c

Observation a0381982-7521-4446-a942-06f7f078b0d3 · outbound

This paper cites Label noise SGD provably prefers flat global minimizers.Advances in Neural Information Processing Systems, 34:27449–27461, 2021.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Label noise SGD provably prefers flat global minimizers.Advances in Neural Information Processing Systems, 34:27449–27461, 2021

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:997420255c8c97deb285b2c0af6391b086f02bd9005614c86c687487541861f1

Observation 563167d2-5db2-4da5-97c5-274a4b6bcf0c · outbound

This paper cites Scaling vision transformers to 22 billion parameters.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Scaling vision transformers to 22 billion parameters

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:dbf2ea995eec26108f14b1812fdcd3d6c17eb5ca9f70af1af2805da3043a51ed

Observation aad33d69-f219-4309-9bc4-20206e40352d · outbound

This paper cites The inverse variance–flatness relation in stochastic gradient descent is critical for finding flat minima.Proceedings of the National Academy of Sciences, 118(9):e2015617118, 2021.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models The inverse variance–flatness relation in stochastic gradient descent is critical for finding flat minima.Proceedings of the National Academy of Sciences, 118(9):e2015617118, 2021

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:d06fa5d2a8f447e4a4f4bb0341c3ea123c6b1c33a05a8f96410e8ae62ff985f1

Observation aa54f2e4-338c-4884-a997-1a193ffbe27a · outbound

This paper cites Shampoo: Preconditioned stochastic tensor optimization.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Shampoo: Preconditioned stochastic tensor optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:485241da90e1d68d7b332913fb7d490e748777be21b28517176a4b9b022748f6

Observation fe55fa24-f041-4dd1-9ed0-1e0cf2177635 · outbound

This paper cites Shape matters: Understanding the implicit bias of the noise covariance.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Shape matters: Understanding the implicit bias of the noise covariance

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:724c81e6de67dc75e3ded5ee4580042329aa1699a3eaba8cf86f67986b043b08

Observation be4249ca-42b4-40b0-bd1c-8f5e91ffe824 · outbound

This paper cites Introduction to online convex optimization.Foundationsand Trends®in Optimization, 2(3-4): 157–325, 2016.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Introduction to online convex optimization.Foundationsand Trends®in Optimization, 2(3-4): 157–325, 2016

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:82ab87af8649363e21eee10bdb63bcb457f9b560ee12e6764983fb5c1c5735da

Observation 6c3e6461-6baa-48d5-bbc6-9586e4540a15 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Training Compute-Optimal Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.389862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:9caa3476b1dd9f2bdea792e9284f9c1486180a213ed17c1efd99b050c049365c

Observation 8086825e-b5f4-4692-b21d-3ce67c91b14d · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.367612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:2bebcf863a8e0f257d7b438563ae37b9028879fb70ddaefb73712442091ad701

Observation 4f71305d-73b7-4d70-882e-fca01f58e58b · outbound

This paper cites Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.360360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:ca8b64ba4048aecc594aac489f8dc057b10745703005a1ea679a68b61c8ce46d

Observation e5b62692-d2c3-4f98-9aad-1a19e2bd022a · outbound

This paper cites an unresolved cited work.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:da9899513dc1dda60b4ae074d0f129e3003184653352d81379aab301437eec10

Observation 7fa44733-cacf-498b-899e-a32e52f58720 · outbound

This paper cites Muon optimizer.URL https://github.com/KellerJordan/Muon?tab=readme-ov-file, 2024.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Muon optimizer.URL https://github.com/KellerJordan/Muon?tab=readme-ov-file, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:5539390ef9ac4183805242e2af5ff07c7f9e6d4b7620c9842b7e916b1058513c

Observation b0795519-9065-46c4-9819-74de45bd95e2 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Adam: A Method for Stochastic Optimization

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.387598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:7736dc658251058335f449691bfc93e5b12ae9c5648496ac29ee4ff62365e8f5

Observation e762af73-3c2d-4a56-aa4e-20f81ae699a5 · outbound

This paper cites Stochastic modified equations and adaptive stochastic gradient algorithms.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Stochastic modified equations and adaptive stochastic gradient algorithms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:3f6b9ab75613008fd7a262ce8d5b11229f37851c3f26bf7ced5a99c3f87e7ec7

Observation 3802c18e-7dd9-4db0-a8fd-24a9e2427693 · outbound

This paper cites What happens after SGD reaches zero loss?–a mathematical framework.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models What happens after SGD reaches zero loss?–a mathematical framework

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:066b1de207a5cd4f87dadc3c4d0ca17d814a4e0570a119618243b0f05cdc4e5c

Observation 2d697157-3b3d-4a0d-84ea-dae74da3b441 · outbound

This paper cites Muon is Scalable for LLM Training.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Muon is Scalable for LLM Training

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.382885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:9caf9a71dc011add408542924ee76d593ea4a8f1b79bc4719f256e9224f17307

Observation ccd897ff-7d7d-4b49-916b-c6a55070a280 · outbound

This paper cites Noise and fluctuation of finite learning rate stochastic gradient descent.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Noise and fluctuation of finite learning rate stochastic gradient descent

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:739b641b3535f4f207d3048daab186c3841f2b3af31d07a4ad4ab0addbce6039

Observation 2eee105d-aa22-4d74-a3f1-9c07e298930e · outbound

This paper cites Decoupled Weight Decay Regularization.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Decoupled Weight Decay Regularization

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.362868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:922aee1a9ed14b5c0e8b4e216e1d0cfd4f2362a584aafe993290ed860d49355d

Observation 8bdfbc38-ee2b-4c5d-a588-e30c4a66f3f1 · outbound

This paper cites Optimizing neural networks with kronecker-factored approximate curvature.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Optimizing neural networks with kronecker-factored approximate curvature

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:fde9daf9913347401377404a76922468438450dbb9b17cd62acde1b63cc2e9f1

Observation 229fa5f6-b91b-4260-b47b-b7cfe68eb540 · outbound

This paper cites Power-law escape rate of SGD.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Power-law escape rate of SGD

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:03:56.372698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:c5058cee40e489fef03754c430374bdb773a1978c5257b0232890a1e9c42d7b4

Observation 43dc79a1-f30a-4837-8e97-59f363d273e7 · outbound

This paper cites Power-law escape rate of sgd.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Power-law escape rate of sgd

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:dda92f72c64b0dc49fd4dd0105b41151d09af076839ca77ca90ae78edea73b56

Observation a20dea62-5991-4a16-9724-79f52a353b8b · outbound

This paper cites Transformers without tears: Improving the normalization of self-attention.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Transformers without tears: Improving the normalization of self-attention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:8a6539c890dc4ab893388be0951b67c8b6f2011fff472f3e4b2602c7d9fbb868

Observation 86de0c30-2a11-40e6-8cd7-179624a610be · outbound

This paper cites 2 OLMo 2 Furious.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models 2 OLMo 2 Furious

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.370226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:e3bbbbf9191b864b96a1853c1c62a7de37982f1da82ac82806e4b76fdd1473d6

Observation 6d02f29f-e9e8-4d49-926c-a25aca2057fd · outbound

This paper cites A unified view of attention and residual sinks: Outlier-driven rescaling is essential for transformer training.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models A unified view of attention and residual sinks: Outlier-driven rescaling is essential for transformer training

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:03:56.348937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:bbbed728e12bce157b5665ea64c1619ef6c18072ec63f35733a1cb00f8b7325b

Observation bc063411-3758-4471-bda5-d7e09e89a366 · outbound

This paper cites Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:a23179d776430bdd4b1e6f2757b4e2b2e6ebd5f9e1c3f3f90f07a8f37d6f04f5

Observation 781a9a10-2291-4cab-94c9-b55e86dea0ba · outbound

This paper cites Exact solutions to the nonlinear dynamics of learning in deep linear neural networks.International Conference on Learning Representations, 2014.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Exact solutions to the nonlinear dynamics of learning in deep linear neural networks.International Conference on Learning Representations, 2014

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:0f27f3b807289e8b78263fa29cb43f7e88de42ad039fe5ea05e011f51f7efb20

Observation fa631fa3-04d9-462f-8e79-a12d46d152ca · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:c4396a92a666e541fb612a2db6a814a39b60b0d92d824b067d92eaa0ac8b2ebc

Observation f8152584-9a2d-4647-9b51-e1493635dfdb · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Gemma 2: Improving Open Language Models at a Practical Size

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.357522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:d261d30e85250c0fdcc6b700052301e8280440e0723d376f1867c2d417973e19

Observation ac12bede-7af8-429f-a75e-bdeab9238f83 · outbound

This paper cites Gemma 3 technical report, 2025.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Gemma 3 technical report, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:500bcd4c46e49c58300d277a19ab00983b56ee9265ed4417638477d8aeafb976

Observation f9400ecb-df4b-4eb8-9543-5a1a02d9ed55 · outbound

This paper cites Tieleman and G.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Tieleman and G

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:c666f176db90bdf933d03cd816db3cfc9924ec048c651cd5d928035f504915b8

Observation 8ef7b968-45f8-458c-b16f-d722427c8cc0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.365142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:56d87073ad7b1aa619abef192fe2954d2d43eee4c2d66eb6469b46450cdd944a

Observation 9f20bf04-4300-44b0-8ac3-b13f92c56f58 · outbound

This paper cites Attention is all you need.Advancesin neural information processing systems, 30, 2017.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Attention is all you need.Advancesin neural information processing systems, 30, 2017

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:cb3b6af29c1ba8eaf42cce1960ae1b79a9203eecc00448931345e76415d5c2a3

Observation c5524de1-3089-4f4f-89b7-d32b035335e4 · outbound

This paper cites Soap: Improving and stabilizing shampoo using adam for language modeling.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Soap: Improving and stabilizing shampoo using adam for language modeling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:bf3a447a61fda60beaa109605c652f6670fb673b708e4d7c479d27e297fe26aa

Observation f927861a-c287-4f22-8414-a935da506baf · outbound

This paper cites Deepnet: Scaling transformers to 1,000 layers.IEEE Transactionson Pattern Analysis and Machine Intelligence, 46(10):6761–6774, 2024.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Deepnet: Scaling transformers to 1,000 layers.IEEE Transactionson Pattern Analysis and Machine Intelligence, 46(10):6761–6774, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:b2c4106088362a635f755da373321599af7f203dcef98d63390fca5a99e809b0

Observation c6ce5507-9cde-4779-b69e-f99be7ff6771 · outbound

This paper cites The sharpness disparity principle in transformers for accelerating language model pre-training.International Conference on Machine Learning, pages 64859–64879, 2025.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models The sharpness disparity principle in transformers for accelerating language model pre-training.International Conference on Machine Learning, pages 64859–64879, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:3dadfd370630f2412d58d9e7a4cc60c983a817e884ea1634defd014af47ef78f

Observation e6704bdd-adb1-4559-a6db-bb215eefb6bf · outbound

This paper cites Gradpower: Powering gradients for faster language model pre-training.International Conference on Machine Learning, 2026.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Gradpower: Powering gradients for faster language model pre-training.International Conference on Machine Learning, 2026

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:df80f1b5ecde2f0560895fc10c1b17edf5be37a2a0a867ebf1292873b3ddffc4

Observation 21d359c4-b7af-48de-9e93-613595c7e38c · outbound

This paper cites A Theoretical Analysis of Noise Geometry in Stochastic Gradient Descent.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models A Theoretical Analysis of Noise Geometry in Stochastic Gradient Descent

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T20:03:56.375366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:04cdff27110d02d73a9209e90122bb1a26c463f5ca8f0354f5be94a62753b765

Observation 492f1ed6-efa2-4daf-823c-e2410eff2bad · outbound

This paper cites Improving generalization and convergence by enhancing implicit regularization.Advances in Neural Information Processing Systems, 2024.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Improving generalization and convergence by enhancing implicit regularization.Advances in Neural Information Processing Systems, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:d90e16ecab7e1162f509857258eba248994f41d6774c13b43879e3481e330c05

Observation cf9f3de2-31a2-420c-bd45-26326b1cff4a · outbound

This paper cites Bayesian learning via stochastic gradient langevin dynamics.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Bayesian learning via stochastic gradient langevin dynamics

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:667749808935e000e55fb5b144e037386a64b31d629bc9aa2389677bd086c616

Observation 832087ea-8f8e-47c5-bc71-c2ec5f8c04b1 · outbound

This paper cites Stochastic gradient descent with noise of machine learning type. Part II: Continuous time analysis.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Stochastic gradient descent with noise of machine learning type. Part II: Continuous time analysis

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:03:56.380331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:47de0321e015453dc75792b385edd7e3015f2acb3291d1d735d1f025d0f410b6

Observation a5ca6b9c-f9f9-4c65-b800-16c329c1264c · outbound

This paper cites The alignment property of sgd noise and how it helps select flat minima: A stability analysis.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models The alignment property of sgd noise and how it helps select flat minima: A stability analysis

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:841a462d6c2d76b46bc6c30e0ff091aec4016569163254958c2781f38620e6e6

Observation a0d7d435-572a-49e8-9dec-ddf14da7dabd · outbound

This paper cites On layer normalization in the transformer architecture.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models On layer normalization in the transformer architecture

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:b9b81e2549db208f834416a391042454ed98f39dc13ee63371a6cfdec8feb5c7

Observation a64532f0-4ff0-4ae7-b228-00a36c377464 · outbound

This paper cites Qwen2 Technical Report.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Qwen2 Technical Report

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.385040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:90b01b8e8f1d70f8b4437a9b68670fe2b9221cf59af089bc4067db559cef6ef2

Observation 2224e8a4-c0cb-4f8f-b368-d9ab2f785c05 · outbound

This paper cites Qwen3 Technical Report.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Qwen3 Technical Report

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.351806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:3ae08f3097ffb84aae51e6a92f755427b57aa18459df5b22849c33b8395efda9

Observation 6445e9cd-4947-42aa-ad06-3609600da68a · outbound

This paper cites Scaling vision transformers.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Scaling vision transformers

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:4a07ab2fc24b4e31afebbd64c496895a7962a7d22d9260107cbf7c72f4e7bd31

Observation ea413e91-d95c-498c-95b7-d2863b9f0c20 · outbound

This paper cites Root mean square layer normalization.Advancesin Neural Information Processing Systems, 32, 2019.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Root mean square layer normalization.Advancesin Neural Information Processing Systems, 32, 2019

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:95173808d782127062d1a9da9d4bb5f8111ea3b2fe85e7ec3150651b4b985157

Observation b3b3fa41-7013-45d3-b3c8-b746c5dfed62 · outbound

This paper cites Transformers without normalization.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Transformers without normalization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:6e5ea4933f6dd41a2ff09c13652a6d447cb3dfa3340dd5dd11c5616c82aa30d2

Observation b8333a4a-3b2d-479b-9b98-5206ef6a6822 · outbound

This paper cites arXiv preprint arXiv:2602.22681 , year=.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models arXiv preprint arXiv:2602.22681 , year=

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:03:56.343296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:fb7f0b07112aec8215ecf3eca325e2a02dd714ce14f9af4c670d9d3d1c829509

Observation b67ff9bd-9274-440b-812b-594310558c26 · outbound

This paper cites HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:03:56.345993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:fdc1053a9e1371feb48883d2102aaa2d641547c101fa2a6291002185ce08fd5c

Observation 3a4233b1-3f6a-4b7e-8454-87f75ec667cc · outbound

This paper cites Parameter symmetry and noise equilibrium of stochastic gradient descent.Advancesin Neural Information Processing Systems, 37:93874–93906, 2024.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Parameter symmetry and noise equilibrium of stochastic gradient descent.Advancesin Neural Information Processing Systems, 37:93874–93906, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:368e822be9953d39cd109358064249e0006521f9a159aa0839c48c6660358a0a

Observation 92c2c495-d03c-4e03-b568-cdf8b149eb2e · outbound

This paper cites Sincew=0π-a.s., we also have a=γ⊙w=0π-a.s.

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models Sincew=0π-a.s., we also have a=γ⊙w=0π-a.s

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-29T19:56:54.047126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:56:54.047126Z digest=sha256:100d3d223f36839c89133c7df4508f96690cb908fc16820599d84b0673454019

Pith citing papers

Observation 5ccad840-ac3a-4815-864f-33ebe42ec849 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

Reference 112

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T20:30:07.720403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:51d9a8fb681f99060cf1b8f5203b9173e80f4ad61267a4a8592eda6483499386

Observation 85f586ff-f2a5-42b6-89f4-830783f91b0d · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:12.034776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:12.034776Z digest=sha256:03cac1f25e406a4cbe689b3c36139d44fc48258a09d7247491efa9718bd38bd2