Pith. sign in

Paper Citation Record · LEDGER

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention

As of 22 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2507.09394.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09394 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:07:38.566888Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T00:52:08.399401Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T05:05:58.687554Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee683523-97b9-4efe-a703-6c44ed8cdb82 · outbound

This paper cites A random matrix perspective on mix- tures of nonlinearities in high dimensions.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention A random matrix perspective on mix- tures of nonlinearities in high dimensions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:07:41.900833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:07:37.168580Z digest=sha256:cdcfd8c4e254c948c5f9ca3d97fe8b32b6aa11e693d49f649a6acafe371094cd

Observation 8cf650c4-73bd-4dc8-86ba-fab5670b39a9 · outbound

This paper cites Self-attention networks localize when QK- eigenspectrum concentrates.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention Self-attention networks localize when QK- eigenspectrum concentrates

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:07:41.736678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:07:37.209065Z digest=sha256:8334a881d982e191eb1d4f5b31abf277d07e6beb64397a449409efb1b84a5b40

Observation 3ee33393-c328-446e-824b-1de74fadeb71 · outbound

This paper cites Random matrix theory improved fr ´echet mean of symmetric positive definite matrices.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention Random matrix theory improved fr ´echet mean of symmetric positive definite matrices

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:07:41.568076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:07:37.253547Z digest=sha256:39cf3927fdecb8d578c979d0de53b51bb5996400774378224af61e1756fdd0ac

Observation 349b43ec-d02e-4ff6-83fb-abf279ddd3c0 · outbound

This paper cites A random matrix ap- proach to echo-state neural networks.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention A random matrix ap- proach to echo-state neural networks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:07:41.319175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:07:37.296250Z digest=sha256:da8ada4e92834705430166fabbef79c6148bb603e79cc6a94aad4585f91d96fe

Observation 81c403de-6cd6-47f2-a51c-71fb29dd1ce1 · outbound

This paper cites A random matrix theory perspective on the spectrum of learned features and asymptotic generalization capabilities.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention A random matrix theory perspective on the spectrum of learned features and asymptotic generalization capabilities

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:07:41.118098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:07:37.345132Z digest=sha256:f2329006e924ae66945ad88774ba53e5237fb09a44043993d00d72627faa60a9

Observation a11f441f-56f7-4e0d-9000-0ac68f94eb44 · outbound

This paper cites Random matrix analysis to balance between supervised and unsupervised learning under the low density separation assumption.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention Random matrix analysis to balance between supervised and unsupervised learning under the low density separation assumption

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:07:40.943455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:07:37.408099Z digest=sha256:629fdcec3ec1bce8a8b4632bf737a0db9830f1a948c0c569f2fbe5f105018872

Observation 98dffd79-1563-419d-923b-573a5551b30b · outbound

This paper cites Maximizing the potential of synthetic data: Insights from ran- dom matrix theory.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention Maximizing the potential of synthetic data: Insights from ran- dom matrix theory

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:07:40.719825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:07:37.472843Z digest=sha256:9268c56b59060a508fcfe1c2248919e276532e14547385a7d9750ed8a0964a9d

Observation 45bcfe2a-a32e-45dd-8d56-451575538272 · outbound

This paper cites Analysing multi-task regression via random matrix theory with application to time series forecasting.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention Analysing multi-task regression via random matrix theory with application to time series forecasting

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:07:40.541879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:07:37.534906Z digest=sha256:d876b0b157f9e4f2888865a8e1ddbad80f87d7820de49b1ce03303ca7550a5ef

Observation 25ea31cf-a862-4943-97aa-7e3c08b5dcae · outbound

This paper cites The Underlying Scaling Laws and Universal Statistical Structure of Complex Datasets.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention The Underlying Scaling Laws and Universal Statistical Structure of Complex Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:07:37.592873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:07:37.592873Z digest=sha256:5d7684f5ddfadaa85cae0b73a1a346f0f8016369ea96e8e4b0425267c8d89d0e

Observation 8971b61d-d2b0-4d5a-b90d-7d9437f98b6b · outbound

This paper cites Mix-LN: Unleashing the power of deeper layers by combining pre-LN and post-LN.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention Mix-LN: Unleashing the power of deeper layers by combining pre-LN and post-LN

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:07:40.349433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:07:37.662358Z digest=sha256:28ad912a2553f24d1aaa6704575f47cfe5d8925507b3572a80bebb0777b221a1

Observation c6ac8cd5-f5f1-4af1-8e41-5c10500bd25b · outbound

This paper cites The dynamics of learning: A random matrix approach.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention The dynamics of learning: A random matrix approach

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:07:40.148126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:07:37.723896Z digest=sha256:645d28717de6255523c8bbfc0dbcc2aba54191ea576c84b380c1b3556af78c6a

Observation 2b056b2e-0d9d-42bc-877e-fa12dc11cac8 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:07:37.781457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:07:37.781457Z digest=sha256:87a960d0d4b5332ca76456ece809309821222b002cebc643028d928761b2e141

Observation 8c6a1f5a-6fe6-46b8-be9a-d5c50b87b7b5 · outbound

This paper cites DeepSeek-V3 Technical Report.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention DeepSeek-V3 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:07:37.848948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:07:37.848948Z digest=sha256:e6138926881650c87180392623646754fc2459904605813426d3d1362aac6d15

Observation 230f14ec-31ac-4344-95e8-7acaf51b5acc · outbound

This paper cites Distribution of eigenvalues for some sets of random matrices.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention Distribution of eigenvalues for some sets of random matrices

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:07:40.021515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:07:37.905643Z digest=sha256:ae69360b872596b88e957c86fed9c4a7c2a8e1a4fd175d30a0e5a92d703fe98a

Observation 43395909-49b0-4b56-84a4-8915f5ab0f54 · outbound

This paper cites Implicit self-regularization in deep neural net- works: Evidence from random matrix theory and implications for learning.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention Implicit self-regularization in deep neural net- works: Evidence from random matrix theory and implications for learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:07:39.837235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:07:37.956487Z digest=sha256:f06f4e9905694079bd5dc083dad83a786cd62dc750a176c6a7c7b3c5cd449ff7

Observation db4c0ef7-315a-436c-89d6-569d992d4d08 · outbound

This paper cites TransMLA: Multi-Head Latent Attention Is All You Need.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention TransMLA: Multi-Head Latent Attention Is All You Need

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:07:38.019228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:07:38.019228Z digest=sha256:8795c8d9d1c8e2fca41dc5727ce187b21260cbb3f50160621d1a9e88ed6fd625

Observation cb10a6e6-f7da-4471-82ed-b3178aad6735 · outbound

This paper cites Geometry of neural network loss surfaces via random matrix theory.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention Geometry of neural network loss surfaces via random matrix theory

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:07:39.600408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:07:38.104880Z digest=sha256:41afd1b4c23e6a0c5b081b1e4e5be35fdfbf9b5d2301c6b04e76f1eef03bc48d

Observation f8ca5be1-61db-472c-a785-0e266df756c9 · outbound

This paper cites Nonlinear random matrix theory for deep learning.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention Nonlinear random matrix theory for deep learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:07:38.183231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:07:38.183231Z digest=sha256:eca08df1bfd3b65568098d17b4bc3d6f5accfb6953b84015fe22c968e540baf6

Observation 2aa809c4-6732-48d9-af69-59694030d5fe · outbound

This paper cites Locating information in large language models via random matrix theory.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention Locating information in large language models via random matrix theory

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:07:38.258595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:07:38.258595Z digest=sha256:5840d4c86605fa22f5d1fa13e4dfa2dd866eba764e528c1208b89ada6a25ca1a

Observation f33cf979-feb5-402e-82c3-4c4caf4073d4 · outbound

This paper cites Random matrix theory analysis of neural network weight matrices.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention Random matrix theory analysis of neural network weight matrices

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:07:39.412438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:07:38.333559Z digest=sha256:cda2643803500109bcbdffe1f1b8b99a1e938b96c28da56b80edef234c6e0b1b

Observation 36013000-9d1b-44b9-aaea-1875b88768bc · outbound

This paper cites Random ma- trix improved covariance estimation for a large class of metrics.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention Random ma- trix improved covariance estimation for a large class of metrics

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:07:39.199843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:07:38.399805Z digest=sha256:37442b7c497d213eb67b04e7dd0d64e2d8ce3cf21fd84377d92988d8d76bc6e2

Observation 5cd7fe0c-24b6-48c5-b213-3872f97779b1 · outbound

This paper cites More than a toy: Random matrix models pre- dict how real-world neural representations generalize.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention More than a toy: Random matrix models pre- dict how real-world neural representations generalize

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:07:39.022322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:07:38.453182Z digest=sha256:02dc801318459827b92d5c3bd8d240289bd7afbc83b5c5ec33a47d3f200c5959

Observation 4d42d40f-dc65-4ff9-b0d8-1b5136db0d19 · outbound

This paper cites Insights into deepseek-v3: Scaling challenges and reflections on hardware for ai architectures.

A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention Insights into deepseek-v3: Scaling challenges and reflections on hardware for ai architectures

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:07:38.566888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:07:38.566888Z digest=sha256:d8e8951eb7dea9e0257de3c0da0f4c67d5773490a1e98908c7190124075f0f64

Pith citing papers

Observation 51d386c9-4ea1-4944-a112-cfc16267f99a · inbound

How Does Attention Help? Insights from Random Matrices on Signal Recovery from Sequence Models cites this paper.

How Does Attention Help? Insights from Random Matrices on Signal Recovery from Sequence Models A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:05:58.694020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T00:52:08.399401Z digest=sha256:8211f88ff949764535186f297ec3b21abda8874c5ec04b628ad2027cc7d5c760