Pith. sign in

Paper Citation Record · LEDGER

Linear attention is (maybe) all you need (to understand transformer optimization)

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2310.01082.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.01082 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:44:47.613549Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T15:04:46.136709Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ec1e4f59-413a-46e9-8bfc-2cae6704f589 · inbound

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization cites this paper.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:47.613549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:47.613549Z digest=sha256:47d6fbde2b26e560569526a220aa784a5deb99ad30a2a1d19aea8dafa2cec8e4

Observation 2829d5ae-d210-43c7-a57f-8b0f2fb32faf · inbound

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization cites this paper.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.118380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.118380Z digest=sha256:325311e9db0b90b752f9dbd48b7c7ea568b736c3abd11489676b3389d3d55002

Observation 3c56be0f-8610-4504-8510-4e9111b6f761 · inbound

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling cites this paper.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:29.838100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:29.838100Z digest=sha256:063fec530be8006f8daff7c13d38caeea707ef765dcee1a12e055d320255f56c

Observation 1afaba9e-c5ad-414e-a3c8-348c35f33682 · inbound

Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression cites this paper.

Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:17.419244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:10:17.419244Z digest=sha256:6e4a92bc839f50abe73c14831b070a0095eff2ebd02e217c953d4ada8303c478

Observation afa979e0-9e86-4032-993b-d27729694443 · inbound

Nonconvex Decentralized Stochastic Bilevel Optimization under Heavy-Tailed Noise cites this paper.

Nonconvex Decentralized Stochastic Bilevel Optimization under Heavy-Tailed Noise Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:15.929444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:15.929444Z digest=sha256:7b0275121f91b1bbaa4ba30d5d2668a6a793b3390f19803c950d6f949f37ca15

Observation f831fb28-cc3a-427a-a346-77d8a08ce436 · inbound

Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters cites this paper.

Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.337503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T07:01:37.086159Z digest=sha256:2df170374fcd0c4fef3b86e5e1f53b9e6375c8f2cb269cdfd62fa52fb02d6391

Observation 611116c3-b893-4fc0-b03c-8dbe8a9f2ca8 · inbound

Zeroth-Order Nonconvex Nonsmooth Optimization with Heavy-Tailed Noise cites this paper.

Zeroth-Order Nonconvex Nonsmooth Optimization with Heavy-Tailed Noise Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:04:46.138420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T15:04:30.204185Z digest=sha256:a0dceeb82227ca14a032ce94f51576dab4d4211bfe60c0311583ecc995e5251d

Observation 4c67f297-5b71-4076-905f-578153a0e7ea · inbound

The Convergence Behavior of Adam under Heavy-Tailed Noise cites this paper.

The Convergence Behavior of Adam under Heavy-Tailed Noise Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T08:37:46.250989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:37:46.250989Z digest=sha256:432124f4b075a9c5ca7fb7e958e2b33d7bce887be4a21a989d955ca5985bbcfc

Observation 05b88bd0-d395-4f1b-8db1-dd7b93dd0aa7 · inbound

The Convergence Behavior of Adam under Heavy-Tailed Noise cites this paper.

The Convergence Behavior of Adam under Heavy-Tailed Noise Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T03:30:54.192015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:30:54.192015Z digest=sha256:c269609cacafc65d4c98625363349b0bc894392e7b3ffe655975205eb9efc204