Pith. sign in

Paper Citation Record · LEDGER

Linear attention is (maybe) all you need (to understand transformer optimization)

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2310.01082.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.01082 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:32:03.704458Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T15:04:46.136709Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ec1e4f59-413a-46e9-8bfc-2cae6704f589 · inbound

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization cites this paper.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:47.613549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:47.613549Z digest=sha256:c22ccc717fda7902ba9f085772d65338af03bf8a87e77b5dae120d14a16f77c6

Observation 2829d5ae-d210-43c7-a57f-8b0f2fb32faf · inbound

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization cites this paper.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.118380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.118380Z digest=sha256:be2b150bade29d8ec8076a11e9cea2feb014b473a542290f031b6eedbc2c6b34

Observation 3c56be0f-8610-4504-8510-4e9111b6f761 · inbound

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling cites this paper.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:29.838100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:29.838100Z digest=sha256:1297a2de77607319d0c2984ec1e39af95932ca97e4b526b1b76c972a5bf5251b

Observation 1afaba9e-c5ad-414e-a3c8-348c35f33682 · inbound

Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression cites this paper.

Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:17.419244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:10:17.419244Z digest=sha256:801e71569e365612f91cecf9b449aa5b3ab66962daf4a81cd6e2e55393ae7ebb

Observation 8a9d747b-9a42-4583-9d1b-02b07b4e6363 · inbound

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts cites this paper.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.704458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.704458Z digest=sha256:2ef031d6296791f405af89086c2561fb1ad8c550afdd4f3c48ef40ee153936cd

Observation afa979e0-9e86-4032-993b-d27729694443 · inbound

Nonconvex Decentralized Stochastic Bilevel Optimization under Heavy-Tailed Noise cites this paper.

Nonconvex Decentralized Stochastic Bilevel Optimization under Heavy-Tailed Noise Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T16:23:15.929444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:23:15.929444Z digest=sha256:ea5222712d2292b4c845b7be65eaa1875acb1bab61158720f525e2b27f9f3a22

Observation f831fb28-cc3a-427a-a346-77d8a08ce436 · inbound

Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters cites this paper.

Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.337503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T07:01:37.086159Z digest=sha256:64f28ba5152b32ba311ea38daebfa3b4b28f4771e61a600e3548433554fea956

Observation 611116c3-b893-4fc0-b03c-8dbe8a9f2ca8 · inbound

Zeroth-Order Nonconvex Nonsmooth Optimization with Heavy-Tailed Noise cites this paper.

Zeroth-Order Nonconvex Nonsmooth Optimization with Heavy-Tailed Noise Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:04:46.138420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T15:04:30.204185Z digest=sha256:55c55f2c12b9305c77e70be135901cc6fca4d3ae9684d9c5351bd4b780168208

Observation 4c67f297-5b71-4076-905f-578153a0e7ea · inbound

The Convergence Behavior of Adam under Heavy-Tailed Noise cites this paper.

The Convergence Behavior of Adam under Heavy-Tailed Noise Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T08:37:46.250989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:37:46.250989Z digest=sha256:589131e9c23f8de55d30ba079803f2a716628ea86dd011aa6c6691ad9ca88614

Observation 05b88bd0-d395-4f1b-8db1-dd7b93dd0aa7 · inbound

The Convergence Behavior of Adam under Heavy-Tailed Noise cites this paper.

The Convergence Behavior of Adam under Heavy-Tailed Noise Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T03:30:54.192015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:30:54.192015Z digest=sha256:7aff6573ac16a7989b286fbb903e34892d021f9da08437d26add5fa4265e9ca9