Pith. sign in

Paper Citation Record · LEDGER

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining

As of 19 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2509.10406.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.10406 v4

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T17:59:27.996885Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 51b48bad-6533-4f1a-885c-1e8812ddcebe · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R \' e.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining Fu, Stefano Ermon, Atri Rudra, and Christopher R \' e

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.928382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.928382Z digest=sha256:d8ba992ed4f1c3697ebafcc154f634fb15cd05d391ee402beafea93d299f7eee

Observation 16afc1bf-737b-4b2d-b6af-6b25aba21b8c · outbound

This paper cites Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.933354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.933354Z digest=sha256:30a76a1649fdcc1c8d3f3b4b2c3b2b114c68556f931ab669da3a5d63617b9403

Observation 8cbd14e4-e064-4743-9aa3-ab6998ff987b · outbound

This paper cites Reformer: The efficient transformer.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining Reformer: The efficient transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.937761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.937761Z digest=sha256:aa10f5a0322941747310d1305bcda26216712884d0be0582ada4b745c22b66bf

Observation b8d331f9-5133-42b2-afe0-e7f4a5cb7314 · outbound

This paper cites Efficient content-based sparse attention with routing transformers.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining Efficient content-based sparse attention with routing transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.942087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.942087Z digest=sha256:0172c21c77487b6bd9e5a352ef85c4e1494f7422090c6a48cb314c57e3bb54b1

Observation e11901c7-2dd2-47af-8e90-466a9fb233a7 · outbound

This paper cites Mahoney, Kurt Keutzer, and Amir Gholami.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining Mahoney, Kurt Keutzer, and Amir Gholami

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.947033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.947033Z digest=sha256:326c3f63673900b5135fd90bc7b8da1eb018c22703698a72d1e2d89982ac0563

Observation 7909d10f-7544-492b-a031-3af7a75f3a0e · outbound

This paper cites A \( ^2 \) ats: Retrieval-based KV cache reduction via windowed rotary position embedding and query-aware vector quantization.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining A \( ^2 \) ats: Retrieval-based KV cache reduction via windowed rotary position embedding and query-aware vector quantization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.951294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.951294Z digest=sha256:1f5334e0140851e520c3dc56272933c75a899d356c802697607df8180a3dfeaa

Observation 96cab984-272e-4abe-b128-6b36094568c7 · outbound

This paper cites Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.956066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.956066Z digest=sha256:a66a7f6e6d017e6608975a20240dc39f121cf5d7e6246e0fe0f90e03009ebd79

Observation 7cfeb38d-f7ae-4e4e-8320-efecfcd54f9b · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining Linformer: Self-Attention with Linear Complexity

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.960702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.960702Z digest=sha256:53a88c2ad757afc3c1839d3208e008406d020263b383a7babd7552a6147f1d2d

Observation fd573cd5-62a8-4da3-9338-aa71cbcf2501 · outbound

This paper cites o mformer: A nystr \.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining o mformer: A nystr \

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.965032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.965032Z digest=sha256:c2ab30b4d24efd143244a1abc8f472c3c93571a1adf4253ba5bdec9022d18168

Observation b89c88d1-e462-41da-977c-e19f7a150279 · outbound

This paper cites Fast multipole attention: A divide-and-conquer attention mechanism for long sequences.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining Fast multipole attention: A divide-and-conquer attention mechanism for long sequences

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.969031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.969031Z digest=sha256:1e37475aafcb7c00e3f59ccbe10e687a9701f3fdb5a4e2d9ab2220cf63e3ff02

Observation b3d73846-290f-445c-bca3-89d9d087950e · outbound

This paper cites H-transformer-1d: Fast one-dimensional hierarchical attention for sequences.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining H-transformer-1d: Fast one-dimensional hierarchical attention for sequences

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.972982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.972982Z digest=sha256:8b4ada8b6ecaa8dd0b9e7347c0ddb05636c0a354cfb9b193edcf2d6e2fc4a8a4

Observation a8fd9131-cb35-42f7-be53-4b3f11baf0de · outbound

This paper cites A hierarchical o(nlogn) force-calculation algorithm.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining A hierarchical o(nlogn) force-calculation algorithm

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.976889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.976889Z digest=sha256:e25a2137d02573292f6ce95a0a403edbd7dde469dcf134046467a147e0875d8f

Observation fc6f9470-8e4d-42b7-8095-31dae9f2a9b2 · outbound

This paper cites Rapid solution of integral equations of classical potential theory.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining Rapid solution of integral equations of classical potential theory

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.980904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.980904Z digest=sha256:7db340244d1db6dd6f553d2036e2a8429ef0e7cd9171fc41aa9c6ec03af1af90

Observation 4f1d064c-50fd-44d3-aaa7-6a83719f9170 · outbound

This paper cites Perceiver: General perception with iterative attention.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining Perceiver: General perception with iterative attention

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.984842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.984842Z digest=sha256:548d66b2ceb013813c08ae580983eb6e3bca62562386c9f2f58ab470e3190a61

Observation 808b1e01-48b8-4b1a-9cf2-da12e7732fc6 · outbound

This paper cites Luna: Linear unified nested attention.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining Luna: Linear unified nested attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.988703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.988703Z digest=sha256:b2359fbc2f5c26c2ab70588fa5154116d79f26fcc72da88c462cc47885183da0

Observation aeabb9ea-973a-4aa3-b9c0-d3cb4871ab9b · outbound

This paper cites Colwell, and Adrian Weller.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining Colwell, and Adrian Weller

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.993029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.993029Z digest=sha256:d9597fe9188f21efc8390074ef861c30328e5290bb95c025b18d12ccf718fc8f

Observation 26d135a2-32ec-443e-8080-122b18a4c19d · outbound

This paper cites Compressive transformers for long-range sequence modelling.

Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining Compressive transformers for long-range sequence modelling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T17:59:27.996885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:59:27.996885Z digest=sha256:af86978c957ff4d2419dcccaed570af877a5c3166341bc9d8c3d1f091451ebd1

Pith citing papers

No inbound Pith citation observations are available.