Pith. sign in

Paper Citation Record · LEDGER

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)

As of 8 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2507.08637.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08637 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:19:59.124623Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fd8dd414-dd24-4ab0-97ee-9560488ff616 · outbound

This paper cites an unresolved cited work.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:19:59.425992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:59.120424Z digest=sha256:da4b12fd1c36417478e370730e3f357eea34bfac1595bb7e4a0f089fe2419d79

Observation a081d3e3-d084-4854-b1c2-170057c786fb · outbound

This paper cites On the Use of ArXiv as a Dataset.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) On the Use of ArXiv as a Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.002779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.002779Z digest=sha256:cd92dc9e0fc90c7a82cc9220e02f3389a06ff8d27ec612ffa42ec5ce4a24ba33

Observation 109a4e65-e163-46fb-aa3f-825e0a6d4ab8 · outbound

This paper cites Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.026663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.026663Z digest=sha256:654635b20fdb849c18a04df027669eaff8935520db674aeb031b4e8c10de28dd

Observation 7c854729-b3e7-41cf-982e-9683292a4243 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.030523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.030523Z digest=sha256:2331bd4653b345693e76f23cb30d5edcaefb03d6ff4024977817d1ff6871acb8

Observation 5e5daeec-0946-4337-a372-26d9a9fed0fc · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.039722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.039722Z digest=sha256:2d6b675cd02feffbc9d937ed5e9e0e32d3c537331f06005df8e5b26ea14206a0

Observation c5daa5ea-e5a2-4432-8753-c109029e39dd · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Efficiently Modeling Long Sequences with Structured State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.055602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.055602Z digest=sha256:6803591734113d5655c3a21cc9766373b051bfc3844d78d2868009b65b2c0a18

Observation 2ba0c6f1-76da-44b3-b17b-a3107a7cd05a · outbound

This paper cites Mega: Moving Average Equipped Gated Attention.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Mega: Moving Average Equipped Gated Attention

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.066945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.066945Z digest=sha256:b8eaf7ced0418da15575dc538d254d615692d0a01dea0987ab409f7d8204b251

Observation 4682cb19-479f-4f81-9aa7-bd7cfd362be3 · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) RWKV: Reinventing RNNs for the Transformer Era

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.080608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.080608Z digest=sha256:9d1ddcf17d0d470153daf8c104b49e41a06b5f09ae237f69d05d3b9634dfc07c

Observation 892858da-22bc-4cfa-b2b6-4f01e6cb0adf · outbound

This paper cites Random Feature Attention.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Random Feature Attention

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.085521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.085521Z digest=sha256:ea14a228b70c03133b51c8152c2e9ab63f8dbc5e999bfe4eef5886ae3296a50f

Observation 8da21433-da9a-4d10-a229-6a4651f9ca91 · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Compressive Transformers for Long-Range Sequence Modelling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.088847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.088847Z digest=sha256:e4aca2bdc603b6179ee0cbfb3406fd84c3281737097d533cbb354136f766b0b1

Observation 5c7cea95-11b6-462b-8566-4de0ae71d0e9 · outbound

This paper cites Attention Is All You Need.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Attention Is All You Need

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.102425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.102425Z digest=sha256:9e1ca633c5b54ce2482602b5b27a4e4d214a68d4d94ff719efe0b9cad67d6ece

Observation 42f71596-37b0-4ead-bc7a-d3545bbf0d89 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Linformer: Self-Attention with Linear Complexity

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.106177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.106177Z digest=sha256:94ecd29145676613b66a48560a8cc8e4c3039bdcf42667f501a98f5f80b1e5ab

Observation bf4b19f1-f964-4045-9576-875964532731 · outbound

This paper cites Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.454835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:59.112324Z digest=sha256:6c379000177c401685fa26b41538a9e94c013fa220ae71c5ce7b3f534dd869ae

Observation 92a9b74e-4d6c-4b38-b474-fc4fc78b8561 · outbound

This paper cites an unresolved cited work.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:19:59.443101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:59.116208Z digest=sha256:4107d56296bbf0cfaa4c2ebb38d09eb0589ac1cb89496cc0588ca48931e0e605

Observation e9d6b606-b5d2-4735-880d-d1ec74f053fb · outbound

This paper cites Here, R ∈ Rd×m has independent and identically distributed Gaussian entries, E ϕ(q)T ϕ(k) − κ(q, k) ≤ C√m , where κ refers to the softmax kernel.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Here, R ∈ Rd×m has independent and identically distributed Gaussian entries, E ϕ(q)T ϕ(k) − κ(q, k) ≤ C√m , where κ refers to the softmax kernel

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.413911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:59.124623Z digest=sha256:388ebfed9469e19b745e57e8139708e98b1580a52f58423f9bfd39f89b975130

Observation 3d6ca54a-980e-4ec0-b14d-1495f126d14d · outbound

This paper cites an unresolved cited work.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Unresolved cited work

Reference 1999

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:19:59.467682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:59.075723Z digest=sha256:420139d6104fd522e2b86fd488f04663487fe171c3a375373150ca273f17f143

Observation cf960b86-006c-4967-81c9-d1efade079cf · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Generating Long Sequences with Sparse Transformers

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.997598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.997598Z digest=sha256:6befcee2896e388b70d831ef0b87f3daad67402fd14229871fd8e800581fbbe1

Observation 1cde1ea8-d033-463a-a18e-5d66e60c7560 · outbound

This paper cites Longformer: The Long-Document Transformer.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Longformer: The Long-Document Transformer

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.989559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.989559Z digest=sha256:d60fd6260a03183fc1fc0fe9cbb9fbb1b05d3fea6bb150879c175be55d31ae74

Observation ec2fe666-c143-432c-9d6c-f682d7b3ee8b · outbound

This paper cites FNet: Mixing Tokens with Fourier Transforms.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) FNet: Mixing Tokens with Fourier Transforms

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.062242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.062242Z digest=sha256:f3c1a77330b4b1691f535c4b917600150a77e6d0d626d4c850980f8a1bf96529

Observation 14bbb4a8-30e5-4edf-bad3-72df948be2a4 · outbound

This paper cites Peter L Bartlett and Shahar Mendelson.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Peter L Bartlett and Shahar Mendelson

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.512551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.976895Z digest=sha256:a51da609e147f972c6c0f6ac12f048cafd8043a1b2d5135e8196adf8d066e9a1

Observation 9ae1559d-846f-472c-85b1-5fd428862434 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.489106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:59.035017Z digest=sha256:ff8c2debdf29fc61bb37ecebd3f6deff3c2ce42e6e7ae4c7c55fe53f688b0282

Observation 45f548ce-8e1f-4ee9-8e50-e8e145103771 · outbound

This paper cites On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning.

Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA) On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.044342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.044342Z digest=sha256:139b337c002f62aa2ba534a2399b3b5d1b9c0504e71305c98a5bb4e1db3174c6

Pith citing papers

No inbound Pith citation observations are available.