Pith. sign in

Paper Citation Record · LEDGER

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques

As of 17 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2502.01659.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01659 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:58:14.675856Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:30:14.955831Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:30:15.054198Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ad0472da-339c-4524-80d8-8a306225c2c6 · outbound

This paper cites A Survey of Large Language Models.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques A Survey of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.552267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.552267Z digest=sha256:12cc4deb653f056a641640aa088043dc8d25a85f267fae1e438cb3d10ea2d2bb

Observation b4660cd1-02b2-46f5-910e-2475dec1acb2 · outbound

This paper cites Enh ancing molecular design efficiency: Uniting language models and ge nerative networks with genetic algorithms,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Enh ancing molecular design efficiency: Uniting language models and ge nerative networks with genetic algorithms,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:15.105507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T19:58:14.558119Z digest=sha256:c9ee58fcce68123e998f3fcd56f3599b556418ab177793052b7e9ea916387171

Observation a873b736-35bf-4859-ae49-63ca1ac51876 · outbound

This paper cites Path-bigbird: An ai-driven transformer appro ach to classi- fication of cancer pathology reports,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Path-bigbird: An ai-driven transformer appro ach to classi- fication of cancer pathology reports,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:15.091091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T19:58:14.562776Z digest=sha256:119bc674e3edea29c9f77619649f244be484616f3f0a3b5c5d682322ebf79328

Observation 55288ae4-3069-4aea-bbf5-72de2c81f4fc · outbound

This paper cites Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:15.076670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T19:58:14.567627Z digest=sha256:dac38d91ca0d600dd8e3f7afe15062d6e81f4c212a07c2fd5f64896d217d3509

Observation 482daa7e-b360-45fe-83ce-ef2c1586d498 · outbound

This paper cites Attention is all you need,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Attention is all you need,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:15.061485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T19:58:14.572624Z digest=sha256:e49bdc080b9fbe373d70d5172672bf08498184e6b5fdde071001402e794e4127

Observation fe2008fa-86f5-4635-9eb2-17a2a97781aa · outbound

This paper cites Big bird: Transformers for longer sequences,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Big bird: Transformers for longer sequences,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:15.046896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T19:58:14.577071Z digest=sha256:b59f032af81917b526c102ca94db1945a35edbfe988ba3a756a077c829501b07

Observation 111e59df-db05-4aa6-8cc8-4fb5eb861679 · outbound

This paper cites Longnet: Scaling transformers to 1,000,000,00 0 tokens,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Longnet: Scaling transformers to 1,000,000,00 0 tokens,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:15.032040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T19:58:14.582358Z digest=sha256:50fe93ab9644750f0fd06d855b9ab33b4be5177ee444f1fe8eae2d392ec46b08

Observation 519b063f-f8c2-495f-b7b6-70a578983ac2 · outbound

This paper cites Longformer: The Long-Document Transformer.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Longformer: The Long-Document Transformer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.591311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.591311Z digest=sha256:77adb0809165217ba8a9605e003725ecb5bdc6103493f72351433484d2f092b6

Observation a142f2ca-da0e-4d9e-ae0b-fc39cdfd7904 · outbound

This paper cites Scaled Dot Product Attention,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Scaled Dot Product Attention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:15.001995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T19:58:14.596125Z digest=sha256:11b768fcd2b969f257445f7b2a5e9b37c9b2c8a3c90b331e04bf3a076aeb9167

Observation 0b861f90-77c0-41b4-80be-1ac1764db6ad · outbound

This paper cites xformers: A mod- ular and hackable transformer modelling library,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques xformers: A mod- ular and hackable transformer modelling library,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:14.987469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T19:58:14.600723Z digest=sha256:a761b1071f32e4566bea406dd93415c6a0d1796310885b0ac033c8de686fffa0

Observation 07c2b4a2-cc94-449b-b72d-7f07304adabf · outbound

This paper cites Reformer: The Efficient Transformer.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Reformer: The Efficient Transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.605321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.605321Z digest=sha256:9b3ec3e8a819493bd7c625cb1f8b24f0877bef53a8f158454381ef7b41aa9063

Observation 31528233-5fe3-4462-bbfb-19c5e332320b · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Generating Long Sequences with Sparse Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.609918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.609918Z digest=sha256:d7cd0b1bab8adc8cd3f8a7a0e0e4246d3b3a95e862a0e300ffcd5ccb0ab9452e

Observation e9c68474-da92-401f-8422-e672b75014f6 · outbound

This paper cites Representing long-range context for graph neural network s with global attention,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Representing long-range context for graph neural network s with global attention,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:14.971912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T19:58:14.615306Z digest=sha256:e898818e95898661e0ab8c9577b37cfe3e0c9a3b8242a185256717e4a1671131

Observation f5b02652-6964-4fe2-9172-75aad3881ef4 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.619889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.619889Z digest=sha256:be52166e333d85adf3248f20be62302c55f4cb1b0399da75f8a2fc6e57bb4157

Observation 4f491a2e-409f-4af4-85d1-321a611e4390 · outbound

This paper cites Blockwise parallel transformers for large context models,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Blockwise parallel transformers for large context models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:14.955929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T19:58:14.624919Z digest=sha256:e9fc708b8bf6101491c619b1050fd8246662edebf79686129bcd757e34c06fe6

Observation d0c6718a-c934-411e-97ea-65eab7e4a81a · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.629276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.629276Z digest=sha256:79a7621536e96e7f5643cfcb8e49d33f07e0e585fedcaf2bfa3b6fde273430ca

Observation 3e7a8ee8-796c-4045-a9d0-b86c1b4377c4 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.634125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.634125Z digest=sha256:1103ce8fca2131300d127efa3467d8ba49e2bce37b5e3654920235a9ac0325bd

Observation dd7d6fc3-1fca-4267-9285-c34428813930 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.638827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.638827Z digest=sha256:475be634700aef0f679eb7d8c9de0d37007a80ffe8a030aa31f1e975fb3b6336

Observation b7c3e28b-d0c1-4d88-a2a9-25be97bb4c89 · outbound

This paper cites Flashattention-2: Faster attention with bett er parallelism and work partitioning,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Flashattention-2: Faster attention with bett er parallelism and work partitioning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:14.941376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T19:58:14.643740Z digest=sha256:1e2c4b2208881d127b47fd276a56c904020b38fa4a585c45bbc45c15674768d4

Observation 5094d809-8176-4923-9b10-210a5d1406c7 · outbound

This paper cites Flashattention-3: Fast and accurate attention with async hrony and low-precision,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Flashattention-3: Fast and accurate attention with async hrony and low-precision,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:14.926557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T19:58:14.648012Z digest=sha256:9be3dc32a1da0dc92a00bfd0861f38e055a2bf1d6918644ebc5d2d8070ac4ede

Observation 15a5d1cf-8792-4de6-a79b-b2838f604f5c · outbound

This paper cites Faster Causal Attention Over Large Sequences Through Sparse Flash Attention.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Faster Causal Attention Over Large Sequences Through Sparse Flash Attention

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.652472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.652472Z digest=sha256:c6e80bf7ddc2772bbd97db4bd176bb6599f38494f21ab375dd9b6d42578258ed

Observation 708e9999-3208-4809-80aa-f22aaee6f15d · outbound

This paper cites Efficiently Dispatching Flash Attention For Partially Filled Attention Masks.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Efficiently Dispatching Flash Attention For Partially Filled Attention Masks

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:58:14.748858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T19:58:14.657264Z digest=sha256:5b996f928e481e21fe54fb22114ed1c2237d3d89ddb889adb754297ef21a0088

Observation fb76bb74-b7da-4650-8099-a2379f0770de · outbound

This paper cites Online normalizer calculation for softmax.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Online normalizer calculation for softmax

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.661834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.661834Z digest=sha256:6d5569e88a5f7e721b0d9e019d05199bc28d247a6bdec43ed8b1da26d43faae9

Observation f45fad95-f6b0-4952-b127-0159a1d532f5 · outbound

This paper cites On the power of some pram models,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques On the power of some pram models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:14.910917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T19:58:14.666325Z digest=sha256:b46140ff8ef6c3725b9f59cc766e4bb3e975ac58556c37427d5470ba3a9d2a1a

Observation 6b0649cd-c7c1-4705-b0cd-421d6f60bb44 · outbound

This paper cites The Llama 3 Herd of Models.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques The Llama 3 Herd of Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.670570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.670570Z digest=sha256:4d470189bf6499613d259c0be05b8f004215b052c81400266bea7b106b78aa19

Observation 57219432-3d10-46c5-8c68-3d6832828071 · outbound

This paper cites Algorithm 10xx: Suitesparse:graphblas: Graph algorithms in the language of sparse linear algebra,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Algorithm 10xx: Suitesparse:graphblas: Graph algorithms in the language of sparse linear algebra,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:14.895370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T19:58:14.675856Z digest=sha256:0e156b2564540307eb57b8f8f343e179ce5e70573b28a048e5092dd3a20b5997

Observation b0074eb4-cd9c-4d93-a543-bcf4e14d4178 · outbound

This paper cites Available: https://arxiv.org/abs/2307.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Available: https://arxiv.org/abs/2307

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:15.016939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-09T19:58:14.586742Z digest=sha256:3a112054b58b082830752ae1c7cbb991e0cc07659c0643703c1974b0fbed1a49

Pith citing papers

Observation 68793c4a-d8bc-43c8-b927-4c223eb6b933 · inbound

Sparse Fine-Tuning of Transformers for Generative Tasks cites this paper.

Sparse Fine-Tuning of Transformers for Generative Tasks Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:30:15.058609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T17:30:14.955831Z digest=sha256:20be7a8051c7cc99efdace01613ece8aa23d3e7f84d437263df57b46c6d004fb