Pith. sign in

Paper Citation Record · LEDGER

Rectified Sparse Attention

As of 9 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 4 inbound Pith citation observations for arXiv:2506.04108.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04108 v2

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:52:48.987113Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:06:32.693476Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:46:58.816966Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 195f9a5b-f13b-4520-a5d3-b22ff14f473a · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Rectified Sparse Attention SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.200672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.200672Z digest=sha256:f1342fa4330800a1c277f5a1f4639042d95458fa8fd6b37b77bce093e6e6699a

Observation bbacd52a-ac00-421f-b26f-472ea2bbf55b · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Rectified Sparse Attention GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.285123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.285123Z digest=sha256:79ba6b2467635ecd19b369c181f15f223e8b3880c6bf3261488bd6e43cea0e4a

Observation 16245e3d-3417-44ff-8c73-564e3e612f0b · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Rectified Sparse Attention Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.349290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.349290Z digest=sha256:4e2d22c5cc53e2af58fe01568d3cd6f8c03e31f9caf7c2b8ed34b78309b6f04c

Observation 6d71bcdd-d3e0-46c7-98e1-921e2eea93a6 · outbound

This paper cites MagicPIG: LSH Sampling for Efficient LLM Generation.

Rectified Sparse Attention MagicPIG: LSH Sampling for Efficient LLM Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.465707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.465707Z digest=sha256:73c97405a092d808beef343760ab2f9a3a7913ca31d3e39e09e4bd0606db6933

Observation 6ca1098b-5ddc-474f-a14e-ce7f8b4b2dfa · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Rectified Sparse Attention Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.527304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.527304Z digest=sha256:0e19a1f4ee8b15f617d6e622d1c6b98b1d0055caa6726843e77d3660b7e46726

Observation ed759aec-6814-40db-a56d-41151166bdf9 · outbound

This paper cites Flash-Decoding for long-context inference.https://crfm.stanford.edu/2023/10/12/flashdecoding.html, 2023.

Rectified Sparse Attention Flash-Decoding for long-context inference.https://crfm.stanford.edu/2023/10/12/flashdecoding.html, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:49.268981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:52:48.550761Z digest=sha256:e40cdcef753559cc8e4369280b84848a78417872e4643edb7d13f74096dc34ba

Observation 22abf7b4-91e0-48e1-b5b4-8b0ae3a4152c · outbound

This paper cites MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models.

Rectified Sparse Attention MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.650576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.650576Z digest=sha256:fdc41d4d3d507e2b2679d18e77122df5305cb8c3465c59a729356f88d30de32c

Observation e892edf0-045c-41f9-81e0-73dd74f82b7f · outbound

This paper cites SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs.

Rectified Sparse Attention SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.764869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.764869Z digest=sha256:e9d6ff612c3b7fc87a8e7489b3e1cd7d754bf9f0340f6b6b35960367c72e4b61

Observation 99217db7-9ac4-4fd3-91b0-9f5e57ab2e5c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Rectified Sparse Attention DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.868706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.868706Z digest=sha256:e94a54600873cbda0b8c4b5dc96b8981db100a17562273056f89dc911827d612

Observation ee8639a1-dc19-4f1b-8861-83e1711016ae · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Rectified Sparse Attention OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.901428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.901428Z digest=sha256:f43abc35a80760a2344da47235efaf9bea79bf1eb2d79f414879a52fb1a6ab2f

Observation b8c5d9b4-69ab-498a-9051-24b27898a446 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Rectified Sparse Attention Measuring mathematical problem solving with the math dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.935265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.935265Z digest=sha256:9144c8b6abc93777cc8b84a14cd808ec320961633087f2c84af31a15a1c15875

Observation 17c4b2d3-7d80-4f7e-a321-478c29739989 · outbound

This paper cites DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference.

Rectified Sparse Attention DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.938385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.938385Z digest=sha256:b8bdafa6450187f61779b5480531061c2153c719ef283a72dc2e63f3c68d7da7

Observation 12ea0f1f-3189-47b6-864f-d1ea47ef7f87 · outbound

This paper cites OpenAI o1 System Card.

Rectified Sparse Attention OpenAI o1 System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.941415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.941415Z digest=sha256:6c220513c06debf437c4a9d94eaa76ee491ebf2eeca3e0d5100f6cd922fff9e1

Observation 35a5dac5-0e4d-46d9-b136-1aecd9c73860 · outbound

This paper cites Fast inference from transformers via speculative decoding.

Rectified Sparse Attention Fast inference from transformers via speculative decoding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.944325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.944325Z digest=sha256:556c815192633d5f274537e331be63401395fda207a154937569a2e909d1286b

Observation be024e36-8e0f-44cc-909d-dd3b550312c7 · outbound

This paper cites Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857, 2022.

Rectified Sparse Attention Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857, 2022

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.947196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.947196Z digest=sha256:2905784186f4fe30157ea0a9ff023ab8b8aa067cf66417fa547ff53f282282fc

Observation 11fa7331-4345-4ad5-ac3d-8fe859d827c4 · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

Rectified Sparse Attention EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.949963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.949963Z digest=sha256:2939de3544dd3b8c83377e82fad5bd17816e773b0968e0f8a7804752334735a4

Observation 6ec6e455-732c-44d0-b336-17cb5fc7c4e6 · outbound

This paper cites MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline.

Rectified Sparse Attention MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.952928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.952928Z digest=sha256:755a94996678d6bdadaf26af50a7ca7e6793b4b88ee685df2d974738cc105c91

Observation 89b1cfc7-6b65-4532-a222-6d99636ed20f · outbound

This paper cites ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression.

Rectified Sparse Attention ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.956989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.956989Z digest=sha256:c83ebaa6e8a76de5b40c4d1fe3c2264eddcd2a254fd1c31f43f68122df011b14

Observation 131d3691-37ae-43d6-81b2-a7d6ccfa2d99 · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

Rectified Sparse Attention MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.960314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.960314Z digest=sha256:bea3d3f60bbfa30a1524b02557d236d561acedb1af561c3bec57fa1f64cf1624

Observation ff56fa26-426a-4638-807d-425e491602f4 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Rectified Sparse Attention Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.963403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.963403Z digest=sha256:0d5905599fedb0141dbe64dac6a943ca9ab07959d5e5ba818e4decbf49e4561e

Observation e9161153-b987-4d15-aee1-1ea702c8f66c · outbound

This paper cites MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding.

Rectified Sparse Attention MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.966152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.966152Z digest=sha256:dba76ec5a11842c70cd70ee41046aa06ff2bf9f4f3fd8d50c5de1c249c444767

Observation ee44efbc-ac4b-45df-b458-0d636cf0ee21 · outbound

This paper cites TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding.

Rectified Sparse Attention TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.969030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.969030Z digest=sha256:c075cd5551cc584d427b75bc52b213adc9a91a9fbf7d92463defa32bc388b9e0

Observation 4e5d68fd-4e9c-4eae-8d71-39bf16189b3d · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

Rectified Sparse Attention Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.972268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.972268Z digest=sha256:41e8ea428a3a4e25ccbca48575d245c888f376cb82c10bc0ef9531efb56f3f22

Observation 1a8181b3-5ed7-4a97-975b-447725659275 · outbound

This paper cites InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory.

Rectified Sparse Attention InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.975137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.975137Z digest=sha256:8dddb91a4ec10b0705e605d502283f166b932c932395621be39f41445cea1962

Observation 524b4230-df1a-4c66-a691-39c2f3aa839b · outbound

This paper cites Qwen2.5 Technical Report.

Rectified Sparse Attention Qwen2.5 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.978094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.978094Z digest=sha256:4fb8874f8b4fc82fdeb7bcf0b26aceaba44ebb9d9f628ef5eb94a5de5a051932

Observation afb1e160-cf1b-40f4-8c89-1d923ac34cf3 · outbound

This paper cites Qwen2.5-1M Technical Report.

Rectified Sparse Attention Qwen2.5-1M Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.981099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.981099Z digest=sha256:9a8d0acad295423691be35c356cb5b35949f1c981633254086a6d488440800a1

Observation e9a5af27-b368-4d42-817a-6675be5981b5 · outbound

This paper cites Orca: A distributed serving system for Transformer-based generative models.

Rectified Sparse Attention Orca: A distributed serving system for Transformer-based generative models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:49.239761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:52:48.984029Z digest=sha256:d4605864df2609ff1195e3635e47bf8ead776a2a6f39ec82611de559627c87df

Observation cb718aeb-3183-4d09-8cca-503448e3a0ba · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Rectified Sparse Attention Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:48.987113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:48.987113Z digest=sha256:73d95f4bb8dea0f0d0c399c65d923f91a4e5986e3eba72cf863e06c748de85fd

Pith citing papers

Observation 973f0b67-f390-49b2-b41c-433cfb715649 · inbound

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning cites this paper.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Rectified Sparse Attention

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.693476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.693476Z digest=sha256:cfc5444271bc66d18ec4e30e7ce63df2692b4d3bfae1785da2e63e339858d70a

Observation 3fd29474-9a0f-4b7a-89b8-11bf887c5093 · inbound

Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants cites this paper.

Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants Rectified Sparse Attention

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:41:30.042725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T11:40:41.762364Z digest=sha256:1dcc953643380135a1c96a2b01418cbd47825b5c42aa3132763ca15e34b849b5

Observation cf2a2230-520c-4f33-b5b1-5d835d2570f9 · inbound

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding cites this paper.

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding Rectified Sparse Attention

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.626267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T22:20:53.856657Z digest=sha256:b9cc2333392597c28ca95aba727adf7626acab86f5746909040d2aec6e9b58ab

Observation 5361cd55-4d42-455c-8d9c-b0a2d1a68d7f · inbound

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing cites this paper.

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing Rectified Sparse Attention

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:58.818402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T01:06:04.896501Z digest=sha256:13fea5adc7dca1ebd9aae0a0bc767c685ed1dc5fdaf70c83ff67f5dd2bb52417