Pith. sign in

Paper Citation Record · LEDGER

Accelerating Prefilling via Decoding-time Contribution Sparsity

As of 5 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 3 inbound Pith citation observations for arXiv:2507.21526.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21526 v4

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T03:05:34.843274Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T06:11:23.742632Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-30T06:14:19.020219Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch11

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6752a959-e23b-4e7f-9a6c-9c3867d5517a · outbound

This paper cites GPT-4 Technical Report.

Accelerating Prefilling via Decoding-time Contribution Sparsity GPT-4 Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T03:06:59.982748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:0b6563d6b8090622a4b8eb3025154091b469bfc8e3202b5e6e7f8e07b06ddec0

Observation 8edc1096-af4f-4018-805d-dd8a33972296 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Accelerating Prefilling via Decoding-time Contribution Sparsity GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T03:06:59.938352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:2dd3ac35054545eaf83acde9f4362612e27c967e310bad86f195f615aa744a75

Observation 2ba95874-c273-4c76-8393-dbff51ef72a6 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

Accelerating Prefilling via Decoding-time Contribution Sparsity LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T03:06:59.943448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:5b03ccf1d30e089444d3b8937c2bfaff68fcf7b0678aeb27ee25c1af1fd9be8c

Observation b6b17967-ac78-4f27-bdf2-2605be464210 · outbound

This paper cites Longformer: The Long-Document Transformer.

Accelerating Prefilling via Decoding-time Contribution Sparsity Longformer: The Long-Document Transformer

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T03:06:59.953216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:aeb74912a4ec0a95c14fdc1da4ef2c8ad366eac9412c99ecb018f0a5f9fda26a

Observation 5f4fce76-b8a9-472e-86e9-0ff971984bce · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Accelerating Prefilling via Decoding-time Contribution Sparsity Generating Long Sequences with Sparse Transformers

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T03:06:59.957968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:01fa192be3ad58154c8c43fdc5998a0e8684fff9102a0b4fa6fefa6f2b6d0e81

Observation bc1e33a6-66c6-4201-be3f-bbe78c148eb5 · outbound

This paper cites LongNet: Scaling Transformers to 1,000,000,000 Tokens.

Accelerating Prefilling via Decoding-time Contribution Sparsity LongNet: Scaling Transformers to 1,000,000,000 Tokens

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:06:59.963072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:f1ed87a76dd0bc28f04b5a98d50c49757e0ce61ae68e588028924b21c61e1866

Observation b2944ae0-d9ca-4c94-bc02-ec369b85ea0b · outbound

This paper cites The Llama 3 Herd of Models.

Accelerating Prefilling via Decoding-time Contribution Sparsity The Llama 3 Herd of Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T03:06:59.899550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:3ff74e4df0c27e8f45497274f32879c25e94bfa654823b45ad738f314e28bf5b

Observation 013ae64f-7454-4a2a-b689-fba98dda6e35 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Accelerating Prefilling via Decoding-time Contribution Sparsity RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:06:59.967683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:45273610a4cf2975ec4468522f581c8d9d2b82fe33f03d2d1818d9e6c486b492

Observation 06440acb-70e4-45ec-9322-ffe2279fb5aa · outbound

This paper cites Mistral 7B.

Accelerating Prefilling via Decoding-time Contribution Sparsity Mistral 7B

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T03:06:59.977245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:bfe29d0120723f785c45277451fbf1481d46f9a83500ac75181e6fb407d81429

Observation fe522093-4b58-4108-9d8d-c7eacb0b20b3 · outbound

This paper cites FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference.

Accelerating Prefilling via Decoding-time Contribution Sparsity FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:06:59.904965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:a87904beb4dfe57d8063166285c75e6c6479f800b7a5779139cff09eada859e1

Observation fbdc757e-8c42-4657-be80-d28b6511d4f0 · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

Accelerating Prefilling via Decoding-time Contribution Sparsity YaRN: Efficient Context Window Extension of Large Language Models

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T03:06:59.932202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:3a7938dde55dbc82f069f42d540e4ef2b0bd9ac437b6dbb6647427d6c1cb26f8

Observation 3d544be1-1534-4c55-8221-f7f6228959eb · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

Accelerating Prefilling via Decoding-time Contribution Sparsity Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T03:06:59.927072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:4aa79e412022a25ac1fad8144f9bd3c4b149e90f87e8def84c52a0350cddb038

Observation fc6f7531-f91e-4297-82df-bdd8a9632fa4 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

Accelerating Prefilling via Decoding-time Contribution Sparsity DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T03:06:59.910968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:d6eb7e3112fec6e5cccc48f6496f32a7bb480b82ceb7663a47195e04b0746764

Observation 52568b56-e0b0-45a4-8a27-c9a2417962d7 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Accelerating Prefilling via Decoding-time Contribution Sparsity Efficient Streaming Language Models with Attention Sinks

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T03:06:59.916469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:d2921c9f0204e57d686e3659932798869d482f6e40fd89a961e27a7631c68b85

Observation 74f1b448-573d-466e-8df6-ad20893a2a52 · outbound

This paper cites XAttention: Block Sparse Attention with Antidiagonal Scoring.

Accelerating Prefilling via Decoding-time Contribution Sparsity XAttention: Block Sparse Attention with Antidiagonal Scoring

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:06:59.948278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:be0300d3d60a148087e03e5dc7e7ffda8f810c55a1dce5477b21518e28664656

Observation 917f7ec0-2177-4215-ba76-cffb91f8fd40 · outbound

This paper cites Qwen2.5 Technical Report.

Accelerating Prefilling via Decoding-time Contribution Sparsity Qwen2.5 Technical Report

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:06:59.922217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:f748ef9d57dea37f29fa6ec396b6eda9c7da300646941b4a24f42997a850377a

Observation 1672b298-ddfa-48c1-8d1e-88ae6677fa66 · outbound

This paper cites Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely.

Accelerating Prefilling via Decoding-time Contribution Sparsity Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:06:59.972943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:05:34.843274Z digest=sha256:e3c7f3187d0418e89e1b2973af494270a6ac557ff9eee8e7f49a01746f7ed6a1

Pith citing papers

Observation 71687192-9038-4f65-a634-37452d4c4141 · inbound

S2O: Early Stopping for Sparse Attention via Online Permutation cites this paper.

S2O: Early Stopping for Sparse Attention via Online Permutation Accelerating Prefilling via Decoding-time Contribution Sparsity

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:36:32.848517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:32:52.948154Z digest=sha256:dce9d6e6b0c026d20549c07dbfdf08854fa36730fe8352dbbe4d7203ef7111b1

Observation 5a3aa25d-3f69-4882-8886-3b6d2f0c7957 · inbound

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation cites this paper.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Accelerating Prefilling via Decoding-time Contribution Sparsity

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:05:58.042049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:3df9ee21e8095403fb0892c82835504d159507929bfaca1a2c2ba57b8e493178

Observation f3f8182f-2bde-423a-92d5-304f40faf09e · inbound

MATCH: Modulating Attention via In-Context Retrieval for Long-Context Transformers cites this paper.

MATCH: Modulating Attention via In-Context Retrieval for Long-Context Transformers Accelerating Prefilling via Decoding-time Contribution Sparsity

Reference 79

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T06:14:19.022993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T06:11:23.742632Z digest=sha256:8738b8bab080495502f8b25adc6b547c45bd1a15c8e1abb5114c51e4d2c66fb0