Pith. sign in

Paper Citation Record · LEDGER

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling

As of 7 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2505.24179.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24179 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:39:10.193536Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f1181232-b07a-4c73-a2fe-3033e696e651 · outbound

This paper cites Booksum: A collection of datasets for long-form narrative summarization,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Booksum: A collection of datasets for long-form narrative summarization,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:27.136294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:01.544466Z digest=sha256:1ba7857f008b4b1dda9a92f65853d4cd573be4c3e349c718e0d6d6a2e3983722

Observation b13ab18e-98c6-42c3-beed-4206e077c3af · outbound

This paper cites Transformer Based Implementation for Automatic Book Summarization.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Transformer Based Implementation for Automatic Book Summarization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:39:12.172785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:01.698014Z digest=sha256:400ca9fc7445d423851e81eea9d6fdfdcbc18ccbc7c85ee163cd4df3cf63f800

Observation 512c5be2-fb3e-47e2-8561-02f765bf403d · outbound

This paper cites Booookscore: A systematic exploration of book-length summarization in the era of LLMs,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Booookscore: A systematic exploration of book-length summarization in the era of LLMs,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:26.814959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:01.862280Z digest=sha256:1f204393a9a43091014f91a2df8a24db1262f9a1a0cf23e70ff875cafe32caf5

Observation 8d871b43-4e3d-458f-adad-b036b31dd6f2 · outbound

This paper cites Peek across: Improving multi-document modeling via cross-document question-answering,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Peek across: Improving multi-document modeling via cross-document question-answering,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:26.607039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:01.995486Z digest=sha256:d16df2ac6eef68d2090ef475a07b7cc0bf260d543b36224a1b3edab6a6907ddb

Observation 0492a883-5869-4d00-b33a-23eaf3cf4da0 · outbound

This paper cites Quality: Question answering with long input texts, yes!,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Quality: Question answering with long input texts, yes!,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:26.360575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:02.119346Z digest=sha256:a46c4055ca9b6de51aeaa7664052047b64a3486b09de347e9e8d164a6b57e3b3

Observation 51bd3625-2748-411d-8673-c8a17d439cb9 · outbound

This paper cites Eli5: Long form question answering,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Eli5: Long form question answering,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:26.148743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:02.225120Z digest=sha256:c49b3b2c8653f9845d6b0c6600f2dd1c558cbd298718b34afedc8b668098a1c1

Observation 11e139f6-855e-441e-9eb1-dee98a1dcd55 · outbound

This paper cites Teaching code llms to use autocompletion tools in repository-level code generation,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Teaching code llms to use autocompletion tools in repository-level code generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:25.930551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:02.354328Z digest=sha256:29e92137bb94f88e92a8ac10f29ae256fee8a92c61d9d50687539f8cf67b5f3f

Observation d394e65c-6b84-4b07-85e6-84e86266e87e · outbound

This paper cites Rlcoder: Reinforcement learning for repository-level code completion,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Rlcoder: Reinforcement learning for repository-level code completion,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:23.459473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:02.439280Z digest=sha256:290990c274cec70e92294783f94cec9cd978fdb502588e676891d4f5f2cf8763

Observation cd7e2f28-2281-477c-ae9e-67a4033939f2 · outbound

This paper cites The llama 3 herd of models,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling The llama 3 herd of models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:21.967928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:02.540264Z digest=sha256:64225a1a1a6f5b6e64b7c63b835b45333a6e05028d51d38b1abde507bd48f543

Observation 3dbd470f-d605-4360-a5f6-439816f638b5 · outbound

This paper cites Qwen2.5-1M Technical Report.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Qwen2.5-1M Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:02.628737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:02.628737Z digest=sha256:055a185ebffbba3c708e5bf78bd0eaeaae7317f2a40611262b30423bc4ce2c09

Observation 7024eba2-1cc9-4e41-8937-4623bc89b760 · outbound

This paper cites Gemma 3 technical report,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Gemma 3 technical report,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:18.840840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:02.777619Z digest=sha256:894d2786475129fe2099dfa443b10718fe594087e79cd3da0324d17744532d3e

Observation e6c144f1-0568-490a-8e08-926250a7c254 · outbound

This paper cites Deepseek-v3 technical report,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Deepseek-v3 technical report,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:18.572954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:02.883538Z digest=sha256:ecc08f3fef1613ac6d5532c01efc70fb88f2eaa3d1d048e3962a943a1a87835f

Observation 3ebd4d0b-4c23-4611-a066-be719b0961b9 · outbound

This paper cites Attention is all you need,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Attention is all you need,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:02.961058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:02.961058Z digest=sha256:2d3a4f713a66b40864f8c18c2e67368eb94fcf7838e96b4c9eae3002577c76c8

Observation e8d31f66-3863-43e5-8cba-31d97e3a4ff1 · outbound

This paper cites Challenges in deploying long-context transformers: A theoretical peak performance analysis,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Challenges in deploying long-context transformers: A theoretical peak performance analysis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:18.317295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:03.066855Z digest=sha256:93f1acb1c6bdfda9161acfd94984949a137ea6ed44aa4de5511768c3c75c67df

Observation dc30b61e-5109-45e4-afa2-82763aa26a9e · outbound

This paper cites Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:18.087881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:03.195707Z digest=sha256:ffc407450a760a54f5ed16693aef1fedbb31d4006b8ee3abecfea4172fbf2616

Observation 9b8de900-60eb-4bab-8ac4-71d127049db3 · outbound

This paper cites How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:03.324221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:03.324221Z digest=sha256:b99f1988afbb2b1393e320e16da95bccccb2044cae13977868b0f9662802250f

Observation f1fbed45-88d9-46ac-9bde-e090760fa7d3 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Generating Long Sequences with Sparse Transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:03.429468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:03.429468Z digest=sha256:8083e3efab0e2775ba129c12236f7186155980d9d15488e6e868e88d96e8b1e2

Observation 3fb3dcee-4063-43f4-9f7a-9f4a119ed900 · outbound

This paper cites Big bird: Transformers for longer sequences,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Big bird: Transformers for longer sequences,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:17.709833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:03.569330Z digest=sha256:3dd8b29e0461b6a1899a9c8a012e5cd38bbe7a79f4d1051ac0993fc002d3cc9e

Observation 1077f3fa-2255-4a52-aecf-c71753c526b4 · outbound

This paper cites Longformer: The Long-Document Transformer.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Longformer: The Long-Document Transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:03.743592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:03.743592Z digest=sha256:7bed3f5104e07bb2dc9bc0060db18a4edf44da79662e19d730f5cdf08dd95f98

Observation dd0ba8b8-d185-4029-98eb-6479edf9cd5d · outbound

This paper cites Efficient streaming language models with attention sinks,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Efficient streaming language models with attention sinks,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:17.413962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:03.876438Z digest=sha256:7ddd53d94f34cb508d496dec4eb31c688b3aaa440dc534c05c48f4dcfd5f143b

Observation 9c1f97f4-31be-47bf-a7e4-89aa14373d90 · outbound

This paper cites LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:03.996288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:03.996288Z digest=sha256:46135c0b0835d87b06360704140139bfb6716d679a4e6206841528e1147fe8da

Observation 880c0f54-d757-48f2-b888-07afa42268bf · outbound

This paper cites SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:04.130346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:04.130346Z digest=sha256:9327af75c214a7ae5a3a0e361076df15ef8a34803f7a4112682a1a04cf8051b8

Observation 23c86693-7505-4dfc-ab59-2ab4c6e8a005 · outbound

This paper cites Flexprefill: A context-aware sparse attention mechanism for ef- ficient long-sequence inference,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Flexprefill: A context-aware sparse attention mechanism for ef- ficient long-sequence inference,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:17.203121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:04.292807Z digest=sha256:ca08d092597955b3eaa8b256c228f51983c26d176e4aff0e6290e68b89f8282d

Observation c91011ad-266a-48c1-a33f-3149672e4140 · outbound

This paper cites Spargeattn: Accurate sparse attention accelerating any model inference,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Spargeattn: Accurate sparse attention accelerating any model inference,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:04.457185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:04.457185Z digest=sha256:82f8517d4a1adbf91e8f3f1774c50eb4342c0162143f59138ca9d1a705815b3b

Observation c552fb85-094a-49c9-ab38-0dc97356c943 · outbound

This paper cites A training-free sub-quadratic cost transformer model serving framework with hierarchically pruned attention,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling A training-free sub-quadratic cost transformer model serving framework with hierarchically pruned attention,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:16.819110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:04.564148Z digest=sha256:667bd2ed18de23f7539f92a5912de6ce7ba7e69cc590759edb7fc3139887fcab

Observation 4d02b2b5-2c14-4ecb-9d69-11e44e8a1782 · outbound

This paper cites When attention sink emerges in language models: An empirical view,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling When attention sink emerges in language models: An empirical view,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:16.468229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:04.678368Z digest=sha256:9e406a12cc8a84be0a4eb21196c2eca87b67e8caee424bbabea2925202d56f2a

Observation 777fd1ba-d746-48d0-8058-35432def5b16 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling H2o: Heavy-hitter oracle for efficient generative inference of large language models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:16.209941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:04.867954Z digest=sha256:4f021ceea3a6fc998e3b44b96679f2db49cd35bdfe0cd08404b4537f3dc25f44

Observation 0b2391d7-3bc3-4e6a-9457-1abb8c98786d · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Snapkv: Llm knows what you are looking for before generation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:15.903699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:05.026810Z digest=sha256:406900a5423f683806339ff64e93e0d7ae1181c1a4db305d230481d39ce62258

Observation f0d3eeb6-e310-4965-99c8-b85aa2d41096 · outbound

This paper cites PQCache: Product Quantization-based KVCache for Long Context LLM Inference.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.152541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.152541Z digest=sha256:13f84c0d400ebbf8e920e739b738797db90aca26a1f9098f6a7bcc57bac59be9

Observation 68f5277a-7d78-4afb-813e-57c80128c8d0 · outbound

This paper cites Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.263720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.263720Z digest=sha256:a98a9c9001cbcabef3525345fb92b17c2894d04cae7484b7a653900687c658d9

Observation d099f893-9734-41ba-994f-43a4a313878a · outbound

This paper cites Qwen2.5 Technical Report.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Qwen2.5 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.385233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.385233Z digest=sha256:282d06a3922ad21c5a878fb5f2d6a17406a863144ac0bbe7deccaba48dfd069e

Observation 8fbdeb21-5945-4ade-a74b-dfe82e553374 · outbound

This paper cites Characterizing Prompt Compression Methods for Long Context Inference.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Characterizing Prompt Compression Methods for Long Context Inference

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.504381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.504381Z digest=sha256:182f1b821a96d943ed2ecdb1c2c961af212f2a64a42ebfb832e994975907c3a2

Observation 0e17897b-9f3b-4af4-bd78-0b15be0fe53d · outbound

This paper cites Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.601985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.601985Z digest=sha256:73f8b032304faeffe4e28b02dbe86bbf71236f0f1b292aaf195302a15a2b310c

Observation 7fb1ec33-d153-40f9-9093-18fd3a040d60 · outbound

This paper cites LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.729533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.729533Z digest=sha256:97671be4391dbfc99cb15a92d5e59ecdedec50679b0c7e670732da00771f4603

Observation 8677bee6-8834-4439-a1f5-446277b5db5b · outbound

This paper cites Compressing Context to Enhance Inference Efficiency of Large Language Models.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Compressing Context to Enhance Inference Efficiency of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.895555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.895555Z digest=sha256:f118d6db306394a03304786c3234b22ecc611445c6b6cd86fa3f492271a4e03a

Observation 54ad1def-52a4-41d4-9065-0738cec69d22 · outbound

This paper cites KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:06.015867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:06.015867Z digest=sha256:bf94dee80e57b764a1152284e0d8073284367f0c2da80d4a6fdc060a1d5e80c6

Observation 73c9d2ff-0379-4a51-80bc-4a1411f3cd37 · outbound

This paper cites Moa: Mix- ture of sparse attention for automatic large language model compression,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Moa: Mix- ture of sparse attention for automatic large language model compression,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:06.153339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:06.153339Z digest=sha256:29e8fb8b136a9e4588855c1258e85962876ba1742e691ad40ee224caf935b2ef

Observation a9985a76-0204-4317-95d5-3abf1cd66568 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:06.317237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:06.317237Z digest=sha256:67479cf1b2d76f6e7e6b2c388d32a3e435d028d6af08468576db95311231571b

Observation 783534c0-5b4b-49b8-bff0-f2b48d7627f8 · outbound

This paper cites SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:06.471625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:06.471625Z digest=sha256:54a9f0931f630709e657cc1ff2c76e67f9c4eb09b25b29aea5300ae49341f057

Observation cb1b6144-7239-4bab-a852-bd4d7f9f8a3e · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:06.633344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:06.633344Z digest=sha256:840adb258b91943762edd7888a3a2cbe48ff6fdf8f4f40cca370d21966908576

Observation 52639c41-fccf-42ba-a00f-b3047216356f · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:06.783759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:06.783759Z digest=sha256:67943a7ad5ac6ce6314c99446733ceb56b29ed3eb3003943e464dec24df1c9b5

Observation 122136bc-932d-4d23-a356-0c9231bfa3bc · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling RWKV: Reinventing RNNs for the Transformer Era

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:06.925763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:06.925763Z digest=sha256:5339ec79939e984e8303cbeb6b93735e829d6cfa47fe81ef9aad09fc0b11352f

Observation 9ffea3ab-fd1e-43dd-ac41-cd4805e84c8e · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:07.074760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:07.074760Z digest=sha256:e9bdca45b70e42bc871da55bc817a04da433d9c7fc6a876fc45d0d0281ce2d5a

Observation 072a6031-85e0-4157-9026-1e82c3954bfd · outbound

This paper cites Mamba: Linear-time sequence modeling with selective state spaces,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Mamba: Linear-time sequence modeling with selective state spaces,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:15.549430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:07.195393Z digest=sha256:8116b9f1c9d7d1896886671533e492e69e547949d4e46de94c8842440a690f8c

Observation 656dcb64-0540-4ede-b338-6e62a40d085d · outbound

This paper cites Transformers are ssms: generalized models and efficient algorithms through structured state space duality,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Transformers are ssms: generalized models and efficient algorithms through structured state space duality,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:15.233636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:07.353698Z digest=sha256:cba8f5c874d5cc37abb2294ba414d3b8fc2ce0c68e005d96c2ed269e6088be1f

Observation 8094898b-a09f-4271-ac6e-1e47a32bf967 · outbound

This paper cites SparQ Attention: Bandwidth-Efficient LLM Inference.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:07.482431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:07.482431Z digest=sha256:f87598b95b039afcb6dd01d32617c677ca081268aab2151b9d1d2a14263feca7

Observation c1ede19b-d24e-4f37-9da5-4551d0f4b7eb · outbound

This paper cites {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache management,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache management,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:14.895773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:07.591420Z digest=sha256:45ec977e68ed77b1c9cd4268e644b1cb79ab19362b31e8f21eb2c6a9364cf30f

Observation c8b8f53b-8397-443e-8325-0edc21cec549 · outbound

This paper cites MagicPIG: LSH sampling for efficient LLM generation,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling MagicPIG: LSH sampling for efficient LLM generation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:14.691354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:07.803444Z digest=sha256:707ea34724e9f845cfc945bedd670970b03f901db7039d6bbbdb44d690d7483a

Observation 421f902d-d766-48e5-9a7c-159250df986e · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:07.925816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:07.925816Z digest=sha256:85beef808c747ab136714cd1aaff2e5ec325f99a599b1470b404f5a0617feee9

Observation 05e50e07-b88e-4883-b6a4-fc715debffd4 · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:14.514745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:08.057556Z digest=sha256:7830e90171184745b3a3e29ece103ca00ba4ed77e952280b82409837f80bdfcc

Observation 815689dc-7de3-4398-8a7d-9d441cf63646 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:08.191015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:08.191015Z digest=sha256:8b29dffff2dbf959fa5d2057455a9ea2767555adbba8458e228db894bacfe5b3

Observation dbc3a49e-e94f-44ed-96cd-72c9db4abe33 · outbound

This paper cites A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:08.367652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:08.367652Z digest=sha256:4883dfe66a9e679a664a6a01c5848298fa13e212a56f39afc47dc0975013f4d8

Observation 7771ceee-8c80-42c4-ab8e-5e1c995482d3 · outbound

This paper cites Cam: Cache merging for memory-efficient llms inference,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Cam: Cache merging for memory-efficient llms inference,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:14.395429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:08.545492Z digest=sha256:f8653ac932f945c5e736c51305e30eb8350aae1e4b7bc84a8e067643b983b2ff

Observation cc45d404-06d6-4b1f-b8bb-d6a0af867bdd · outbound

This paper cites SubGen: Token Generation in Sublinear Time and Memory.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling SubGen: Token Generation in Sublinear Time and Memory

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:08.668655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:08.668655Z digest=sha256:9ba1926d173b819592612dc41fe9e778ee91eb69fc7c49b4a1d7359f273a3c91

Observation 7d3e88fc-7553-4943-bd0e-8208da72cbd5 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Flashattention: Fast and memory-efficient exact attention with io-awareness,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:14.199089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:08.838348Z digest=sha256:1915101798434951ad19dc5a9983a355ed8bf1ce5fd23cf7a25f850dfad61055

Observation cc81e05c-8df2-4ac4-9c0f-1c127d93a221 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:08.999098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:08.999098Z digest=sha256:d03d687516be03e564d9441e5640bf77decccd5936e1dcf8f45d87e715bd9b0d

Observation 0a37b6a5-5f58-4758-9f2b-8da82c070851 · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Flashattention-3: Fast and accurate attention with asynchrony and low-precision,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:13.945945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:09.165987Z digest=sha256:c1b492cc087f713e29018a3bd09f479de8e94f6917abb525f92fad35a024a563

Observation d15aa6d2-c9ba-4662-8b4f-9464d76ad745 · outbound

This paper cites Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:09.239572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:09.239572Z digest=sha256:4ad51db6f7e724cdf85d90965f5cdc3cf9794130b02b1dccd3d2ccc1bf0b67bd

Observation b210d8f7-8e27-4561-b883-0e2d0f162c3c · outbound

This paper cites Sageattention2 technical report: Accurate 4 bit attention for plug-and-play inference acceleration,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Sageattention2 technical report: Accurate 4 bit attention for plug-and-play inference acceleration,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:09.376740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:09.376740Z digest=sha256:732e11bc7938844f8e95e7ba8555a7de84fb2ec2900bd9cb290c208a31442c5c

Observation fb3fe215-278a-4b2a-a516-3d74a2c82088 · outbound

This paper cites Sageattention: Accurate 8-bit attention for plug-and-play inference acceleration,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Sageattention: Accurate 8-bit attention for plug-and-play inference acceleration,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:13.794860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:09.493235Z digest=sha256:59f4449630c1c5fdc0d79bc6010c1595798b6976665b3735158552b01232ba0e

Observation 0a3992ac-f256-4aa1-87be-50971fe84b26 · outbound

This paper cites Triton: an intermediate language and compiler for tiled neural network computations,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Triton: an intermediate language and compiler for tiled neural network computations,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:13.573170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:09.618175Z digest=sha256:cd1e961308849c020fbff7e1fd9e4c47f541de159bf0cddeed861119f91793a3

Observation b2475b7e-d19d-470d-bab2-1864008f8952 · outbound

This paper cites Transformers: State-of-the-art natural language processing,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Transformers: State-of-the-art natural language processing,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:13.287003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:09.724622Z digest=sha256:13fa96f8bede3eb2c91eced69fdd88754883812b54bd87f0e121b98615c9bd78

Observation b7f2c074-e5c4-4711-9955-6e96db117fca · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:09.826529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:09.826529Z digest=sha256:46b80590c8648501ec52a04eea21cc23ac9efcc40474805c5676c08752c021d3

Observation b44475d5-4d5b-4b6a-aa3b-fe28cae2bbc4 · outbound

This paper cites DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:09.890077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:09.890077Z digest=sha256:14721defc69776b139ddf3ad1ba39f1da3855be8e77eb7d369aea565ec246eab

Observation 3b2b6213-5e68-4e18-8bb8-1ed454617470 · outbound

This paper cites LongBench: A bilingual, multitask benchmark for long context understanding,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling LongBench: A bilingual, multitask benchmark for long context understanding,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:12.976678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:09.973087Z digest=sha256:9cb3d2042b62dcfa977812c05479c842f2464a366ad90bffbbb9df702aaf9134

Observation 01963edd-4a23-4919-b6cd-d2123191a12a · outbound

This paper cites ∞bench: Extending long context evaluation beyond 100k tokens,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling ∞bench: Extending long context evaluation beyond 100k tokens,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:12.732486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:10.082023Z digest=sha256:abbd9c519a740c6ef9a519a4390fe7b1375eee360f2d87b6476858c23fbc5d03

Observation a29cb238-8b36-448d-a202-e4a809b09dea · outbound

This paper cites Needle in a haystack - pressure testing llms,.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Needle in a haystack - pressure testing llms,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:39:12.426025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:39:10.193536Z digest=sha256:a9da9fb22aed385f70fd91bbd4b72cdc07b7a66e3251796751c4f2a915876585

Pith citing papers

No inbound Pith citation observations are available.