Pith. sign in

Paper Citation Record · LEDGER

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention

As of 19 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 1 inbound Pith citation observation for arXiv:2605.18753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.18753 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T10:50:12.926232Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T05:40:56.557646Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

78 of 78 outbound references displayed

  • verified exact28
  • verified fuzzy42
  • unresolved4
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6d5ba521-272d-448a-8673-71cacd619d3a · outbound

This paper cites Is it really long context if all you need is retrieval? towards genuinely difficult long context NLP.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Is it really long context if all you need is retrieval? towards genuinely difficult long context NLP

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.146562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:820a078195b7b9213acd8272c6e7c9bbc1a39a187b364146e0cb65537d32ec77

Observation 929d89d6-30cd-40ee-8a68-1b6381af3f11 · outbound

This paper cites an unresolved cited work.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:53:26.135459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:2c222f22b6e6059c762f37bd18038c567cfc06cfe168dec24717a7868b7da6fe

Observation 60e5711f-49c1-4cba-a6b4-9e445e33c176 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Attention is all you need.Advances in neural information processing systems, 30

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.133698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:3ba01885cb52d849a87745d2ec37b08a67878c3068ba8fd12d3cbcb0a10c5860

Observation 7cde45ae-baeb-4fda-b8f7-8d05985f8c39 · outbound

This paper cites Softmax is not enough (for sharp size generalisation).

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Softmax is not enough (for sharp size generalisation)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.144002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:9cd9d5326bb457fc892ff61311107308fc659088bb00d8241c65a28f02f4c0ab

Observation c3e2f403-7f6e-4f7d-8a65-294a897559f5 · outbound

This paper cites Native sparse attention: Hardware-aligned and natively trainable sparse attention.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Native sparse attention: Hardware-aligned and natively trainable sparse attention

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.129008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:555e6556d2b2ae8c8a977e0402913a62243d05fcdda6c8047fb214e02aa97c3b

Observation 03f36460-b66b-4402-8a46-86c456e2fb4f · outbound

This paper cites InfLLM-v2: Dense-sparse switchable attention for seamless short-to-long adaptation.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention InfLLM-v2: Dense-sparse switchable attention for seamless short-to-long adaptation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.137533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:638386fd9652558d630804e76e4f4738565eb5e3f46780b9c41a509a24dc5f3a

Observation 0068619b-f1ab-4ccf-b610-bbd56a088321 · outbound

This paper cites an unresolved cited work.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:53:26.141925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:3a8c020f4756ecccfdf6a6f30f67f1975dae13adf522b51abb044eacb53cf1fd

Observation 13a84dfa-7a41-42fd-85bd-749ddcd466e9 · outbound

This paper cites Long-context generalization with sparse attention.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Long-context generalization with sparse attention

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.133509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:598f431ea7c87be68588c4800e7c5c161d24e025db5e851245b2d7a6856d9aea

Observation f54f691a-91c0-4bc2-8ee2-a6fdd31a56a4 · outbound

This paper cites MoBA: Mixture of block attention for long-context LLMs.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention MoBA: Mixture of block attention for long-context LLMs

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.102675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:1bd856cdf271d64336fe89d918bfed41612b0136d06a79c43555f1a4029f2fa1

Observation 96437640-d12c-414b-a20d-477cfe944867 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with IO-awareness.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Flashattention: Fast and memory-efficient exact attention with IO-awareness

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.094757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:5ac15899eab8abb758324509b73063a7698e1fd00c04fa476847d2f968ad38bc

Observation c1e71adb-e692-4559-91b5-c1e49de57ea7 · outbound

This paper cites From softmax to sparsemax: A sparse model of attention and multi-label classification.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention From softmax to sparsemax: A sparse model of attention and multi-label classification

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.090382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:0e50b25b8de41185bf70bc9cf2fdceb8bb6756bc115798a31955eca4eeb6e7c5

Observation 6d29233f-6a4b-4841-bc28-040776463dbf · outbound

This paper cites Adasplash: Adaptive sparse flash attention.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Adasplash: Adaptive sparse flash attention

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.122761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:544503172fa05442f51fc36c309da202c40771ac28fccff19414d6d59ac8d43f

Observation 28d35cc7-d532-45bc-bb60-7d3e98144cba · outbound

This paper cites Adasplash-2: Faster differentiable sparse attention.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Adasplash-2: Faster differentiable sparse attention

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.092660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:27492056cd1cf6b8996f94642e4cfdb80760f8fade7415bf1ba215182a433a34

Observation 411b1c93-e3a7-4f79-8df0-f8e783e4eea6 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.109441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:7cb53b441c2b919ca8224ac3b9aa1094b8f3e53c1408dce40d2826ef3af8d04c

Observation 6c096a2f-2461-42cd-a942-6d7e043880d6 · outbound

This paper cites Learning classifiers with fenchel-young losses: Generalized entropies, margins, and algorithms.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Learning classifiers with fenchel-young losses: Generalized entropies, margins, and algorithms

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.135662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:43d2e624e80f70713fb13c8cb9b1f57f4a46e8d635f7054eb5235a5a3adce793

Observation 2127a77a-bee5-4144-8add-8b4dc98c9bba · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.107670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:437f7d25f5c877662f866c811f2cadbe6d37ae28fc140468a2f7b1e75d6da23e

Observation 745df542-c48f-46b2-90d7-bb65983a9acf · outbound

This paper cites an unresolved cited work.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:53:26.107447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:e2061c5ad0e406efdaf28e836a87aa99b5b054c09e2deb266a324bd40ee368b7

Observation 382caca0-f26e-4ce1-a1d1-42e43378e158 · outbound

This paper cites Infllm-v2-data-5b dataset.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Infllm-v2-data-5b dataset

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.111009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:cfbc223ff2e36cd7f1ca8e185c4d15194bfc5b50fdc7f6344f945ae714a70565

Observation a219ed25-0cbd-4564-a9bb-26d52baf25a1 · outbound

This paper cites MiniCPM4: Ultra-Efficient LLMs on End Devices.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention MiniCPM4: Ultra-Efficient LLMs on End Devices

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.541414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:b4ef28f7189a2cd23b7cce4c94663e5e44317c909e4c3ca197ab2483386e488c

Observation 5c9972ef-8e6a-4392-8a57-6ab1c3edf280 · outbound

This paper cites RULER: What’s the real context size of your long-context language models? InFirst Conference on Language Modeling.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention RULER: What’s the real context size of your long-context language models? InFirst Conference on Language Modeling

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.112768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:0f25d11cea34c8b9b8690cb5fadfcf2f27e5ccf64c188372291222d854a9358a

Observation c454c62d-091b-4d39-a19f-f9b27fb03b8f · outbound

This paper cites HELMET: How to evaluate long-context models effectively and thoroughly.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention HELMET: How to evaluate long-context models effectively and thoroughly

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.109072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:99bf08420d3e21f0121dac18d9393fd120354ebb593a617d045e844db2072ba4

Observation 1e11a913-1c19-45e9-b9e3-b527eda3301f · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.142415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:e76635cd06444678b94b96154d76c21476002f22ee9f728b4291444d8042ff86

Observation 066a030b-ac45-4354-b6a3-e1b65e8d6471 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Measuring Massive Multitask Language Understanding

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.574632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:b04a4bdafe0cd0e2de89b9225dc67278f0593ff6b8400298d26abf0c528f39b9

Observation 2ea2c98e-808c-43b1-8933-551ad4bf317a · outbound

This paper cites Commonsenseqa: A question answering challenge targeting commonsense knowledge.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Commonsenseqa: A question answering challenge targeting commonsense knowledge

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.124815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:8a7ca8b851e424a12eeef8553b7e12ecf4ab3138af9067998cf0a6411a663419

Observation 4c26b643-b332-4ed3-92a0-779477d99aff · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Instruction-Following Evaluation for Large Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.559200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:7a48234affe42a062835c9dedb75e9b3675f953b9afd90104c3d6190c85f17dd

Observation 721e1ca2-1129-40ab-a4fa-c36d7c65a36c · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th annual meeting of the association for computational linguistics, pages 4791–4800.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th annual meeting of the association for computational linguistics, pages 4791–4800

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.116548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:3bccf9a488e25c2f5b26a5790c99a4cb1d6fbfafaaf5df353c8ffa2e7b633d43

Observation 51b7de01-03e3-4df9-820f-149602f4de4c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Training Verifiers to Solve Math Word Problems

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.502184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:0d521ef57392277b0768b4d08be69cebdd5f03e9229666850b96343d6049d8aa

Observation 4fce58cb-1e38-4b55-9c22-2bccc3a88622 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Measuring Mathematical Problem Solving With the MATH Dataset

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.577323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:fa7141c4aff063acc3d3646df026c6f63930e737676bb533103ce736af97081c

Observation 6af8292c-5eeb-4c28-a865-35a21d11dbfe · outbound

This paper cites Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.126760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:5b52149f45abbe7a3857a3474f7ddfe14d1df7e39125fa133a3050b13a2251dc

Observation 4431f30c-6e91-48fc-8b48-4d39681f5306 · outbound

This paper cites Program Synthesis with Large Language Models.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Program Synthesis with Large Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.556324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:43afbfd8093832e8fe7af81f61dd5354a95fd83c0716abc3dd9c855c2637b242

Observation b9bd9f9c-2486-4072-8da8-92df5d05d670 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Evaluating Large Language Models Trained on Code

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.519507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:cd447d5d80665fe2f5d0d87ea53d41fe38e0037e052d1661ba6ffa4c6f68f60f

Observation d210fddf-4bd7-43fc-9f7c-54e5f20d2873 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Efficient memory management for large language model serving with pagedattention

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.118486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:b587b09e6e2e342136ea603ab60462f7d624d5a3486e477ae0a926f94893d0e1

Observation 308a8b91-229c-4f72-a63b-f3d9af120b49 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.144440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:b5073d34990953a993c795e9a98d07b4f02f6bb9472ff49bac801993b98d7246

Observation 61c50e82-deb7-4d3b-83e3-0f96536ea180 · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Flashattention-3: Fast and accurate attention with asynchrony and low-precision

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.119438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:f49bc366d93e3c426bb394da35a7a9c789c544fa2e08882693574a4f0b1ed3fc

Observation 5991786b-911d-483f-b8fe-2a03898e7ed1 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.596117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:f3a90e9bd234c29730543a7bef3f1353c3933c94c3246316801660bb18c2f79b

Observation 7fd93b9f-ad9a-46b0-a97e-c4efac714fa1 · outbound

This paper cites Pyramidinfer: Pyramid kv cache compression for high-throughput llm inference.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Pyramidinfer: Pyramid kv cache compression for high-throughput llm inference

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.115897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:80fb202a3b6afd64b26e4f28b4b96706a66b169c0ed9f313666e5bc97fdcb70f

Observation 89fc6086-267b-4d24-960d-16c253156d00 · outbound

This paper cites Efficient streaming language models with attention sinks.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Efficient streaming language models with attention sinks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.075215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:ff7301259b1c0518d97562c33775226c850ce99a41e5a545cb2180a3fd79a2bc

Observation 856ad4ff-07df-434d-8a66-40f646d9527c · outbound

This paper cites Longformer: The Long-Document Transformer.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Longformer: The Long-Document Transformer

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T10:53:13.553401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:3fb11d8a206a91dac3cdc10aec52ee238122ef9db287b6dd086548e17091c9e3

Observation fcd1287e-80be-43f2-bb8a-2b4996452c3e · outbound

This paper cites Big bird: Transformers for longer sequences.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Big bird: Transformers for longer sequences

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.073490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:ccdf43ddf07f12c7968cc26b5c2e98ec6bb00533f65541693b265d9690496f7d

Observation 4af70864-ffa5-4506-b5c9-9afcf20246bf · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.114210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:8fad75249336a65e5850c3e4cf1ed13360966d854a95c6123c0f1a01eaa62a2d

Observation 44cbbdd6-45e9-46b0-93d7-d27c7a5d104e · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:53:13.547437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:708f149924183e9062aecaf95d0f6d5b907b8604edb4b7baccfa802097876e75

Observation f444d31c-07f9-4262-a33b-523a679281c1 · outbound

This paper cites Reformer: The efficient transformer.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Reformer: The efficient transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.117767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:d064f51ed053a245f16685c9dd9cbce34ecc6d87a855d4f7eb38dfd60be72a9a

Observation 61ed1033-3856-4c6c-932f-6bc89d8bd961 · outbound

This paper cites Spargeattn: Accurate sparse attention accelerating any model inference.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Spargeattn: Accurate sparse attention accelerating any model inference

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.565342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:81cab377c7adb8e067017b784d3166fc86949391cfc51503d6c148c78d688d5d

Observation 2de1d86a-f818-4c0c-924d-118e52d8a62a · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.550501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:06f94a335e69d6f5b0e9b3ac824d76fbf2d97eeed976455631293bc28b564a36

Observation f76f756c-f240-4be3-b87e-f23f48ce80e8 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.571507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:861b23c9b71eaffb2bc42bdcb04a061515cf7cc43a53c4ace125e9e2dcb9b3c4

Observation a4ec7517-e142-4d4d-be6b-264149bff7ed · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.580322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:92f67836d283ed56066c27251220375eca645e4e994588d2079a31ea0e1a949d

Observation 5bea1bae-d609-4e79-8070-001742a9a948 · outbound

This paper cites Lycheedecode: Accelerating long-context llm inference via hybrid-head sparse decoding.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Lycheedecode: Accelerating long-context llm inference via hybrid-head sparse decoding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.568362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:598a8b5285c9897038d1d1df5862903cc78100176c5b7da6a3c79bf22477f2cf

Observation 0a4c63d7-30fd-4472-a06f-48ab893eb12c · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.105830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:a3a79e410630e66faef4acf29cc46b6ea29c8943c69dabeaca2d729e77e6756b

Observation 242bae82-e641-4f8c-994f-c90ba31b61d1 · outbound

This paper cites R-KV: Redundancy-aware KV cache compression for reasoning models.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention R-KV: Redundancy-aware KV cache compression for reasoning models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.583802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:47c2cc464018b24b2eae593195f234c1c847667d45a3faad152d9d3a83fdddc8

Observation 84c5d80e-72f7-430f-b5f6-4a638ee239d4 · outbound

This paper cites Indexcache: Accelerating sparse attention via cross-layer index reuse.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Indexcache: Accelerating sparse attention via cross-layer index reuse

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.529441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:e93e99a415498469d061898fb056a0629ccc8eca6d97842810d29bce6086c9f4

Observation 5aa1b51d-5b99-4871-8de2-8d4d7dc237ee · outbound

This paper cites Infllm: Training-free long-context extrapolation for llms with an efficient context memory.Advances in neural information processing systems, 37:119638–119661.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Infllm: Training-free long-context extrapolation for llms with an efficient context memory.Advances in neural information processing systems, 37:119638–119661

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.121386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:579cfe84e2b68773fd26fa6219b29d1215f77de329017e4c2c9f9b1924cd47c4

Observation 2de173d6-f8da-40a1-b937-7a2692db3acb · outbound

This paper cites ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.593270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:46cea6c57ff63916999127a46348c238ec7e3ee36ff6cb04f8621341d27feb3e

Observation 0a017d96-554a-4abc-852f-ab2b4b2a37fe · outbound

This paper cites Nosa: Native and offloadable sparse attention.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Nosa: Native and offloadable sparse attention

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.589945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:0db25619d4f7f4468c0e42a8bb14e5e4ca95bac0320988ad7dd77dca200e238d

Observation 178987df-f6f3-4739-a48d-eabf26415659 · outbound

This paper cites an unresolved cited work.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:53:26.112291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:2f78588aa846ba63563e4b5a9f60e8d4b4334c604a8ad80eec78920b12d56b2b

Observation d5245d6c-a565-4287-8f3e-f21f15e45ead · outbound

This paper cites Inference-time hyper-scaling with KV cache compression.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Inference-time hyper-scaling with KV cache compression

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.148533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:be5d3abf3d2e5807bfa81586c8cc8b655dd1b6b4de8645f6e49cb8ce4bef57a8

Observation 6be08aa3-d628-42e7-8402-535a22e41e32 · outbound

This paper cites Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.140118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:e60a5838f302f61f85897d165cea08936aa68964dbd78399c1bff446b74d6edc

Observation 4b0f79cf-da3a-4097-840d-827433bdaddb · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.598927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:270ec7e466291c06c442f639dcabfc1d76a193d0565f35976d6cab8e6e1caa89

Observation d54f0181-1c18-4a4f-b48f-72eeb58fc5ae · outbound

This paper cites Pqcache: Product quantization-based kvcache for long context llm inference.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Pqcache: Product quantization-based kvcache for long context llm inference

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.114446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:b63319960dfaefb84fc43689bd0d790ba5e393f59e7c7d510def55aa885b1e9c

Observation a8990c60-f268-4918-920b-394a6500e38c · outbound

This paper cites MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.513638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:b554039be98fc2359754b086e72754345bc39b12069a620dedea9b4125151956

Observation 015213c6-84eb-4fbb-bb4f-1b04d4a97998 · outbound

This paper cites SeerAttention-R: Sparse Attention Adaptation for Long Reasoning.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.586873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:3fcdce7f3fbc0ee72a14068300a5dcbcaa0cabe4384132b9ea80bb35dfdf4ea6

Observation 22512576-7a2a-4be3-8518-2108e49e5c76 · outbound

This paper cites SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.535272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:6d052d2c445a39ba1e68d5644c8a9558863d044768b95df1cfea6e7b93b1f11a

Observation 8c463151-981c-4626-9282-0d893fa8f0bc · outbound

This paper cites Flash sparse attention: An alternative efficient implementation of native sparse attention kernel.arXiv e-prints, pages arXiv–2508.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Flash sparse attention: An alternative efficient implementation of native sparse attention kernel.arXiv e-prints, pages arXiv–2508

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.099196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:543c0ee0d8f139685d791f7e2aebdedbd15fc00f17038688c460a60bf9f52dbe

Observation 523dda38-9d8c-4fb4-bf91-921b201947fe · outbound

This paper cites Hsa: Head-wise sparse attention for efficient and accurate long-context inference.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Hsa: Head-wise sparse attention for efficient and accurate long-context inference

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.065707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:b26cdfb1bfa85d485aadac1294453c68acce2c07e202a593763d298201cc67db

Observation 8bc6fa7f-847b-4318-b89a-b5337e39fa49 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.538082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:ce4190715a63d079bed08c891981c93a58b4667913b607dcf7ccc3daada1eaad

Observation 28691fa5-f1df-4e3a-a74a-f9d6191c0ed2 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context in- telligence.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Deepseek-v4: Towards highly efficient million-token context in- telligence

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.063949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:9ee1ad3f021c7fc894363cbdec8ff32d6e1ca83693d5cfed4355f7c98c5215fe

Observation a06cb1cf-cf5e-4d2e-ac87-937f1b96d7e7 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.562151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:6b2485fc5b82d3c6961da4f953a353a5dfa4fdfe4ef479a294c0dffabc4c7ac4

Observation 6840e049-0ae7-4fbc-92b4-1063b144f783 · outbound

This paper cites Spargeattention2: Trainable sparse attention via hybrid top-k+ top-p masking and distillation fine-tuning.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Spargeattention2: Trainable sparse attention via hybrid top-k+ top-p masking and distillation fine-tuning

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.510469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:c15a25288aea233e599faa5f6dbf12c58ed32f90047dc52d20d64f11e57a8b6e

Observation b802fea6-bcc7-4764-a155-64c75f823508 · outbound

This paper cites Double-p: Hierarchical top-p sparse attention for long-context llms.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Double-p: Hierarchical top-p sparse attention for long-context llms

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.526500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:2fd171320e3f2a2655a08afb61ab0c49ff78632df34c720bed6f74e2249c0136

Observation b387206a-3673-41b1-9161-664b978492c4 · outbound

This paper cites Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.523245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:af041251ecf6818192cdc4d03de21cae2f38fe1f4e5f20f7cacdeb7dc5314ff0

Observation 5e5eac3a-8f1c-4e36-af15-f69d8cda4fb0 · outbound

This paper cites Possible generalization of boltzmann-gibbs statistics.Journal of statistical physics, 52(1):479–487.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Possible generalization of boltzmann-gibbs statistics.Journal of statistical physics, 52(1):479–487

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.102971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:fb3eff964da003a5342289d19addadc60dbdc2bb308d86b35390e90c1e41afc2

Observation 52831df3-351c-45eb-bef1-1cf126c01724 · outbound

This paper cites MiniCPM: Unveiling the potential of small language models with scalable training strategies.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention MiniCPM: Unveiling the potential of small language models with scalable training strategies

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.077431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:08dd28c75886bbc49289835b574cfedfbaf1e713be32333f0e81b390aa272962

Observation 8909391e-6eb6-4e71-bdd6-9f125bc24e1b · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.516609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:e1925646e98b358788e74070e2b6dc6792083818c7ea233873748265e41db63d

Observation 15a95ccd-30b8-4cc5-a967-a66f05443e93 · outbound

This paper cites Olmes: A standard for language model evaluations.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Olmes: A standard for language model evaluations

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.110754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:17b248f246bb1683e41792f439a11103304c9348e52ed5e75608ff1dba0bc2c8

Observation 0ae6d7d9-f2eb-44a5-81e2-f604dafe161f · outbound

This paper cites The language model evaluation harness, 07 2024.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention The language model evaluation harness, 07 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.101138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:89298a9b05e392ac852503970f864a1673da5ad2336f005c145b310d92545249

Observation 1a315cbe-b00b-4ae8-90bf-c80d01e79514 · outbound

This paper cites Hardware-aligned hierarchical sparse attention for efficient long-term memory access.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Hardware-aligned hierarchical sparse attention for efficient long-term memory access

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.544336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:b150444c41a2c4de5ac2f209478b5042a532c71dff4f9e1039558a2543c06b9c

Observation 63a076b0-92e8-4a4b-ae1c-c3535cfb464d · outbound

This paper cites Every to- ken counts: Generalizing 16m ultra-long context in large language models.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Every to- ken counts: Generalizing 16m ultra-long context in large language models

Reference 76

Resolution
malformed identifier
arxiv_id, observed 2026-05-20T10:53:13.532290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:7f0861591389643559646cccf868621a954193096050a6de1d9c5d615c6e38f2

Observation a1a735b0-6b28-4105-9879-e05470164930 · outbound

This paper cites lim n→∞ H aggrsoftmax z(1),z (2),· · ·,z (H);θ logn = 1.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention lim n→∞ H aggrsoftmax z(1),z (2),· · ·,z (H);θ logn = 1

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:53:26.060183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:f56d3e98b086352214006d22365648c73031b22bad025804f015042cf5228169

Observation 92d0096f-f18d-4542-a5a9-224759a06778 · outbound

This paper cites Proof.We first prove that softmax head aggregation is dispersive.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Proof.We first prove that softmax head aggregation is dispersive

Reference 78

Resolution
malformed identifier
raw_fallback, observed 2026-05-20T10:53:26.066170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:fef0d00940721fc6c950afe07af4e3e6416606f1dd40b303e6208ce93d549476

Pith citing papers

Observation db6737fa-c185-44b1-a272-8815764a89af · inbound

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling cites this paper.

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T05:40:56.557646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:40:56.557646Z digest=sha256:1d4480cf9d3f82cb5271b898286fcada6b148745896977ef97bf9dd79bdbf36d