Pith. sign in

Paper Citation Record · LEDGER

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration

As of 19 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.11104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11104 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:03:19.387482Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:13:27.644720Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T05:13:28.150379Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact3
  • verified fuzzy1
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 39b75541-84d0-436c-a687-1d832d3d8a51 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.875957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.265408Z digest=sha256:9de17883f24495648ff6fe397cd83e394d08eb33108e99c4c584d88f6d3ed1d3

Observation fd441fa5-0ea3-4132-a475-96c32bb8aacd · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.866737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.269172Z digest=sha256:27ca8d41c454da72d5bf68594d87ebddaf7306e04b71790b5aeee39df232ebb4

Observation 3260fa7d-8bec-493d-a95d-2eb25927ded6 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.857783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.272325Z digest=sha256:89b8d81d7706c5f0faf1c8808879eb0f818df5e57ce3a5b10d5e53676a8d1217

Observation 0aa85b62-32dd-445b-8034-ac4fda8bcde0 · outbound

This paper cites Longformer: The Long-Document Transformer.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Longformer: The Long-Document Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.278649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.278649Z digest=sha256:9ddf4b562a7e6c68636a78ca529a273e8a99dbe5127d1a2effaecbea6cbc34e3

Observation 2a75ea7a-6630-4698-aa96-2d9059574388 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.281563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.281563Z digest=sha256:772c4e550fd196d26a8d8cda76ef9ea0cea04052903e62c0d5b58f87ec0a1513

Observation 109f8c1b-18d4-40e9-9e3d-309081fc6931 · outbound

This paper cites NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.284561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.284561Z digest=sha256:d5f4385201c63dfdbccabade7c82c8e13d799f75b2c0b13e24e9729025998f3d

Observation b5e93402-0bcf-4b35-aa7d-d1b0f66e2241 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Generating Long Sequences with Sparse Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.287621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.287621Z digest=sha256:046e27c236ec3048458fbe8dfb6bc14ac622d0930233235500a9dc6b5d57b17f

Observation ace05112-41fc-445a-8fff-53ea8b05380c · outbound

This paper cites Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.290731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.290731Z digest=sha256:bb6338b1728d550b8f7f83c3c42dd8589b9b7f76cf394fa1e746d308e202cdc7

Observation 2ed08cd9-9e48-4ab4-a16c-a33e2b3f3383 · outbound

This paper cites Adaptively Sparse Transformers.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Adaptively Sparse Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.293695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.293695Z digest=sha256:3fe49fc30952a29dc2989804dcbbf4bb1833db78fe08b9ceac396037acb7e701

Observation 72b51fed-4cae-4cf4-a006-764cc2c0b555 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.296945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.296945Z digest=sha256:c509bb7555374d51681ab850321aa7409159084f0593e61447eadb894a76716a

Observation f462fe05-3d87-4621-b1b1-f213bc3d6e00 · outbound

This paper cites Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.299889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.299889Z digest=sha256:931f213755d56937ddd3cada0bc6be9962dc86b928be62835ab0be90ef5ce7fb

Observation a38fa083-ed9b-4ed8-9083-7a3cbef28524 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.302732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.302732Z digest=sha256:5276c008be10c46410cde8f0e39191fa6d5e7c1cfb2a1e3ff6d5ae2d34032684

Observation b593f040-8eea-4bca-9588-07bbd603a0e1 · outbound

This paper cites Semsa: Semantic sparse attention is hidden in large language models.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Semsa: Semantic sparse attention is hidden in large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:19.849396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.305392Z digest=sha256:bf0b8287b40d3b3539d749fcedad274dc749a2909d1ca1deb4364670f377f5db

Observation 38126815-345c-449c-97d8-caf3610bff33 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.840241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.308260Z digest=sha256:eb90224ac0f411c3848f5e336df652b10bdae051e433f4edc2ec7eec9c3665fc

Observation 11123903-98f1-4399-8696-71a83d51e448 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.311000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.311000Z digest=sha256:e1be4c27aa7d5c3a2903cf20ed1c39c170563043c549ce3faf7e69b6ba04a372

Observation 7dfb9399-0a44-4a50-9657-3da18bb4e273 · outbound

This paper cites A Dynamic Head Importance Computation Mechanism for Neural Machine Translation.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration A Dynamic Head Importance Computation Mechanism for Neural Machine Translation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:03:19.616366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.313839Z digest=sha256:75ad0378d9ac27c26fdd1815b874efde81d8f151bd6300b2d9ea0874fdb4f2ab

Observation a9404664-de59-48ed-99fe-06cc174f009e · outbound

This paper cites Axial Attention in Multidimensional Transformers.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Axial Attention in Multidimensional Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.316871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.316871Z digest=sha256:3ad5e744083dc2929b8b2ad949543b4bd5371cde7db25e1e558216ddeeba882c

Observation 964a02ee-e3b0-467c-b77d-b3cd16ae2297 · outbound

This paper cites MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.319925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.319925Z digest=sha256:8fe328375dfbdc39262037eff3e9afd1ab80bf52d508fbc322205810afb26d79

Observation 60a43518-9018-4fbb-a493-0f4020fed5a6 · outbound

This paper cites Reformer: The Efficient Transformer.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Reformer: The Efficient Transformer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.322704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.322704Z digest=sha256:91c13e9e1a36fde9dda589ee4d85271488637f3cbda13fef7f7341a521915616

Observation 77f0c218-0e94-4bb1-80cc-3958813350b3 · outbound

This paper cites LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.325755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.325755Z digest=sha256:d037f1ed24e4284c7758146199997f6a0f2da6648560a2957861b10d15bc38a7

Observation dff2d4cb-47cf-4d2b-82dc-4e4ec51871a2 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration SnapKV: LLM Knows What You are Looking for Before Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.328603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.328603Z digest=sha256:3577931fe1053511ecb5da3bd97a8a3940242b9b592cd7bd10554daee2c0b331

Observation c5eacb4a-1b9d-4430-91be-00d18496f4a5 · outbound

This paper cites Global Attention Mechanism: Retain Information to Enhance Channel-Spatial Interactions.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Global Attention Mechanism: Retain Information to Enhance Channel-Spatial Interactions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.331633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.331633Z digest=sha256:a23de9b4cdf2d061728054810d85f7d38647be24a3080e86675076fb1b8d20d2

Observation e8386c03-f072-4bf1-82c1-b59efe1bbb23 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.831550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.334602Z digest=sha256:e9e9b0740785c3931e86dffe3fd07ba0c3aae4c76c280f9a9fb886e0eb8db2f8

Observation 47463b5e-831e-45b2-a3ba-4524dcb270d4 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.822960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.337352Z digest=sha256:d745d186dda3cbcf7e33ccd0a575c0711c1aa17e7c9970ceb5928828f2fd0689

Observation b849773a-ef26-49c0-b491-92f618feca64 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.340101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.340101Z digest=sha256:fe3f384a5cb93c1b30ca4839c6744a86ad01355ce59b3bda0b2f185c87d884bc

Observation 723e7c91-852a-4b51-baeb-20f365ac3591 · outbound

This paper cites Lightweight and Efficient Neural Natural Language Processing with Quaternion Networks.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Lightweight and Efficient Neural Natural Language Processing with Quaternion Networks

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:03:19.557693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.342944Z digest=sha256:e1604324e406931ff89739248e081b8c3f8cec225f9bc91d3a72402208c8939d

Observation aee541e8-2104-418f-88f2-f80226d18deb · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.345983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.345983Z digest=sha256:6acfd075c93f3b25059c6d5753a04739deae98f8a96118e1103a8494e21a4882

Observation 53206f46-2ae0-49ae-9df8-68285976e6a3 · outbound

This paper cites Multi-Head Self-Attention with Role-Guided Masks.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Multi-Head Self-Attention with Role-Guided Masks

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:03:19.545146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.348741Z digest=sha256:e00184850764d8670d15672e46ef5bd810e8a1ecacc78ec81925dbf021765f3a

Observation 78a2e262-6928-4375-ae3e-d37f98c80da9 · outbound

This paper cites Improving Transformers with Dynamically Composable Multi-Head Attention.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Improving Transformers with Dynamically Composable Multi-Head Attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.353231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.353231Z digest=sha256:d707ce5e30527268e48dc63ba117cfa5c63a026a9985a8423e3705389720d645

Observation 4ad6a6f0-2888-4cc0-bc53-f5b9ee758c26 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Efficient Streaming Language Models with Attention Sinks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.357102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.357102Z digest=sha256:035c06ce6eb71a5134123d8ae3d004e0baa5099d4e3ab41a9346d9e82b6ed6de

Observation f770c5da-57e3-4456-a076-a60800a339bf · outbound

This paper cites LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.359687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.359687Z digest=sha256:5d9ec2bdd3615d4f7ad9a6ff0892cc2c3dbf75d85fb1c78b0ce13cf0d2cf0d5e

Observation d96aaaba-19ec-46e6-b609-794c52746810 · outbound

This paper cites ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.362463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.362463Z digest=sha256:2852bf5ffbf1d1b3b9c825676781f1b280eaded22198b14531cb11ba199b8970

Observation 582542cf-e636-4fcf-a262-cdc2a48eb119 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.365339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.365339Z digest=sha256:9e4f3febd1ec32bc9b103bba78586a6f47002292b6ade709dcb7954c9b350738

Observation 92cf65c7-5490-460f-8c4f-d8e12ac71434 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.803606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.368297Z digest=sha256:4a48bea7e9d40280827dc810b4324cbb45663d9d5394723f11cd66d1eda81e64

Observation 6af78f03-c580-41bb-a8e2-244859840541 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.370887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.370887Z digest=sha256:146acc8c5300816b34bcb9537e2dc6c33220b70fa998adfc7b7127e348902d6a

Observation 5814662f-8955-4c90-9230-7afdb02297a1 · outbound

This paper cites DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.373594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.373594Z digest=sha256:fefa00a59d845c7c36381adac56426229018719812fdf7a4b8f7a14cfa761261

Observation e15d8fcf-d490-405b-bfa2-f6f4f5218f1c · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.376481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.376481Z digest=sha256:fd2c8e4d8f53c5b2860d53bd415fe3950085a528a1d7bdfa41f36363cfb33110

Observation ac3977f9-50d3-4bac-a657-5d6f08e176a4 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.784410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.379033Z digest=sha256:bde7b3090c4bab82b2bfe72145796e6c606553c5105d2c6a78669f76387c7615

Observation 0fa2375a-4a90-4325-82ad-523cb12c7c76 · outbound

This paper cites BUZZ: Beehive-structured Sparse KV Cache with Segmented Heavy Hitters for Efficient LLM Inference.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration BUZZ: Beehive-structured Sparse KV Cache with Segmented Heavy Hitters for Efficient LLM Inference

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.381771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.381771Z digest=sha256:123245b466c0825f0ecde076438d6a281b9b43c14fb517dbc893c7fe7a0f93d9

Observation 8ad58d27-eea1-4510-8dc9-59d415f56a26 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration SGLang: Efficient Execution of Structured Language Model Programs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.384593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.384593Z digest=sha256:a557f3ce23d659b7456a5d25d62524eb0a6b01bd1bf8b55eac0301a35ac11ce4

Observation a6411a1e-082c-4e41-a447-5238a7d44b9c · outbound

This paper cites BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.387482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.387482Z digest=sha256:71aaaaaeba267fad4d9c11d8e76cbba4de981d8d6e827ebb83f58e5db939f23b

Pith citing papers

Observation 3d0b4a51-b3ee-462f-9e79-f2da65fab369 · inbound

An Overview of Algorithms for Contactless Cardiac Feature Extraction from Radar Signals: Advances and Challenges cites this paper.

An Overview of Algorithms for Contactless Cardiac Feature Extraction from Radar Signals: Advances and Challenges DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:13:28.210636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T05:13:27.644720Z digest=sha256:855f4e9713ad3361d4c9c5cc8ed59e4cba94448238f43b7de2de711a0209a67c