Pith. sign in

Paper Citation Record · LEDGER

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration

As of 7 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.11104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11104 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:03:19.387482Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:13:27.644720Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T05:13:28.150379Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact3
  • verified fuzzy1
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 39b75541-84d0-436c-a687-1d832d3d8a51 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.875957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.265408Z digest=sha256:d9b31b3d3bb0e82fbfce9acbd1df3f139eda440e31111bdb35af9851795db8d3

Observation fd441fa5-0ea3-4132-a475-96c32bb8aacd · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.866737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.269172Z digest=sha256:0bd0b3bcc153f2b2f20adc2bfaef560381e5eb6b4f46d6e4453746fe2fb9c6b2

Observation 3260fa7d-8bec-493d-a95d-2eb25927ded6 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.857783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.272325Z digest=sha256:6239a138bb93257020be52bd972dc58c78f675f63d5a517750c396c5fec7d6f6

Observation 0aa85b62-32dd-445b-8034-ac4fda8bcde0 · outbound

This paper cites Longformer: The Long-Document Transformer.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Longformer: The Long-Document Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.278649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.278649Z digest=sha256:5df291450f9f3d07ea0dfc2097d3f3bed0c3851dec05507af92738cf26c3a027

Observation 2a75ea7a-6630-4698-aa96-2d9059574388 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.281563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.281563Z digest=sha256:146f349fbc7714bb7f848a2e35549e3d0f6429fc19de254a14ab1c4011554899

Observation 109f8c1b-18d4-40e9-9e3d-309081fc6931 · outbound

This paper cites NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.284561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.284561Z digest=sha256:23188989e5a6a6b6cd4dd43f4c4ab31d63def19d03742576b9027575ec08c5ca

Observation b5e93402-0bcf-4b35-aa7d-d1b0f66e2241 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Generating Long Sequences with Sparse Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.287621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.287621Z digest=sha256:919842c04942c0fe9b759654886914b05755fe069d1a5cb51bb3d8766d2c0833

Observation ace05112-41fc-445a-8fff-53ea8b05380c · outbound

This paper cites Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.290731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.290731Z digest=sha256:566f5a7bb03b1bdb7f5bac0c1c807aee5be5fe2a9a538f476b54eadd45d95d0b

Observation 2ed08cd9-9e48-4ab4-a16c-a33e2b3f3383 · outbound

This paper cites Adaptively Sparse Transformers.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Adaptively Sparse Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.293695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.293695Z digest=sha256:e6739f9ea233df17eac7292b1f929744b5f190d619bf6384adb2158487a33539

Observation 72b51fed-4cae-4cf4-a006-764cc2c0b555 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.296945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.296945Z digest=sha256:e2b54d45bbdc2726a49155c727c1dea5326e65c6fdf5b3b7435edd166e738418

Observation f462fe05-3d87-4621-b1b1-f213bc3d6e00 · outbound

This paper cites Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.299889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.299889Z digest=sha256:e089c99b4c617fab77b81579a4167db759157cbf3f2a9eb7b680b29f68f8de9d

Observation a38fa083-ed9b-4ed8-9083-7a3cbef28524 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.302732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.302732Z digest=sha256:d211da54056a966172172dbed538e4b7c258c457fc9419ddbb0ffaa94e9ab1d6

Observation b593f040-8eea-4bca-9588-07bbd603a0e1 · outbound

This paper cites Semsa: Semantic sparse attention is hidden in large language models.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Semsa: Semantic sparse attention is hidden in large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:03:19.849396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.305392Z digest=sha256:99bd7679edc6d54ed4da3444455987fd120a33e655dc5de07f085abf22246c3f

Observation 38126815-345c-449c-97d8-caf3610bff33 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.840241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.308260Z digest=sha256:0b731bb20d53d920e68fe08de235eaadf30d10bae7e13edd89ef7bee060b7c59

Observation 11123903-98f1-4399-8696-71a83d51e448 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.311000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.311000Z digest=sha256:a95988b1b368762cd02ec1b61456091998f7539e43f649d06a2af8e61195d5ff

Observation 7dfb9399-0a44-4a50-9657-3da18bb4e273 · outbound

This paper cites A Dynamic Head Importance Computation Mechanism for Neural Machine Translation.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration A Dynamic Head Importance Computation Mechanism for Neural Machine Translation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:03:19.616366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.313839Z digest=sha256:14d3c9f84100f759b3c8df28c188a5152b33bb563e922dcb45f60d835d52dff3

Observation a9404664-de59-48ed-99fe-06cc174f009e · outbound

This paper cites Axial Attention in Multidimensional Transformers.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Axial Attention in Multidimensional Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.316871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.316871Z digest=sha256:0eef89e92d65f7a5b18eba9d468fb383b8ea1d8d2b62de0fede1e8def9ac2020

Observation 964a02ee-e3b0-467c-b77d-b3cd16ae2297 · outbound

This paper cites MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.319925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.319925Z digest=sha256:0750246203eefe5a9205b939da423cf9ec365aa19553109f994bed4bc2cbdf11

Observation 60a43518-9018-4fbb-a493-0f4020fed5a6 · outbound

This paper cites Reformer: The Efficient Transformer.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Reformer: The Efficient Transformer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.322704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.322704Z digest=sha256:07381c199bb6bbe6111bbe39004bffc34553fa1a3614e43ee5b780abaf01a13b

Observation 77f0c218-0e94-4bb1-80cc-3958813350b3 · outbound

This paper cites LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.325755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.325755Z digest=sha256:c883d2df579a97450e5561a1c21eecb04d55d6c000a3795b89ea9ca1797df257

Observation dff2d4cb-47cf-4d2b-82dc-4e4ec51871a2 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration SnapKV: LLM Knows What You are Looking for Before Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.328603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.328603Z digest=sha256:3a3f20c287f8414b35df7e223e042d761e65ad6d688d5672e4cf15a4e7f9e91e

Observation c5eacb4a-1b9d-4430-91be-00d18496f4a5 · outbound

This paper cites Global Attention Mechanism: Retain Information to Enhance Channel-Spatial Interactions.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Global Attention Mechanism: Retain Information to Enhance Channel-Spatial Interactions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.331633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.331633Z digest=sha256:7012e1f3654ba3bde98d12e4b7ece6e86eb952be1a823e72573f55fa11069730

Observation e8386c03-f072-4bf1-82c1-b59efe1bbb23 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.831550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.334602Z digest=sha256:8dfa05b6b2920eecfd258c3991f68653b05e81ff22616ee7aea0ef591bb65064

Observation 47463b5e-831e-45b2-a3ba-4524dcb270d4 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.822960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.337352Z digest=sha256:aa16f5d42b1019bbbb456c1f051433b651caa54b6b067142edc5a602fe0763a4

Observation b849773a-ef26-49c0-b491-92f618feca64 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.340101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.340101Z digest=sha256:ddbc00a15838b856c1192fbe5a4036570b95d34baaa64f2dc5e03243c25c1a70

Observation 723e7c91-852a-4b51-baeb-20f365ac3591 · outbound

This paper cites Lightweight and Efficient Neural Natural Language Processing with Quaternion Networks.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Lightweight and Efficient Neural Natural Language Processing with Quaternion Networks

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:03:19.557693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.342944Z digest=sha256:2c5e62ecd53ee0477a5b4e810f75a3f34a125e9ffd1a27ca7d3feb3b7aa80248

Observation aee541e8-2104-418f-88f2-f80226d18deb · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.345983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.345983Z digest=sha256:0ffcbc6c0508ee965abac1fba5f0164eec0d314524bc670dffe1a11ba5911086

Observation 53206f46-2ae0-49ae-9df8-68285976e6a3 · outbound

This paper cites Multi-Head Self-Attention with Role-Guided Masks.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Multi-Head Self-Attention with Role-Guided Masks

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:03:19.545146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.348741Z digest=sha256:384d81f299625fb70e5c598450beb01cc202b9b2d6e03557ca271a30b200b92c

Observation 78a2e262-6928-4375-ae3e-d37f98c80da9 · outbound

This paper cites Improving Transformers with Dynamically Composable Multi-Head Attention.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Improving Transformers with Dynamically Composable Multi-Head Attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.353231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.353231Z digest=sha256:e62f8e0ceb92f82a310c3582afdaaf05cdcd24d9f160eb3ffd1790a32c18f62c

Observation 4ad6a6f0-2888-4cc0-bc53-f5b9ee758c26 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Efficient Streaming Language Models with Attention Sinks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.357102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.357102Z digest=sha256:f257432b02b98b65b77bf8becde1b8a9495003009c6bb01afecef1e4e66aa345

Observation f770c5da-57e3-4456-a076-a60800a339bf · outbound

This paper cites LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.359687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.359687Z digest=sha256:6df4a224d6d748314fd9c5143039d7653e84e918f27840f2daba4d3cbc72c7c7

Observation d96aaaba-19ec-46e6-b609-794c52746810 · outbound

This paper cites ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.362463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.362463Z digest=sha256:d973358f6d332c2c7999a65114b6cf7718de69da234a1d7c9739c97b211428f5

Observation 582542cf-e636-4fcf-a262-cdc2a48eb119 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.365339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.365339Z digest=sha256:87007d52320bc1d3c6a81d2d961de79365c87676db069d5af2b82a87ffc36e75

Observation 92cf65c7-5490-460f-8c4f-d8e12ac71434 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.803606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.368297Z digest=sha256:17562994fb3ded2f053b018248c6a091d7c9535ea9d157ea68495f1fcbac5255

Observation 6af78f03-c580-41bb-a8e2-244859840541 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.370887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.370887Z digest=sha256:c3a4d20d374f157e16199f93553b36d286c7e026e6dc8745c1205362346e80ff

Observation 5814662f-8955-4c90-9230-7afdb02297a1 · outbound

This paper cites DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.373594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.373594Z digest=sha256:533dda1f7a4e2b34e9da8e8da57cfafc88fa24ef92158252f30f01e19b785c59

Observation e15d8fcf-d490-405b-bfa2-f6f4f5218f1c · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.376481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.376481Z digest=sha256:8c8d4bf637bbb2f47169209a7b440e8f036fecef13a7724d5fb2a33d310c543b

Observation ac3977f9-50d3-4bac-a657-5d6f08e176a4 · outbound

This paper cites an unresolved cited work.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:03:19.784410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:03:19.379033Z digest=sha256:cf0aff7b92d38924cbe72ec7e11068df98afe23bf90035621e58545a87ae426e

Observation 0fa2375a-4a90-4325-82ad-523cb12c7c76 · outbound

This paper cites BUZZ: Beehive-structured Sparse KV Cache with Segmented Heavy Hitters for Efficient LLM Inference.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration BUZZ: Beehive-structured Sparse KV Cache with Segmented Heavy Hitters for Efficient LLM Inference

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.381771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.381771Z digest=sha256:bf57928f658563e99845a12f6021253c53605445804054fb2998b70c595a5204

Observation 8ad58d27-eea1-4510-8dc9-59d415f56a26 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration SGLang: Efficient Execution of Structured Language Model Programs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.384593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.384593Z digest=sha256:6956b915ec992f6c0ee56773bc76677922bcfeeec4705e385d3e7b099a171759

Observation a6411a1e-082c-4e41-a447-5238a7d44b9c · outbound

This paper cites BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching.

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:19.387482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:03:19.387482Z digest=sha256:ab5a105c1753d23fb762c630f4bc658289a0461c144cfe02ba491be3f4a9cb73

Pith citing papers

Observation 3d0b4a51-b3ee-462f-9e79-f2da65fab369 · inbound

An Overview of Algorithms for Contactless Cardiac Feature Extraction from Radar Signals: Advances and Challenges cites this paper.

An Overview of Algorithms for Contactless Cardiac Feature Extraction from Radar Signals: Advances and Challenges DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:13:28.210636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T05:13:27.644720Z digest=sha256:3681e0db290824fe4513db6fff4b4e08c4b6e195512fd3ff1a05a9a2ceafb600