Pith. sign in

Paper Citation Record · LEDGER

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding

As of 18 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2607.24555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.24555 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T11:49:11.813142Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved76
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 959d079c-d81d-457b-b730-c0597f54713c · outbound

This paper cites LongBench: A bilingual, multitask benchmark for long context understanding.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding LongBench: A bilingual, multitask benchmark for long context understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.599287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.599287Z digest=sha256:a2e2cc7e8c5090cb7838cae511d449752a8b652b13950b35ef184dec6a900627

Observation 6bfc83ab-cfa0-4dc1-b86d-6a03fe53ac7c · outbound

This paper cites RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.602883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.602883Z digest=sha256:d109969481d8c7cb5d4e8948909afadae5d6bf4b1586edbfe4aac8c630662888

Observation 88edf86c-4b9d-40e6-bfd6-61030eb40ecf · outbound

This paper cites Longformer: The Long-Document Transformer.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Longformer: The Long-Document Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.605941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.605941Z digest=sha256:931223a73f69812ff673302bb2bb2e0a0cb018fa3c440a3e2738c77340c8409d

Observation fd2a06c7-b6e3-4a44-bf38-7a295f928a31 · outbound

This paper cites R-KV: Redundancy-aware KV cache compression for reasoning models.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding R-KV: Redundancy-aware KV cache compression for reasoning models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.609016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.609016Z digest=sha256:cb861e4b20463737dbf37342bf97225b314326808851b26422f1731dd4595418

Observation 24c27a4f-b41a-4cb0-9434-542160ff4baf · outbound

This paper cites Runtime-Certified Bounded-Error Quantized Attention.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Runtime-Certified Bounded-Error Quantized Attention

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.611913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.611913Z digest=sha256:776502620ab503a2c3a3f1cab800783558080ecdc83f86c8f07d07bc3d7d8801

Observation 9efb6cf7-0d03-46c7-bd1f-c19c47410427 · outbound

This paper cites Palu: Compressing KV-Cache with Low-Rank Projection.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.614843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.614843Z digest=sha256:f3a4401483486ab989c2265ebf1e2d7bee424ca03222de0d5b6eace85f174a0b

Observation 6c8c0567-ff9a-46fc-b902-333781a667e8 · outbound

This paper cites Value-Aware Stochastic KV Cache Eviction for Reasoning Models.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Value-Aware Stochastic KV Cache Eviction for Reasoning Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.619336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.619336Z digest=sha256:828e16bb06651191016c31fb47ef2c413b2d70b31b1e60bb5ea15ee3b14e751a

Observation 105a00ad-7f73-4dd8-960c-490ae7a98afb · outbound

This paper cites MagicPIG: LSH sampling for efficient LLM generation.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding MagicPIG: LSH sampling for efficient LLM generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.622393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.622393Z digest=sha256:d9ea1508aa686979b4010bdbfe2d03511f9d2440231496b4ebd083c7df367fc0

Observation 787f2b05-3c08-4b5c-83b6-d1a1e38031e1 · outbound

This paper cites Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.624951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.624951Z digest=sha256:ca3ea536d43084d183d4c45e55114079a01381f54feac5d0b081872655901094

Observation bf45fbee-d439-4b21-8200-1a204d6aed2e · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Generating Long Sequences with Sparse Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.628245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.628245Z digest=sha256:9bc65d80c71a9d0fb56df62a7647dd5f701ee3904f60d1fb5e0c893c06daf9eb

Observation 4bd23b9f-fc37-48dc-8cc2-2631cb030d7a · outbound

This paper cites DeepSeek-V4: Towards highly efficient million-token context intelligence.arXiv preprint arXiv:2606.19348, 2026.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding DeepSeek-V4: Towards highly efficient million-token context intelligence.arXiv preprint arXiv:2606.19348, 2026

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.631097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.631097Z digest=sha256:e372a97a967c0e84b27fcc114d3a47dabdf998d6fe9cdead7d6de021c3fe6545

Observation f14d6af5-3a8f-4761-ad5d-786073bccf16 · outbound

This paper cites vAttention: Verified Sparse Attention.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding vAttention: Verified Sparse Attention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.633572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.633572Z digest=sha256:c08bd026a13c27eb1e1484c33722a8e006377325331c6edfc9d15f343201d90e

Observation a7af074a-ce98-491a-ad7b-294fc8fb8aef · outbound

This paper cites Kascade: A practical sparse attention method for long-context LLM inference, 2025.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Kascade: A practical sparse attention method for long-context LLM inference, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.636314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.636314Z digest=sha256:c171109e7f2d0c8597239f736702ea2c6656cc90e5afcde3a33d9bc7d4c7cd6c

Observation 5a79ae37-134a-419b-9838-285db21d3a9e · outbound

This paper cites Expected attention: KV cache compression by estimating attention from future queries distribution, 2025.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Expected attention: KV cache compression by estimating attention from future queries distribution, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.638817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.638817Z digest=sha256:584d8040445acef755498eff3b963c35fbc8f06a9f853ed0276cdaa249e3ad23

Observation 576f0597-61dd-4d3d-bb0f-3d1c41622851 · outbound

This paper cites The approximation of one matrix by another of lower rank.Psychometrika, 1(3): 211–218, 1936.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding The approximation of one matrix by another of lower rank.Psychometrika, 1(3): 211–218, 1936

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.641335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.641335Z digest=sha256:f76fd3e7f3759e891e79bca1a062c6736e8a17af9c0f6e9779be22b580a581c5

Observation 4301edce-d036-4326-a5ac-7512a5da7da9 · outbound

This paper cites Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.643900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.643900Z digest=sha256:8324be16e744a055f14e03fbc349b8ea52648724b9eb9d3707b5145203002aab

Observation 5f24662f-4ecb-4e8a-8539-5011aac93f39 · outbound

This paper cites SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.646757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.646757Z digest=sha256:abd5a0bbccc489f2d3f3da8e6d28239d3b32f78bca477da27fe5cfd8f53f01d3

Observation aaab944d-c458-444a-a56e-f7e5f59bbc3a · outbound

This paper cites The Llama 3 Herd of Models.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.649772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.649772Z digest=sha256:a8d7b43a4d3c4c1feec8dd6332907bf9d1d00e42d0b38a3a6785b42b965de047

Observation d64abe1e-c04d-4ccd-b828-ada98883c038 · outbound

This paper cites Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.652489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.652489Z digest=sha256:b6da5a59339604fa1f0002d373c6aafdbfd1e73a4af8060b9af03a294682c562

Observation ddb8a54c-fef2-44d7-bffc-93230c7d777b · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.655479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.655479Z digest=sha256:7cfd5dfe935b29806be3c3ff0153429196df9b489aab35efcc604400f5b8a99b

Observation 774ad2d5-f9fb-48ef-8bbf-3ca35b8fef0f · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Measuring mathematical problem solving with the MATH dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.658299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.658299Z digest=sha256:380c717c59ec4a52aef57015bc6af4d3cbf37628bd571a2bdf3dd0e8e638e67f

Observation 8018d232-67d2-497e-93f9-0e6a4ca29532 · outbound

This paper cites SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.660856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.660856Z digest=sha256:30e32a0062db4baeeee8886c12201a1039c1977d2aac0ce99c4f5143aa2d6eb6

Observation b5e1e7c3-cf4f-4b72-9352-3f3c2b20fe4e · outbound

This paper cites KVQuant: Towards 10 million context length LLM inference with KV cache quantization.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding KVQuant: Towards 10 million context length LLM inference with KV cache quantization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.663751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.663751Z digest=sha256:b2a0965cd512d75d71ca3189b1d69d2e902ac69ff3d313c5091325a670609cfa

Observation 34800a76-c8d5-494c-93f1-e544608b5d5d · outbound

This paper cites Multipole attention for efficient long context reasoning.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Multipole attention for efficient long context reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.666428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.666428Z digest=sha256:53cdb9b5bf269b670d7b75a8d3213c527910cdd27fd39bb7d6094c313960df6f

Observation 65566cb2-fd52-4269-92f0-99ac7392a2ec · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.668924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.668924Z digest=sha256:aad0dea9121a714185d48fea5c631a3057701bb050309a31f041a93c3923bb2c

Observation 14757c3e-1a75-4357-91d4-012f5450664d · outbound

This paper cites Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.671759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.671759Z digest=sha256:fd5ce203ab73b1fdef318a1ee944d7cdbdbcbbfb4ff21ff7616b873dcd8edcc7

Observation 2eb3acca-ace5-466f-a1ee-345b70edc513 · outbound

This paper cites MInference 1.0: Accelerating pre-filling for long-context LLMs via dynamic sparse attention.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding MInference 1.0: Accelerating pre-filling for long-context LLMs via dynamic sparse attention

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.674420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.674420Z digest=sha256:5fdb82b7e1c84c1bea965e831f239f9a218386d414bb5aa0aa998ab4f126c888

Observation 516b3f31-b1e3-4d4f-a77e-dd5e6e19b868 · outbound

This paper cites Lee, Sangdoo Yun, and Hyun Oh Song.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Lee, Sangdoo Yun, and Hyun Oh Song

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.676916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.676916Z digest=sha256:2a7ba9031a375430885cdc4c001baee840c889851cc9bde6e3ad848878933a1f

Observation 7a346038-c55e-4124-bf64-fa3a8adbaf57 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Efficient memory management for large language model serving with pagedattention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.679444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.679444Z digest=sha256:679836ce6ba345114a9d12b86e2c0dbc0f26e1f9b0886d8070936d4d681997b5

Observation 04a72802-65a1-421c-ac88-b62ee91cb951 · outbound

This paper cites InfiniGen: Efficient generative inference of large language models with dynamic KV cache management.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding InfiniGen: Efficient generative inference of large language models with dynamic KV cache management

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.681959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.681959Z digest=sha256:ef7287be73c1bf769ef343decb874e4a8d2c5bcf282ffd33ffc75604c28c65e6

Observation 1ffff939-3524-4b72-b215-9f3326ba376a · outbound

This paper cites Rakhshan, and Guillaume Rabusseau.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Rakhshan, and Guillaume Rabusseau

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.684394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.684394Z digest=sha256:39fcc595b1572cf5f4e209f33955fd961732ea3400d780fd0c9af84488ce6de3

Observation bae6cc65-27c4-4f78-b7bb-8094a25943fe · outbound

This paper cites Efficient low rank attention for long-context inference in large language models.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Efficient low rank attention for long-context inference in large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.686847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.686847Z digest=sha256:cbefd4b01c004ea66f8dc6e8308cd806086d8630fb975cca5f26f7517fce4066

Observation 806acae9-9c8d-4f33-881f-b309b9aff0f1 · outbound

This paper cites Snapkv: LLM knows what you are looking for before generation.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Snapkv: LLM knows what you are looking for before generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.689334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.689334Z digest=sha256:b59e0948cbe4f8e630004b74fb8a787fc8b989663583e8e787a2a481a266a194

Observation 542fa98f-0031-4df5-8222-cfe9b0fafef2 · outbound

This paper cites Let’s verify step by step.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Let’s verify step by step

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.691973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.691973Z digest=sha256:1a9452113ed6b5f202d6b7bb924b4f6eef10da968a46b42649962bec097bf47b

Observation cc221710-858e-4c41-88a6-e2e7264ce6fe · outbound

This paper cites Twilight: Adaptive attention sparsity with hierarchical top-p pruning.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Twilight: Adaptive attention sparsity with hierarchical top-p pruning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.694489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.694489Z digest=sha256:538c64b7dd67860a1858e558781bc8d6afb024dc7392fb5b38a1fcc252738bb2

Observation 7f953dad-68ab-447c-8805-5da810eadb73 · outbound

This paper cites RetrievalAttention: Accelerating long-context LLM inference via vector retrieval.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding RetrievalAttention: Accelerating long-context LLM inference via vector retrieval

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.696980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.696980Z digest=sha256:e208d88629e541917dea5beb145c6a3b4d7665bc45f60e8c1669c5d94b0619df

Observation 2e0d35b9-7965-439c-b3f0-21a020d041ed · outbound

This paper cites ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.699513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.699513Z digest=sha256:56e91237a6e13330986cc9307dbf944173e4409048348c8dcf8d7a786f1bc846

Observation 48f52fbd-9abb-48af-8f7f-82002a348382 · outbound

This paper cites ReasonAlloc: Hierarchical Decoding-Time KV Cache Budget Allocation for Reasoning Models.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding ReasonAlloc: Hierarchical Decoding-Time KV Cache Budget Allocation for Reasoning Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.702163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.702163Z digest=sha256:8aa616b925c7a6fe0bf2046743d3462eee5553b721550a7e988803c6ca004a1a

Observation ff8b9d3c-845f-47c7-9d77-745d50a3c062 · outbound

This paper cites KIVI: A tuning-free asymmetric 2bit quantization for KV cache.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding KIVI: A tuning-free asymmetric 2bit quantization for KV cache

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.704874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.704874Z digest=sha256:beea7b4da0de8ecfe7d8f1079925bcf3dd497b1cd4bf396df2312d56e6f3cead

Observation 78245182-d1d3-4e19-9e05-613aed85bdd3 · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.707395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.707395Z digest=sha256:5839ebde472172a4ff6f30d7fd9ad6e50d4c63f98898a5d4ed356471117399e9

Observation deb03427-f860-4341-bbe7-00aa29be4af6 · outbound

This paper cites TriAttention: Efficient Long Reasoning with Trigonometric KV Compression.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding TriAttention: Efficient Long Reasoning with Trigonometric KV Compression

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.710201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.710201Z digest=sha256:cf586e34ee0f8e3f75c8780d976d73d9da61d78b122a9a1b728e3b1e0faec8ca

Observation 46f7380e-cb26-4594-8697-62cd46e165ea · outbound

This paper cites Landmark Attention: Random-Access Infinite Context Length for Transformers.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Landmark Attention: Random-Access Infinite Context Length for Transformers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.713173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.713173Z digest=sha256:7e665b2a19a521c306de9e00abf7da95aa3acdea6ffa91c2ceb37347d9ab0c48

Observation 8ac1f2ad-c643-4606-8f66-82abbc90c30e · outbound

This paper cites SALS: Sparse attention in latent space for KV cache compression.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding SALS: Sparse attention in latent space for KV cache compression

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.715929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.715929Z digest=sha256:f8e5416e466043c70333bae0e89dd4ed7ac626180a00d1a18adfca47a2fc3278

Observation 18f71b2b-c7d9-42c9-b8ee-0e1a11a0fbc7 · outbound

This paper cites The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.718361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.718361Z digest=sha256:8145fae021e8896405e566b99394d4ec382caef0a82a7c8dff0e0d4947d0d8ed

Observation 2774982c-d33a-435e-a1c2-e43837eae884 · outbound

This paper cites Qwen3.5: Towards native multimodal agents.https://qwen.ai/blog?id=qwen3.5, 2026.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Qwen3.5: Towards native multimodal agents.https://qwen.ai/blog?id=qwen3.5, 2026

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.720982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.720982Z digest=sha256:67cdd6fac7fe365a0563edcb5901a400fd77571cfb881fbb14fc0030c3d5eb49

Observation 34034361-9ed1-4385-9431-18af41f3191b · outbound

This paper cites SparQ attention: Bandwidth-efficient LLM inference.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding SparQ attention: Bandwidth-efficient LLM inference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.723462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.723462Z digest=sha256:f30a25de8ae62a04301d2ae2a57ba11ab495cb2041ce1646e060577d571d4edc

Observation 6288f059-36a3-481d-a87d-9ad3c4cd5bdd · outbound

This paper cites Eigen attention: Attention in low-rank space for KV cache compression.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Eigen attention: Attention in low-rank space for KV cache compression

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.726565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.726565Z digest=sha256:18ac8f494fc2e0ea92fba12cfb238e8c149e0cc39ff65958fb8872b421a309a6

Observation 09739156-8774-4047-9d85-b6e8cb20a7b8 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Fast Transformer Decoding: One Write-Head is All You Need

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.729328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.729328Z digest=sha256:f2c47f7e5cb5200b8b437cbcc968aa0fa38e8ff74fac37eaa0a7f3b0a36ba3de

Observation 5393c3b5-e078-4c43-8730-3640dc0c7287 · outbound

This paper cites Loki: Low-rank Keys for Efficient Sparse Attention.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Loki: Low-rank Keys for Efficient Sparse Attention

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.732389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.732389Z digest=sha256:787d4f29d1d59cdced716909d8203e47b5421443c495919498624d66ad91e815

Observation ea58ec9e-d60a-44c2-8139-1c3909ba88a2 · outbound

This paper cites LongFlow: Efficient KV Cache Compression for Reasoning Models.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding LongFlow: Efficient KV Cache Compression for Reasoning Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.735297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.735297Z digest=sha256:d305b595448ca01afee9b10a9762e16327de1063a34975f21993e0e8f0acaec4

Observation 2eb77442-e559-441b-916b-5bdf1a8bbaa1 · outbound

This paper cites ShadowKV: KV cache in shadows for high-throughput long-context LLM inference.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding ShadowKV: KV cache in shadows for high-throughput long-context LLM inference

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.738457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.738457Z digest=sha256:f049f07991089dff2b62e2e9a757999f8c31f42ff7cda6238bf2d74d95f1ceed

Observation c6baaa99-e09f-489d-b42e-91aa819d0403 · outbound

This paper cites Quest: Query-aware sparsity for efficient long-context LLM inference.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Quest: Query-aware sparsity for efficient long-context LLM inference

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.741061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.741061Z digest=sha256:c6135b2451c1275a331b4991a2b4489c8483c36a8ac415a1dc8ef25cbeac8324

Observation 02fc6b07-76db-4b02-96cc-313ceb3ed263 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.743571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.743571Z digest=sha256:407b6e3125a0465ccd1212206329d2a686897b388fed901e3a2d4ab4cc9d0c07

Observation 1366120e-6eef-44b1-808d-fdc969bc4e17 · outbound

This paper cites COBS: Cumulant Order Block Sparse Attention.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding COBS: Cumulant Order Block Sparse Attention

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.746726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.746726Z digest=sha256:34870924da29a7ab4fd18cf806dfd13c84a3638a439878d2df06f10c2760b01d

Observation 3e65cea9-3327-4e76-b100-6617e0255937 · outbound

This paper cites A mathematical theory of top-ksparse attention via total variation distance, 2025.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding A mathematical theory of top-ksparse attention via total variation distance, 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.749782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.749782Z digest=sha256:9f106e8bad473c03f953f822aa30c68836aad5a0f60e10e123a87ce0ce74deb2

Observation bbd7d912-9cc9-43cf-afa1-5c54d7176c1d · outbound

This paper cites SPLA: Block sparse plus linear attention for long context modeling.arXiv preprint arXiv:2601.22379, 2026.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding SPLA: Block sparse plus linear attention for long context modeling.arXiv preprint arXiv:2601.22379, 2026

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.752410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.752410Z digest=sha256:fb7ecb51e895cee005989921ce26bd1e3cfe6fcd3893eec86f0605fd6bea44f8

Observation a7b3e8db-b6d5-4784-8d5b-44d9dc530c8c · outbound

This paper cites Prism: Spectral-Aware Block-Sparse Attention.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Prism: Spectral-Aware Block-Sparse Attention

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.755340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.755340Z digest=sha256:7ce41b9a3b0b92b37c4a87c8641e0295b9ad990e863a6ba6f59c60eabb098318

Observation 21c42903-c6bb-47aa-b73f-20a3c4c0d651 · outbound

This paper cites TokenSelect: Efficient long-context inference and length extrapolation for LLMs via dynamic token-level KV cache selection.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding TokenSelect: Efficient long-context inference and length extrapolation for LLMs via dynamic token-level KV cache selection

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.758862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.758862Z digest=sha256:58444d8c03887efecb0afbbb2846f1314b94990d3df44da2a2b8755bea51ac2b

Observation 6ed4e946-e7ea-498f-8aee-42a4d25a38dd · outbound

This paper cites InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.761477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.761477Z digest=sha256:b9d52f4f3cfff635f137ed963c635d18a44a31b1f1944a3a0385ec5d90aee05a

Observation f01c3bbd-793b-43ec-a26d-5f1b1da8b0f5 · outbound

This paper cites Efficient streaming language models with attention sinks.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Efficient streaming language models with attention sinks

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.764973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.764973Z digest=sha256:606128e1efac8a3b0448dbbc0fd648a498a1db5e48dbd2d000d9a535f791c564

Observation b90b59e2-f2af-4706-9a62-2710cc046603 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.767862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.767862Z digest=sha256:c27cbd19a8d32680e965ed8f312e1557813e965789002d577c73150ba74551e8

Observation 60e153d1-c2fd-462a-8230-af37e924c857 · outbound

This paper cites Qwen3 Technical Report.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Qwen3 Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.771041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.771041Z digest=sha256:b2a950f2956634fa29997b4a6f1fa26327efd8717b3f547cb954bc0806d395fa

Observation bf064771-150e-419f-8c46-194fb053990b · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.774088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.774088Z digest=sha256:f892637e99efaca9c35757626e6e785a26edf33432bfb0225c69524bfd7bd028

Observation fb304ee9-3e1a-4e57-9269-b67b116de627 · outbound

This paper cites Adaptive Mass-Segmented KV Compression for Long-Context Reasoning.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Adaptive Mass-Segmented KV Compression for Long-Context Reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.776915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.776915Z digest=sha256:844b19856dbad82817716ad48f13e169d58e2292942c93f2187c79b17ddddd66

Observation 4dc332e9-ffdb-43bb-90e4-b8a915f52fd7 · outbound

This paper cites TidalDecode: Fast and accurate LLM decoding with position persistent sparse attention.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding TidalDecode: Fast and accurate LLM decoding with position persistent sparse attention

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.779925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.779925Z digest=sha256:2d53c2149e95efd071b6dd4385d6a25d05d9183235b2978323a678389e3f24ad

Observation deca8bf0-5b1a-485b-b49e-cecbdb30daac · outbound

This paper cites Self-indexing KVCache: Predicting sparse attention from compressed keys, 2026.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Self-indexing KVCache: Predicting sparse attention from compressed keys, 2026

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.782655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.782655Z digest=sha256:58920d43d6ef222348219299c5f2f5db126cecec679d1b950b01ca2c8d3c4338

Observation 12701088-aa65-445c-b456-6b4edbff2b8d · outbound

This paper cites HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.785240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.785240Z digest=sha256:fc2ceeeab8d7c83c1e017f7a0af73d3e15cc57992b891cdfddca0cf1c7080744

Observation d35f19ad-4206-4fb0-bf15-52471124beb1 · outbound

This paper cites BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.788006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.788006Z digest=sha256:962d707e3062640683104db29c2aada5e6a3b7a2e95de291fa6741f508078a15

Observation 22443527-7156-4090-adfe-56a04d49be16 · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.790902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.790902Z digest=sha256:0eb5c7288204260d59361a7818f6d49672d72cc08bfb550dacd5c420a48718b8

Observation 134df698-5dae-4c79-8d7b-b21fcd450986 · outbound

This paper cites Big bird: Transformers for longer sequences.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Big bird: Transformers for longer sequences

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.793719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.793719Z digest=sha256:e2752425879d34ac51c8b40a5b40d9e6584a86fbff1a665211dc6e82eb5762c0

Observation 28ebee9f-58ba-4a69-b2f6-6cc784a6b235 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding GLM-5: from Vibe Coding to Agentic Engineering

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.796378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.796378Z digest=sha256:861b9f69ac7d7ccf90e20eedf7ad4dca5577d8fcf840830d3f930eb9e475883e

Observation e1473abf-d04b-484b-a419-926e0cb80548 · outbound

This paper cites PQCache: Product Quantization-based KVCache for Long Context LLM Inference.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.799232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.799232Z digest=sha256:a1dc89ed459458f545c5f13d83b443395fada68d17eb1a6490d44faa8b86df96

Observation 3b8fbe2d-954a-4314-b91a-3424b16c5da2 · outbound

This paper cites LazyEviction: Lagged KV eviction with attention pattern observation for efficient long reasoning.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding LazyEviction: Lagged KV eviction with attention pattern observation for efficient long reasoning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.802040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.802040Z digest=sha256:038131625cd2da362de2b8afc0642a8ee045b18565a271947823ec6107f7e504

Observation 26379618-dee9-499f-bbb5-d9352f94721a · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.804916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.804916Z digest=sha256:654bd71445935916cb685da3542ffc517da716c2cf8762245ea18108da23784e

Observation 9bfb584b-155d-41ad-8434-49a2871f16f4 · outbound

This paper cites Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.807382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.807382Z digest=sha256:8bb24ff945ee0f55d8a491ba8767e96aefdb13cf7dc9c6ecedffbeec42275a54

Observation aabcb895-f888-4f90-877b-1a621e302597 · outbound

This paper cites OjaKV: Context-Aware Online Low-Rank KV Cache Compression.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding OjaKV: Context-Aware Online Low-Rank KV Cache Compression

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-31T11:49:11.810189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.810189Z digest=sha256:5400ac4837334a9aeddc2b4dcbc6e689613e9098046f15a3dfddab91caccf273

Observation 8c14fd66-b31c-4811-b6f9-ee63def766db · outbound

This paper cites Fast KV Compaction via Attention Matching.

LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding Fast KV Compaction via Attention Matching

Reference 77

Resolution
malformed identifier
no resolver link, observed 2026-07-31T11:49:11.813142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:49:11.813142Z digest=sha256:554e9e1a081067ae9fd5fa60006f402ff951519330d8f5a6593746938320419a

Pith citing papers

No inbound Pith citation observations are available.