Pith. sign in

Paper Citation Record · LEDGER

SCBench: A KV Cache-Centric Analysis of Long-Context Methods

As of 15 August 2026, this Paper Citation Record lists 100 of 108 outbound references and 18 inbound Pith citation observations for arXiv:2412.10319.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10319 v2

Coverage vector

measured 100 of 108 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:14:06.284069Z

measured 118 of 118 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:38:46.746093Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 108 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved81
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 18bfbac8-1a26-4962-a053-34f6e6c8b3d8 · outbound

This paper cites Star Attention: Efficient LLM Inference over Long Sequences.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Star Attention: Efficient LLM Inference over Long Sequences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.493385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.493385Z digest=sha256:20ab21d51077e6c6fd8984cfe4a5a81c1c70193c4c981715974d577e7408bfda

Observation fe095b35-d5b3-4ae7-9bb4-449b65ffcff4 · outbound

This paper cites Chan, Biao Zhang, Ankesh Anand, Zaheer Abbas, Azade Nova, John D Co-Reyes, Eric Chu, Feryal Behbahani, Aleksandra Faust, and Hugo Larochelle.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Chan, Biao Zhang, Ankesh Anand, Zaheer Abbas, Azade Nova, John D Co-Reyes, Eric Chu, Feryal Behbahani, Aleksandra Faust, and Hugo Larochelle

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.503321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.503321Z digest=sha256:b702b6a79b5637df90e89c1ffb6af71cd0d620fc01f105646c5bfc1bf90f8e27

Observation f3d474b6-4284-4ba3-956c-71b50c3013e0 · outbound

This paper cites GQA : Training generalized multi-query transformer models from multi-head checkpoints.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods GQA : Training generalized multi-query transformer models from multi-head checkpoints

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.508289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.508289Z digest=sha256:b83ff064cc5b91890c5e8cccb7b2d0e5e034689fef3691f143b00b571c7e4789

Observation 7e4e5e93-2df0-4c68-8909-a77e8e23a17c · outbound

This paper cites Just read twice: closing the recall gap for recurrent language models.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Just read twice: closing the recall gap for recurrent language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.513213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.513213Z digest=sha256:9bdaa6db3660db4c41640f2b10bfa8f008e28e0c95281ea3f528340aaf31ccc2

Observation 778e19cc-45a0-4def-ae16-b81f1c555fad · outbound

This paper cites Prompt caching.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Prompt caching

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.525308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.525308Z digest=sha256:bff66e70b3f4e9bf637b16badfe96d565172dd3dc45bf5c1d39a42fc707a9137

Observation 80301961-0e49-4895-b67b-ec1c15b0f668 · outbound

This paper cites MT -bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods MT -bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.529601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.529601Z digest=sha256:85be678f4b9c4cea3d3b555c0aec85fc23868727d0857a54abb102bf513beac4

Observation 2d8fde69-847a-4c0c-848e-abd4a8714e7a · outbound

This paper cites L ong B ench: A bilingual, multitask benchmark for long context understanding.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods L ong B ench: A bilingual, multitask benchmark for long context understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.534148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.534148Z digest=sha256:bcc6fecb517ae6ce1c95f4828aaf013e776148a0c51d5cbf2ecb9280008f8509

Observation 17b9cd6d-5fb8-4c25-84d3-7230803902e4 · outbound

This paper cites Codeplan: Repository-level coding using llms and planning.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Codeplan: Repository-level coding using llms and planning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.538651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.538651Z digest=sha256:9701d6fc608141578788ed3e558860d37d58ea61edf85578ada3148b47716999

Observation aa38c05d-75a8-4446-a94f-03b638986935 · outbound

This paper cites Longformer: The Long-Document Transformer.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Longformer: The Long-Document Transformer

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.545976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.545976Z digest=sha256:448ecdf0a05a20bbc2da0e796ba4426c410ff5f738575f793e835a3cc36817be

Observation 0ed946b5-9a6f-4b91-99ed-b776ff242134 · outbound

This paper cites Peek across: Improving multi-document modeling via cross-document question-answering.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Peek across: Improving multi-document modeling via cross-document question-answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.552095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.552095Z digest=sha256:94f30965bfc5b188e7c6227574874cfce6632068456cba1f18b7305412a70752

Observation 73932485-a9cf-43e2-b151-a82d3efaa9b2 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.559074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.559074Z digest=sha256:2ccba3a82067376334914ce81ace28654872fd1b0df54a602a4b5300ba83a8f8

Observation 3c7b7770-b0a0-42dd-8a93-46b381e2c2c9 · outbound

This paper cites Magic PIG : LSH sampling for efficient LLM generation.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Magic PIG : LSH sampling for efficient LLM generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.569975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.569975Z digest=sha256:c8d3e35a12c77360a6669b19903d4a57fd1af3a5c140628109a5cefe25981a0c

Observation d706461c-637b-47df-aaff-cc0366b35892 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Generating Long Sequences with Sparse Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.581212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.581212Z digest=sha256:ddd333ad00ea028e83f4b86a4ca6f24e9ada948b88dfa2f397460f006ce43882

Observation 7bf95e5c-4ba6-4af8-a9e8-a30482e60893 · outbound

This paper cites Prompt caching with claude.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Prompt caching with claude

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.587485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.587485Z digest=sha256:60f21ab623645485fece965619fe4a523bec9eceb82dfee89efc8077df377bf8

Observation 10688121-8db5-442a-8783-bade2f68d766 · outbound

This paper cites CORM: Cache Optimization with Recent Message for Large Language Model Inference.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods CORM: Cache Optimization with Recent Message for Large Language Model Inference

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.592172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.592172Z digest=sha256:00275944393bb92ccf98dae5a4f4755f049684b214113f6f3038c11861247f0a

Observation 40383359-a7e9-456b-8b68-d751478e04a3 · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.597645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.597645Z digest=sha256:18a2d42d06217a3753d81790f3a7bd94f763d99803fe03b0cc6b13fad239da0c

Observation 177686b2-3b3b-456e-8163-95a03c989e42 · outbound

This paper cites Transformers are SSM s: Generalized models and efficient algorithms through structured state space duality.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Transformers are SSM s: Generalized models and efficient algorithms through structured state space duality

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.602037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.602037Z digest=sha256:2f231fee8388088c07877f6527e4ca3325ce9def0b65e7a52c537461d3d026d3

Observation 271fafd0-d81b-455e-a805-d2108fa7a407 · outbound

This paper cites How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.607282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.607282Z digest=sha256:d6eb234ed69546166140046d0b32af9e0c9641d0c3c72c1cc2fa7cef7e7fd2bd

Observation ec997631-019a-481a-a033-8cbb04b67af7 · outbound

This paper cites HashAttention: Semantic Sparsity for Faster Inference.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods HashAttention: Semantic Sparsity for Faster Inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.613937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.613937Z digest=sha256:167495682f8bd5c2fb0161d7fe7e517c2775f0b67418877482e15982f1b697e0

Observation 798bc6bc-32cf-45df-979d-45d6da57d34b · outbound

This paper cites Domeccleston/sharegpt: Easily share permanent links to chatgpt conversations with your friends.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Domeccleston/sharegpt: Easily share permanent links to chatgpt conversations with your friends

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.619168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.619168Z digest=sha256:1c3db7525b44af8004c0ce912b40a2ed94ae93970756093a4e7ab73066a80f01

Observation 08aea4bc-d5c9-464e-a316-4f81c94c17c1 · outbound

This paper cites The Llama 3 Herd of Models.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.623861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.623861Z digest=sha256:b52d010be40eae8c99ba512a4cb3ebc281e9339e25b46a6ba1eeb35a965eedda

Observation 26e6dfca-560a-45fb-a6fc-12cb7371dd4f · outbound

This paper cites Finch, James D.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Finch, James D

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.629895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.629895Z digest=sha256:18300af7a65b62475fd9887d45c890ba32df9666ab0c4fa1df5e4157cc757fed

Observation d15a8e5f-af64-4ccf-9c5d-0fad4fefbf19 · outbound

This paper cites Moa: Mixture of sparse attention for automatic large language model compression.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Moa: Mixture of sparse attention for automatic large language model compression

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.635486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.635486Z digest=sha256:69604f0fe1629a0df1364db0c6b08d2ef571d0136af9aa982e710eb1f81aca70

Observation f1c321da-c372-4aae-a834-93fbed9de7a6 · outbound

This paper cites Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.640877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.640877Z digest=sha256:5554f2d6b4d0c6c8a079f5dc8a85f384b73ef0dc95ed607d7684f8df607a5b40

Observation 511a06ce-6c18-463f-930b-602a4d9d3434 · outbound

This paper cites Model tells you what to discard: Adaptive kv cache compression for llms.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Model tells you what to discard: Adaptive kv cache compression for llms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.647156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.647156Z digest=sha256:4dd147b3ae591450d61cfcebfded136f4c0d6e1069e1582499c420ae4abac32a

Observation d30ef9fe-2a14-4169-9ae6-8e330e7e30e3 · outbound

This paper cites Context caching.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Context caching

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.652480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.652480Z digest=sha256:e0755b8fc714af901f5e60776154a582d98bd387a522944ecd011d6a86bce26b

Observation 1a7dda80-cd62-457f-ac49-5dd93d54be83 · outbound

This paper cites Prompt cache: Modular attention reuse for low-latency inference.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Prompt cache: Modular attention reuse for low-latency inference

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.657245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.657245Z digest=sha256:042474128e98d8c553e3bf1bd7dcff6a71313ed07a18d9d8e786dbe6b87e1698

Observation 1c18c550-1c1e-4c77-9fb3-a4e1570cf463 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.665083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.665083Z digest=sha256:4fb23fab1661666af756cec74b4a772829ad50f95a9b180e530e0f7d63c49b07

Observation fdb52d04-7c0a-44a2-8ffb-3a156f50195b · outbound

This paper cites Llama-3 8b instruct gradient 4194k (v0.1), 2024.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Llama-3 8b instruct gradient 4194k (v0.1), 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.672513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.672513Z digest=sha256:ff8bbefd9842df31620b7024e66eca39d2cb02d9344fcb2209afa68c8be5dd91

Observation 988a1d49-9e67-410c-bc1b-fb2154793c53 · outbound

This paper cites Mamba: Linear-time sequence modeling with selective state spaces.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Mamba: Linear-time sequence modeling with selective state spaces

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.678164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.678164Z digest=sha256:3a61b312a198ed13c42369c1449d7c7b96cdfe1e1ac70994bd6f0098156ee82e

Observation 2840c81b-80dc-4b34-87ca-5cc2f189b002 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Efficiently modeling long sequences with structured state spaces

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.683448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.683448Z digest=sha256:dbb313d37fb0f933d0ab967a803fff7c41191d33a015c387deafc8bf7f10a7d6

Observation 39ff0f01-ef68-4669-9e3f-bfa802775362 · outbound

This paper cites LM -infinite: Zero-shot extreme length generalization for large language models.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods LM -infinite: Zero-shot extreme length generalization for large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.690698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.690698Z digest=sha256:ff875ebc628710724ed937db16d1957fc0296d3ba00dd8b07f43048f9bc2461c

Observation 25bf92da-2bfa-4482-b0fe-fd924d4b3a86 · outbound

This paper cites Structured Prompting: Scaling In-Context Learning to 1,000 Examples.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Structured Prompting: Scaling In-Context Learning to 1,000 Examples

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.700159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.700159Z digest=sha256:afed92f6de23a2ebcf7e9d252dc002185a84c01253acd7fb6532280dd39e957c

Observation 25410f7b-8665-4a5a-b8cb-112703d47632 · outbound

This paper cites Liquid structural state-space models.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Liquid structural state-space models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.720243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.720243Z digest=sha256:1523b2265787a6de045b9f34ad7db1c2142789c162708ca54eeea2d1362e8d99

Observation 42757b06-926f-4a90-8808-b624a3571c08 · outbound

This paper cites Block transformer: Global-to-local language modeling for fast inference.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Block transformer: Global-to-local language modeling for fast inference

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.735108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.735108Z digest=sha256:8a9bec27c5d8be615c779ef85d0d5d98d4cb7b790a55f450b83358915df2584e

Observation 23d9cb6d-f612-4719-b520-b8d8907a0ddc · outbound

This paper cites Squeezed attention: Accelerating long context length llm inference.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Squeezed attention: Accelerating long context length llm inference

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.745921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.745921Z digest=sha256:46f50b6e51cf71fab7fffe974af7d36a8af69c3a15339fdd2a8436ca56506636

Observation b651caa1-ee07-4f2a-9acd-2c460a63bf33 · outbound

This paper cites RULER : What s the real context size of your long-context language models? In First Conference on Language Modeling, 2024.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods RULER : What s the real context size of your long-context language models? In First Conference on Language Modeling, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.753247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.753247Z digest=sha256:8ac876aa24576926700a84012e5dd43e67b553d726d913f1472171af75f949a2

Observation b41d645b-fac8-422c-a97e-7963cec17c38 · outbound

This paper cites Kakade, and eran malach.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Kakade, and eran malach

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.759270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.759270Z digest=sha256:6130c9d146577cc42ebb5a71de9407da02521ad21c0cdc5672d8b492e1ad5e36

Observation 9cef39e2-c13b-4ca8-81e8-75edb810893e · outbound

This paper cites LLML ingua: Compressing prompts for accelerated inference of large language models.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods LLML ingua: Compressing prompts for accelerated inference of large language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.766023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.766023Z digest=sha256:0c67afe0d7714a356225e19e08ef0bcadbdaf299c4cbcf7817b07036c8633ea7

Observation 8bb5c15d-81cb-48e1-91d0-6474f38e97f7 · outbound

This paper cites Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.783041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.783041Z digest=sha256:0decbae9dc5e60dfe9d6a859942968bfb7f5115a27572e070b579546aec7d4b8

Observation 8d3ddb03-18bc-483b-b4e9-cfdf9b52352b · outbound

This paper cites SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.792099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.792099Z digest=sha256:dbbde4f4b3ef3f9c0fedf3ff9f5f271372578a0fbd6292758c72f723691561cf

Observation d56d8df0-3248-4452-b84f-35c3115b04bd · outbound

This paper cites RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.801823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.801823Z digest=sha256:2aa073f6b94c0adfa30b368fb163c98ab1588411d8e19f5e3eafe6b3ff05568f

Observation c69b9e64-bd9a-49ad-892b-78695d0c61ba · outbound

This paper cites Hydragen: High-Throughput LLM Inference with Shared Prefixes.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.806863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.806863Z digest=sha256:6da4b5f77acaad37b9a6954aece4304b8aa2180cee04cea14cb2a6e4d4a77e8d

Observation 4c4f7554-03f9-4849-a93e-f6ca12eb31bb · outbound

This paper cites Needle in a haystack - pressure testing llms, 2023.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Needle in a haystack - pressure testing llms, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.815247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.815247Z digest=sha256:75231b9ebd5333131259c3992744fbd455ee6b7bb5926f85bed45a311a106754

Observation 9d8f3ddb-240e-4c67-87e2-8732c273e502 · outbound

This paper cites MT -eval: A multi-turn capabilities evaluation benchmark for large language models.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods MT -eval: A multi-turn capabilities evaluation benchmark for large language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.825768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.825768Z digest=sha256:47784cb61c819f384c7c0710792667b8000ead19529aae26f1b904ac3cb0d804

Observation b17a2e77-4efd-492c-b86d-abc664291541 · outbound

This paper cites Flexprefill: A context-aware sparse attention mechanism for efficient long-sequence inference.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Flexprefill: A context-aware sparse attention mechanism for efficient long-sequence inference

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.844002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.844002Z digest=sha256:aae159852278548551787090355d30aed16774755c37b2e7b7d71217b8760fae

Observation 91252970-082f-436a-b988-b8afdb9de1d0 · outbound

This paper cites an unresolved cited work.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.849167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.849167Z digest=sha256:e6d4452e9a746078c560dfed32fcb730f4ec150ead3fc4ba89289c0962addf48

Observation 22c5b984-f9f3-48c7-94db-7e96b6bac95e · outbound

This paper cites Long-context LLMs Struggle with Long In-context Learning.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Long-context LLMs Struggle with Long In-context Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.857578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.857578Z digest=sha256:5e76a5d3d150bccea64bf14cd6c95d435fb8ff99495c72204b50d4496a5e2249

Observation 1e06a6e8-1e1f-48d2-9422-5b606cd61a2f · outbound

This paper cites Alpacaeval: An automatic evaluator of instruction-following models.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Alpacaeval: An automatic evaluator of instruction-following models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.863375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.863375Z digest=sha256:bb02bb0c0dffeed74910ecea32f74abe1cd242df86047c98221f7ba6ef42fa44

Observation d02dc9c8-06c0-4244-9eb8-390e0182d5dd · outbound

This paper cites Compressing context to enhance inference efficiency of large language models.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Compressing context to enhance inference efficiency of large language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.869749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.869749Z digest=sha256:abb2042dfb91b4550bada0e16430e4591404856d6021766e0465667c76698f79

Observation 6a8e5304-f31f-400a-99ea-92ac8fbcce98 · outbound

This paper cites Snap KV : LLM knows what you are looking for before generation.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Snap KV : LLM knows what you are looking for before generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.874814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.874814Z digest=sha256:6c83ea943a8a6b653aa2d00b8e579e2e21be949a66d971fa846f3223d3b3bd8c

Observation 658faf26-dac0-465a-b45c-4a31b19df96a · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Jamba: A Hybrid Transformer-Mamba Language Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.881945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.881945Z digest=sha256:e40d84489dcee2e6ecda2cca72898c797cfeef10ff507d3c11f713968232f38c

Observation c225df22-e0e7-48de-b586-3e2de20edeba · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.889712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.889712Z digest=sha256:5b1c647cc733a4f4442927a781607196972b8d23424c32f9175c0188a7d2854b

Observation 527ca086-9439-4b0c-a600-1840b4307bc6 · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.896609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.896609Z digest=sha256:79e79ab23db7c1072c156831dbd4db49bcc0f728d198ae494ddf0df244e42e7f

Observation b80d5d13-cd3d-472b-90b4-69eccce7945f · outbound

This paper cites RepoQA: Evaluating Long Context Code Understanding.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods RepoQA: Evaluating Long Context Code Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.900988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.900988Z digest=sha256:6d7ab4a5e03125e6100ef7879399f1e5e272a1086c4cbf1b5176eda8ca10feb8

Observation 528e4dd8-7aea-4814-8304-7e4008d8eec8 · outbound

This paper cites Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.908218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.908218Z digest=sha256:7ba27c4bfbca3847ba7915164a2c468a93af1e39f82a784ec53461b5efd8d10b

Observation f85d36f6-f265-4014-8850-9183cf7a44dc · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.915658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.915658Z digest=sha256:0cad05ba94d5f8e844077a94bf381ff6b62f6f2d9e4562a8007d47b6e600be6d

Observation 889c9ea3-8860-4e6d-9b37-0aba53f14722 · outbound

This paper cites KIVI : A tuning-free asymmetric 2bit quantization for KV cache.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods KIVI : A tuning-free asymmetric 2bit quantization for KV cache

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.925832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.925832Z digest=sha256:9986e8d080838bb8f343828b805998fde5eed1d3e2f5ebbf49858c5988414bc2

Observation 198d9c45-aebe-4cf7-94b8-19a502c710ed · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.931089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.931089Z digest=sha256:514eb0670074f19d53df5a313fa05b7456df8f8cb65f853362a6648580daaab5

Observation e1793aa7-df82-4c94-bfbb-173c9766c5d5 · outbound

This paper cites Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.937098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.937098Z digest=sha256:b22742b62cc13f4d502ee39d303813edb381f902cda9dd03bc55954f1597dadb

Observation 9c22a376-d722-4ce0-a32e-563170ac771f · outbound

This paper cites Dynamic memory compression: Retrofitting LLM s for accelerated inference.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Dynamic memory compression: Retrofitting LLM s for accelerated inference

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.817249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:05.941952Z digest=sha256:709717f311fcdc7311636dcea038704c71587d4ed0c0689f43e2ebc05baaaffa

Observation 2deae104-22b9-4784-98a1-a821e77acbf0 · outbound

This paper cites In-context Learning and Induction Heads.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods In-context Learning and Induction Heads

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.947431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.947431Z digest=sha256:8175f10ab7d3dcb1ebad564ad79a6357b7262e2b3c1bca242d6458ab247bf107

Observation 17ac8824-9b32-44bc-9055-29430d43b639 · outbound

This paper cites Learning to reason with llms, 2024 a.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Learning to reason with llms, 2024 a

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.955388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.955388Z digest=sha256:0f1d4f44a7381f1568b490386f8d56d1cc954dff1c1e63cb9b4c87d68dbd2d76

Observation 2969f56b-40b0-46e9-ba54-2f456969c045 · outbound

This paper cites Prompt caching in the api.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Prompt caching in the api

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.768455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:05.964250Z digest=sha256:8678bad1f92026b2208d18b24bebeb39a9f42ffb83ea61f73ad81528d6f83b46

Observation 440829c8-ba64-465d-80c8-b0e52d2a7e6c · outbound

This paper cites Vicky Zhao, Lili Qiu, and Dongmei Zhang.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Vicky Zhao, Lili Qiu, and Dongmei Zhang

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.971476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.971476Z digest=sha256:b63e525c920359baf53a12bfdf51e7d820ee54ebbacb3456f00dd06e4368de97

Observation b4462b04-5a87-40b8-9c9f-3f69d08ad7f7 · outbound

This paper cites O'Brien, Carrie J.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods O'Brien, Carrie J

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.726914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:05.979232Z digest=sha256:c6c9e1d8471eca6f1880b59ab899ef5b894c1929259dae54907de26b0139e8f6

Observation 8943bcd6-a291-471e-9472-6c713bc89683 · outbound

This paper cites Wind, Stanis aw Wo \'z niak, Zhenyuan Zhang, Qinghua Zhou, Jian Zhu, and Rui-Jie Zhu.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Wind, Stanis aw Wo \'z niak, Zhenyuan Zhang, Qinghua Zhou, Jian Zhu, and Rui-Jie Zhu

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.705590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:05.984994Z digest=sha256:a4ea9aebce4962a1f88056351d2f4a3da11fd502baa4ea0044f32005f62f0df0

Observation f24dec16-aa6e-42a6-a9aa-ed107cb8f74e · outbound

This paper cites Mooncake: Trading more storage for less computation a KVCache-centric architecture for serving LLM chatbot.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Mooncake: Trading more storage for less computation a KVCache-centric architecture for serving LLM chatbot

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.619077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:05.992398Z digest=sha256:b509de1c47636e69188e6e34bb0f5c1b465afe3743bb7269bb1d57f8abcb12c6

Observation 3ff945e3-5e2c-469c-bde2-1663884c73e6 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.996739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.996739Z digest=sha256:bb75275357c195322231c6138623fe4fcb4feddd00a0e6b94eafba2a9d277ba5

Observation a65b7ef8-78f9-455b-ab0e-e1f3982b3294 · outbound

This paper cites Samba: Simple hybrid state space models for efficient unlimited context language modeling.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Samba: Simple hybrid state space models for efficient unlimited context language modeling

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.006947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.006947Z digest=sha256:1449a391432e23527ffb0b24fc0277f1d0f5e198d71e72ac6e3755e1a30a1c92

Observation 494b5aa1-fe8b-4634-bfdb-73fd36a4771f · outbound

This paper cites Sparq attention: Bandwidth-efficient LLM inference.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Sparq attention: Bandwidth-efficient LLM inference

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.485959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:06.011708Z digest=sha256:7f8885be25ec29e23962cd431df35c46f5f4c6f677e95d80b2b41d7f1be21cd1

Observation 4587d32f-fbb7-401d-9932-70542f847662 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Fast Transformer Decoding: One Write-Head is All You Need

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.017822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.017822Z digest=sha256:95585183768d24447c7a1bc955de4031901e2a93d84400711d0ed42099ff81ad

Observation 222e6612-2b0c-494e-8b0d-08ba82ccf2bf · outbound

This paper cites F lex G en: High-throughput generative inference of large language models with a single GPU.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods F lex G en: High-throughput generative inference of large language models with a single GPU

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.437291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:06.023125Z digest=sha256:9018aa545583c2a75f7a3fc50382abc57fff93c75f21c564f7a729bdb62a108c

Observation b7c57564-bfa8-40c4-a9eb-b1c323735f24 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.031774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.031774Z digest=sha256:c383bcf53be3d5e9ef482a38f16fe13b5dfbf3c713e53edfc5a7229951664a48

Observation ab500c6e-b0ee-429e-a6cc-3c8966cf4871 · outbound

This paper cites Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Beyond the imitation game: Quantifying and extrapolating the capabilities of language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.419525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:06.045641Z digest=sha256:c6a34ad43d5b37e2be87a5d61b376969e30a10ccfdf79e6f8d74642f41ffbd06

Observation 26a87e2a-968c-4121-9028-452a777bf42d · outbound

This paper cites ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.050044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.050044Z digest=sha256:7324d765006cd69d950c432e64a1bb054ce3e83db8b67c50619a6f3f22910458

Observation 05320c79-bdd3-476b-bc40-cd0afe5fecd7 · outbound

This paper cites Triforce: Lossless acceleration of long sequence generation with hierarchical speculative decoding.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Triforce: Lossless acceleration of long sequence generation with hierarchical speculative decoding

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.388689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:06.057656Z digest=sha256:8cb47dc8355604f78fddb085c715296a914eb32919c83fc34162a674f6542167

Observation fb10823b-db14-4a4f-a061-1da966191dec · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Retentive Network: A Successor to Transformer for Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.064028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.064028Z digest=sha256:ea9b506dc690982ec4a359e4b46d59efb3ecc071420a21dc29270e24e92f2dca

Observation 6adc473c-32e0-46bf-9f20-d770813298b6 · outbound

This paper cites You only cache once: Decoder-decoder architectures for language models.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods You only cache once: Decoder-decoder architectures for language models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.364731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:06.071191Z digest=sha256:377ba2643140a2bbdf1a2a0a169814897f63e0611d76685611130932afdbfacf

Observation 569071c9-4b33-432f-900c-ea4174e6cbd3 · outbound

This paper cites QUEST : Query-aware sparsity for efficient long-context LLM inference.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods QUEST : Query-aware sparsity for efficient long-context LLM inference

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.080839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.080839Z digest=sha256:c97332b911eff701f6def8aad546a7814c506ee7382600a8ca173da88e19aa5b

Observation bdfa4aaa-db5f-4930-88ae-08d1bfb603ac · outbound

This paper cites Codestral mamba, 2024.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Codestral mamba, 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.312561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:06.089393Z digest=sha256:202c033e39b43b21dc7c125c7a1d8cbd7aa97c15c9bdf0fc738daee2b5e46c75

Observation c5483b3e-6342-483b-85f1-8874d6af226b · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Qwen2.5: A party of foundation models, 2024

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.285181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:06.098248Z digest=sha256:281f9e8a03943085db72921f0a5ce5c5e00f8879228e8d575218f13a2accc0f2

Observation 6686e7be-bda2-46e4-8e3c-7538d5f4cd44 · outbound

This paper cites an unresolved cited work.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.104081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.104081Z digest=sha256:ed751bb6e437fcf2ba9711db8e5fcb01e97f7c696ea711221f6d26a22e0292e8

Observation d2efaf0b-cc4c-46aa-893e-ef2e967eb006 · outbound

This paper cites Automatic prefix caching.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Automatic prefix caching

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.254447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:06.110138Z digest=sha256:e85752ae9dc5f2d6e8dc9c01f89cffac2152ce8c724d0c5f6244a8a4e6b9603b

Observation 2dece6c6-3f32-4f9d-8fa0-3a181e4607b1 · outbound

This paper cites Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.121591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.121591Z digest=sha256:46eb6cdaeaee80509b8db4c37c9ba4a1988da716dabe9be723f0f2d3fdb6126f

Observation 4524286c-7793-4a71-b162-50af9049f9d4 · outbound

This paper cites An Empirical Study of Mamba-based Language Models.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods An Empirical Study of Mamba-based Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.129767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.129767Z digest=sha256:eb7d27c71dc215a3d43cf871e89c7ee5b7aaefa6dd763541e2b743539f6f5a0f

Observation c9d6a03d-a463-4065-aa62-e37f107167b5 · outbound

This paper cites MINT : Evaluating LLM s in multi-turn interaction with tools and language feedback.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods MINT : Evaluating LLM s in multi-turn interaction with tools and language feedback

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.234551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:06.136052Z digest=sha256:c7d5fb58a9de0ee87afbf52fad41207dddf220dd30f0fb3bc941ec82f73f6938

Observation 96ffaab2-c48c-4ab8-af55-f441d82bd9cf · outbound

This paper cites Efficient streaming language models with attention sinks.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Efficient streaming language models with attention sinks

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.145396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.145396Z digest=sha256:8cef4ae13155755e6f1364b72fbd9f3b2dfc31987a904c378472bb918430119c

Observation 4fe1dbbf-88fe-4b1f-8e35-874bba0b8c44 · outbound

This paper cites Duoattention: Efficient long-context LLM inference with retrieval and streaming heads.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Duoattention: Efficient long-context LLM inference with retrieval and streaming heads

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.178682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:06.152536Z digest=sha256:2744cc77993d7959e68f2a36056cb042ca82bf517501c153d9c578315ee02e18

Observation ff884504-f70b-4e81-b39e-a6ec9928ad9b · outbound

This paper cites LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.161696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.161696Z digest=sha256:335a15091a2594c58861c656bce963097973be2200e3a002bab5de0b2d496bf8

Observation 353866e5-2091-459b-b15f-85bdcc7a0be0 · outbound

This paper cites CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.173124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.173124Z digest=sha256:ce48659423164c3184f26a9ff02b3cb522f4de075eaabd38eac86b60abceed8d

Observation 68bef232-e66d-465d-97c5-ca891e8d482a · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Tree of thoughts: Deliberate problem solving with large language models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.189873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.189873Z digest=sha256:f636aa7b1c10d97a10c6f5ac562270bb7757e319536d1e94c94aa52ece834c2e

Observation 0f393cad-5ce7-45e6-a2d5-9b7e224cf771 · outbound

This paper cites Cascade inference: Memory bandwidth efficient shared prefix batch decoding, 2024.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Cascade inference: Memory bandwidth efficient shared prefix batch decoding, 2024

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.117891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:06.215053Z digest=sha256:c7028447fdeff2d3d08dd628aac40cfb1f4e64b124385273cac7625780b76b1b

Observation 9a1d06c3-cacf-4c4b-a352-842dc6d3a7de · outbound

This paper cites HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.222150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.222150Z digest=sha256:5a26d4199403d29747dd01f261e38cf729b19ac7f059b753022ecd29150cdc4d

Observation 8468a9e5-3470-401c-9df6-18a13595b0e3 · outbound

This paper cites Dialogue-based relation extraction.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Dialogue-based relation extraction

Reference 96

Resolution
verified exact
doi, observed 2026-08-11T16:14:06.420317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:06.248401Z digest=sha256:b77c0137b8e06447bf6a9ce5a1fc92fb9b4f2a34534e5740721c843e823e3f20

Observation fe36ebce-ce0a-402f-b777-c622d49dfa7b · outbound

This paper cites KV cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods KV cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.252705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.252705Z digest=sha256:10c82843728485cd0564d0b60359d88918d6fc7b87884a5a771e9d07fd581fee

Observation 11bb78d4-17cc-4a5f-97d9-7ff4be83815c · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.258692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.258692Z digest=sha256:34e0ea06c4cd7b00db5b73978f9d7a32128463fff9cd2bdfab9d2aa899ed1253

Observation 05cdc766-3668-4039-ab4d-41075a5500cc · outbound

This paper cites Reddi, and Sanjiv Kumar.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Reddi, and Sanjiv Kumar

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.081543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:06.269118Z digest=sha256:3257cdc7daff01e286a1a389a2cc9245ea7fb2f6cded7c9585dc2cfc7b0ce201

Observation 20c3e22f-3acc-4b9c-8b39-c99bea5ed461 · outbound

This paper cites Infinitebench: Extending long context evaluation beyond 100 K tokens.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Infinitebench: Extending long context evaluation beyond 100 K tokens

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:14:08.060930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T16:14:06.279971Z digest=sha256:b7c591baa8c64b97298060c508874583f4f612e0f97f1d628a3156a184c20d37

Observation f847b6cc-af3e-4ab0-ad3d-f1195fb45c2b · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:06.284069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:06.284069Z digest=sha256:e756c9b04585af10d259b128776ed36769b5f9c4a7fc7b5eee0be91680af5be1

Pith citing papers

Observation 7b080277-58e4-4e45-8560-57bac9655726 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:46.746093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:46.746093Z digest=sha256:d49113277ce022630133cdaba93e4a7efdceb32e40ba17ffad19489b07692051

Observation ef36e9a9-22d8-4947-a948-1d7fb234a758 · inbound

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models cites this paper.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.265626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.265626Z digest=sha256:55be45bb3ca7b9dee0e2ca992ad7dd4121c7ecd70ddb8871bff34fdf518bc6ed

Observation 73f6391c-4166-407d-973e-5850a203fe1f · inbound

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation cites this paper.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.722512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.722512Z digest=sha256:ba63a8d1efe414957b5dc218ce86dfd33edde462571118b116fb018a09831f97

Observation 59b9fc34-209f-4705-a7f5-c96c010a895c · inbound

A Survey of Scaling in Large Language Model Reasoning cites this paper.

A Survey of Scaling in Large Language Model Reasoning SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:22:09.420903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T21:20:07.238992Z digest=sha256:c6220fd57ee3518db9cf11cf034f3f205201dd67fc12dc0c5e0dda4a5674b115

Observation 1d7ae517-d409-4357-b5c4-6d3b18cdd479 · inbound

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning cites this paper.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.642988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.642988Z digest=sha256:11d709254fb712f1ba9c5793886606ffe66f9917b81a0edbaa853d8faa3fd5fa

Observation e5d9f31e-b613-4a3c-8b4c-a3c8fbe621df · inbound

Lag-Relative Sparse Attention In Long Context Training cites this paper.

Lag-Relative Sparse Attention In Long Context Training SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:10:46.759948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:10:46.759948Z digest=sha256:154159f7bb065a1e4b3fe057edc36f5097f605d4b80910f49dd0ea1a659e9727

Observation a7cc8d2f-75ec-4ffa-8fc0-5f5277e9658c · inbound

ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing cites this paper.

ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:17:00.821916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-19T03:14:05.109011Z digest=sha256:97bb6d3d16db87274576bbf7fe44ed00665c9a17547e2e78e91469a56e8f609e

Observation 6805502c-372d-4a5c-9805-371e773c449f · inbound

IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs cites this paper.

IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:51:01.223918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:46:43.682254Z digest=sha256:aa04bb49af378b855e85591d82fd471d0803f755a8a4d5385a21b861f4821718

Observation dd591b7a-ea57-429c-adbd-4eb123ac75b4 · inbound

Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM Serving cites this paper.

Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM Serving SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:17:37.454956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T08:15:45.474618Z digest=sha256:346cd55443c21e747ef13f44ae0c87f438ecb0fe2d0aa736258fac4c603f46dd

Observation 2cefd318-562d-4ea8-802c-06367d0fedfc · inbound

PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training cites this paper.

PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-09T21:43:30.254592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-09T21:42:16.848186Z digest=sha256:260aa76f81879a91d9f9f3cfccd9a2e6427922136b951649cc4ca3204f8a162e

Observation ab6cb96e-8163-4cda-b3d5-5d1bab322634 · inbound

PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training cites this paper.

PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 131

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:31:06.204775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-09T21:42:16.848186Z digest=sha256:fd2953072aa81348ee20a41413a44845fe2434fbaffbfbaebe98fa8dec170927

Observation eb610b38-d85c-4f9d-b2ec-fe3f114e1d93 · inbound

Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective cites this paper.

Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-09T01:54:34.022760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-07T16:41:23.234607Z digest=sha256:d4819c658f0459204ccdd8f25beb7e34ce95f735580d4a342e4970b6cbf56eb1

Observation a5420cf9-fe87-4123-bd0a-133290fc9353 · inbound

Tangram: Unlocking Non-Uniform KV Cache for Efficient Multi-turn LLM Serving cites this paper.

Tangram: Unlocking Non-Uniform KV Cache for Efficient Multi-turn LLM Serving SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.130137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T02:25:46.234576Z digest=sha256:d7e148fe0ec3a423d9bb8a9a5e9da2a2c04ca79d52165b9625d5d857d72e905d

Observation 5f1e4339-8f7e-41f6-b85b-7663bea7bd5b · inbound

End-to-End Context Compression at Scale cites this paper.

End-to-End Context Compression at Scale SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:17:31.591780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T16:36:54.699174Z digest=sha256:313b005192c671f273c28d77af0fefa4d6d08ac9473af4f4550a5a364d2ce1a6

Observation 200ced63-88ec-4fe5-803b-61f3175d24dd · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.018523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:3b89b4e71bcaf2f1049647854d1041d3bac3897427c65ccdf72d1377d27a45be

Observation e4281599-93db-4aae-92a1-5122263b59ed · inbound

Understanding Sparse Attention Selectivity in Long-Context Foundation Models via Counterfactual Evaluation cites this paper.

Understanding Sparse Attention Selectivity in Long-Context Foundation Models via Counterfactual Evaluation SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T23:14:59.988064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:14:59.988064Z digest=sha256:8591a4a20b7e1024be3f4bba94e12f4df455da2896b58311a8e7e5146dd04442

Observation e9ba8123-f1c0-410d-9398-a20799742211 · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:50.964348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:43:50.964348Z digest=sha256:6d8fcd033b764b732fb92133e2c8ffa8af3ca3eeb48c80d4ab16cba379412c33

Observation f0409599-70a5-44b9-9c7f-193024812f5b · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:11:48.319572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:11:48.319572Z digest=sha256:7becfdfaeacdbe86f731cbcc13c29ec72e066049bf2938758ad2dc4a01f1ea2f