Pith. sign in

Paper Citation Record · LEDGER

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference

As of 8 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 3 inbound Pith citation observations for arXiv:2510.07651.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.07651 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:57:44.554457Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:58:45.892043Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:06:33.691924Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cead4cdf-912d-450a-ad98-cab3b41a2ead · outbound

This paper cites Qwen Technical Report.

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T10:57:43.002729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:57:43.002729Z digest=sha256:d8d7171f413dcc7dc4c9501b295d66c39f5e52b30e53ee893bd7c9b5a660976d

Observation b330c0fc-b402-4df2-bb98-85756f245a3d · outbound

This paper cites The Llama 3 Herd of Models.

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T10:57:43.283901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:57:43.283901Z digest=sha256:c99cf292953f1a0bcba40ce53b14ac2ec60c28d93567f5cdcc2b2a619de08885

Observation 6fccc0d0-f496-408b-aa8e-b062405847ae · outbound

This paper cites GPT-4 Technical Report.

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference GPT-4 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T10:57:43.626012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:57:43.626012Z digest=sha256:ca8fe92abb989492cc8257c47994a6aac4f10fae5d0a2d8e84b188b8e09b04a0

Observation 0da759d6-e3a3-47f4-8bed-dee80839d12a · outbound

This paper cites Transformers are Multi-State RNNs.

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Transformers are Multi-State RNNs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T10:57:43.739935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:57:43.739935Z digest=sha256:fcc51c5c62aa59ec961cc0bd713c4b94b33c49e0a6cc75ad2709e08079d6da6c

Observation c9beb50c-1ef1-486c-aab2-5730030b5a4b · outbound

This paper cites Code Llama: Open Foundation Models for Code.

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Code Llama: Open Foundation Models for Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:57:43.957435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:57:43.957435Z digest=sha256:d03506a2d30f49e38d041a59ceea506106a7a883b47bf5665b8df6c581499381

Observation 4aa486f3-a59e-4e2b-9ed7-5fd3f474efd7 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference LLaMA: Open and Efficient Foundation Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T10:57:44.070569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:57:44.070569Z digest=sha256:805aff7638ed03bee4bc7caa3fa6c305ba696d0553b154c0ceaf78e9f9427427

Observation 5f6ffa20-de3a-4f7c-b2aa-542b49f58d15 · outbound

This paper cites an unresolved cited work.

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T10:57:44.336299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:57:44.336299Z digest=sha256:46ef7dc3cd929437d6c81a7ce22713e1349eaca17c35253e20a276c09b515eed

Observation bb5cb6f7-5a78-4f87-973b-cf7cacc05b40 · outbound

This paper cites Task-specific evaluation metrics (e.g., Exact Match/F1 for QA tasks, ROUGE for summarization tasks) are re- ported using LongBench’s official evaluation script3.

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Task-specific evaluation metrics (e.g., Exact Match/F1 for QA tasks, ROUGE for summarization tasks) are re- ported using LongBench’s official evaluation script3

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:57:44.427450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:57:44.427450Z digest=sha256:ea6898ad710536ed97fba676b9592417921ddf342831f6c723bf14bfc7318c72

Observation 15003227-2cf3-436d-a8e9-81b665cdc891 · outbound

This paper cites All eviction methods are evaluated with a fixed 1024-token cache budget.

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference All eviction methods are evaluated with a fixed 1024-token cache budget

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:57:44.476767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:57:44.476767Z digest=sha256:076d0e684f7fff71552b3e5b03ea20f550463931668557b606c2be4590cc77e8

Observation 50848417-4499-4d46-be35-52a4545ff997 · outbound

This paper cites an unresolved cited work.

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T10:57:44.554457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:57:44.554457Z digest=sha256:8f07e70fc88cf965702d0b64c9efe76a23f89347be0d57691bb377bbbad2305b

Observation 5e24562f-9ff6-4b16-9c8f-fc0e680c0989 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-04T10:57:43.388072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:57:43.388072Z digest=sha256:17ae0beec77fb4d40fe8f449a2f398857cfc76a57c9eb0362eb4cd88cd663a75

Observation 3fafd378-c5e3-4f4c-8aea-981b1e5929ff · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Compressive Transformers for Long-Range Sequence Modelling

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T10:57:43.845157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:57:43.845157Z digest=sha256:fca7aaccbe9ae7bd8277dde961ef18615779764cf638b3afe972b0771037c886

Observation e6381fab-c13e-4ae1-afc0-450ed921c73e · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T10:57:43.165745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:57:43.165745Z digest=sha256:cc006ce56ce05766442065b113306aa5cb1b16b0d1319d4be0d53c58759f6c28

Observation 0e88f5d3-3700-40e5-80ad-3f88f9bff7bf · outbound

This paper cites Evaluating Open-Domain Question Answering in the Era of Large Language Models.

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Evaluating Open-Domain Question Answering in the Era of Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T10:57:43.507915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:57:43.507915Z digest=sha256:4cabd5b16a125a4d5edae4aedad907df62a6ba8075d2e66e06f988a8c41e2e84

Observation b280c9da-d9b7-4e5e-8bf7-e2c4f0a1cbf2 · outbound

This paper cites Transformers: State-of-the-art natural language processing.

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Transformers: State-of-the-art natural language processing

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T10:57:44.194526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:57:44.194526Z digest=sha256:d8f907de30b26eccbca86f7cf35f1dce56e35b2923495752e08566b371b05a85

Pith citing papers

Observation e446f852-d717-4982-bde7-a17c29e5d233 · inbound

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation cites this paper.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-01T02:02:20.220338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:736f781220b06d75d1d91e42e1f5938bed6c25f254d31c3695e7a33e76f97bc5

Observation c67c7406-a3a2-41df-9a04-98c6fd5e1b13 · inbound

RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache cites this paper.

RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-01T02:02:20.220338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:16:15.869198Z digest=sha256:f647ccb1926ae740079d49fa6d4942047bdabf0d9a3ceec03cf5a0e08ed0f567

Observation 362591c2-7ed1-4e33-9052-e37188acaa17 · inbound

Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference cites this paper.

Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T05:58:45.892043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:58:45.892043Z digest=sha256:b1fa294d4d1d460e1fcfefb9766be0cf551bc4ef1cc984eb2a792d7ee62a84df