Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T10:57:44.554457Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 3 inbound Pith citation observations for arXiv:2510.07651.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T10:57:44.554457Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T05:58:45.892043Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T08:06:33.691924Z
15 of 15 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cead4cdf-912d-450a-ad98-cab3b41a2ead · outbound
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Qwen Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b330c0fc-b402-4df2-bb98-85756f245a3d · outbound
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference The Llama 3 Herd of Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fccc0d0-f496-408b-aa8e-b062405847ae · outbound
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference GPT-4 Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0da759d6-e3a3-47f4-8bed-dee80839d12a · outbound
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Transformers are Multi-State RNNs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9beb50c-1ef1-486c-aab2-5730030b5a4b · outbound
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Code Llama: Open Foundation Models for Code
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aa486f3-a59e-4e2b-9ed7-5fd3f474efd7 · outbound
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference LLaMA: Open and Efficient Foundation Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f6ffa20-de3a-4f7c-b2aa-542b49f58d15 · outbound
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb5cb6f7-5a78-4f87-973b-cf7cacc05b40 · outbound
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Task-specific evaluation metrics (e.g., Exact Match/F1 for QA tasks, ROUGE for summarization tasks) are re- ported using LongBench’s official evaluation script3
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15003227-2cf3-436d-a8e9-81b665cdc891 · outbound
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference All eviction methods are evaluated with a fixed 1024-token cache budget
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50848417-4499-4d46-be35-52a4545ff997 · outbound
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e24562f-9ff6-4b16-9c8f-fc0e680c0989 · outbound
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 1992
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fafd378-c5e3-4f4c-8aea-981b1e5929ff · outbound
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Compressive Transformers for Long-Range Sequence Modelling
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6381fab-c13e-4ae1-afc0-450ed921c73e · outbound
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e88f5d3-3700-40e5-80ad-3f88f9bff7bf · outbound
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Evaluating Open-Domain Question Answering in the Era of Large Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b280c9da-d9b7-4e5e-8bf7-e2c4f0a1cbf2 · outbound
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Transformers: State-of-the-art natural language processing
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e446f852-d717-4982-bde7-a17c29e5d233 · inbound
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c67c7406-a3a2-41df-9a04-98c6fd5e1b13 · inbound
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 362591c2-7ed1-4e33-9052-e37188acaa17 · inbound
Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.