Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:34:37.733484Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2412.05693.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:34:37.733484Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
15 of 15 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 963a373e-16be-4569-9080-a6dc6da82bf7 · outbound
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 508a6444-f331-4883-9d78-7eb294fba689 · outbound
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression Training Verifiers to Solve Math Word Problems
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19a13875-65e4-463e-9e1f-3287b54e49e4 · outbound
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d4fbcd0c-bc39-4110-9c63-5c11a70464d3 · outbound
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression M., Melis, G., and Grefenstette, E
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8f6f03a4-cec3-4e8c-867e-d4b889a9cc9e · outbound
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression SnapKV: LLM Knows What You are Looking for Before Generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2a8f24f-d5e4-433e-9aba-4eac34e30592 · outbound
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e00cfda5-5e02-4673-87df-6db009a5884d · outbound
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression N., Çaglar G \"u lçehre, and Xiang, B
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3f40174b-6fd9-4fbc-82dc-47b7f45733c4 · outbound
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression Transformers are Multi-State RNNs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3a5e6ca-a5ee-4335-86df-af172c5f7409 · outbound
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 003183d3-cc6b-4da6-9b94-767030bda512 · outbound
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression Fast Transformer Decoding: One Write-Head is All You Need
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27f0f9fb-7aae-4a7d-844d-c4cd1e238341 · outbound
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a09cf4c7-5f1d-4d66-87a5-74f0953b076a · outbound
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f3ad42-38c2-4579-8dba-43b31ac61714 · outbound
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression Efficient Streaming Language Models with Attention Sinks
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 18a88eed-9915-436d-a280-34256d028a1d · outbound
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9003b91d-f4b4-4eb8-a36f-95c003c52961 · outbound
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
No inbound Pith citation observations are available.