Pith. sign in

Paper Citation Record · LEDGER

Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression

As of 18 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2412.05693.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05693 v3

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:34:37.733484Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 963a373e-16be-4569-9080-a6dc6da82bf7 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:34:37.662105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:34:37.662105Z digest=sha256:d41901cd0056c82cfa72d7c3502c8c7c161bd2d003d889da1d4d6b32ffb88e59

Observation 508a6444-f331-4883-9d78-7eb294fba689 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:34:37.667150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:34:37.667150Z digest=sha256:b8e42821f8ea7203477dc5894366647828c09c7d4fd5eb073ee66797cd4ed3a3

Observation 19a13875-65e4-463e-9e1f-3287b54e49e4 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:34:37.975935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T20:34:37.671700Z digest=sha256:84451bbc1f90f0b44eef8583820494934b9b6cde4a8c990c94fbc839b4e9c102

Observation d4fbcd0c-bc39-4110-9c63-5c11a70464d3 · outbound

This paper cites M., Melis, G., and Grefenstette, E.

Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression M., Melis, G., and Grefenstette, E

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:34:37.962138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T20:34:37.675836Z digest=sha256:0b01d656c33ee5c4064790294041957b7ad93462727089eff48737b055fd0b4e

Observation 8f6f03a4-cec3-4e8c-867e-d4b889a9cc9e · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression SnapKV: LLM Knows What You are Looking for Before Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:34:37.680973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:34:37.680973Z digest=sha256:eb8f0fe9a7332f6002031d9c3ed8a4c27b8f6e67b0435d1258cd8907a3dcdae7

Observation e2a8f24f-d5e4-433e-9aba-4eac34e30592 · outbound

This paper cites Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time.

Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:34:37.947527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T20:34:37.685647Z digest=sha256:a4dbec06c0df8305c1bae3b199b788aa48e6a9b8b4d9a1a8c64884bf4adec874

Observation e00cfda5-5e02-4673-87df-6db009a5884d · outbound

This paper cites N., Çaglar G \"u lçehre, and Xiang, B.

Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression N., Çaglar G \"u lçehre, and Xiang, B

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:34:37.933923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T20:34:37.689843Z digest=sha256:8f824f04428408375846b525f53a36e2db2fe7da9480368201c0154c1951d6af

Observation 3f40174b-6fd9-4fbc-82dc-47b7f45733c4 · outbound

This paper cites Transformers are Multi-State RNNs.

Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression Transformers are Multi-State RNNs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:34:37.694365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:34:37.694365Z digest=sha256:c79d1bdfb7c4875b79f06fe43158d6e0ef51abaf6fdb0ac2d9b453ac9ded8a82

Observation f3a5e6ca-a5ee-4335-86df-af172c5f7409 · outbound

This paper cites On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference.

Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T20:34:37.699018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:34:37.699018Z digest=sha256:bd5fa4afeffb1bf6f0ed1f9b607989a2b98ddc10670d75d98b208b733684adfe

Observation 003183d3-cc6b-4da6-9b94-767030bda512 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression Fast Transformer Decoding: One Write-Head is All You Need

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:34:37.704378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:34:37.704378Z digest=sha256:2c996578f8114001e1b91969dfccd45e05a5f13566fd27ddc48554ac26c27c9d

Observation 27f0f9fb-7aae-4a7d-844d-c4cd1e238341 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:34:37.709785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:34:37.709785Z digest=sha256:112f7618200c9f092160b13766b668c06f4da87b4a932a7f93f3676263fa1404

Observation a09cf4c7-5f1d-4d66-87a5-74f0953b076a · outbound

This paper cites SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget.

Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T20:34:37.718187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:34:37.718187Z digest=sha256:0ef80d6bd132bca041001569e342f253440f0d31ccf1779f36229eb13d7a48ae

Observation 77f3ad42-38c2-4579-8dba-43b31ac61714 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression Efficient Streaming Language Models with Attention Sinks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:34:37.918092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T20:34:37.723302Z digest=sha256:607c9194adea07ea158d42d80f1579702c63b5bbe07a1500c2b2a051592294ed

Observation 18a88eed-9915-436d-a280-34256d028a1d · outbound

This paper cites LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation.

Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:34:37.728742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:34:37.728742Z digest=sha256:3904243c794d2df977369894fa1f66d8f01ca06eb7f0a765215e36226b6c4e84

Observation 9003b91d-f4b4-4eb8-a36f-95c003c52961 · outbound

This paper cites H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.

Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:34:37.903408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T20:34:37.733484Z digest=sha256:35f0ab316318aeba61e4ad26e9acaf908625478524b720c26e7bcda04e22b74e

Pith citing papers

No inbound Pith citation observations are available.