Pith. sign in

Paper Citation Record · LEDGER

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache

As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 3 inbound Pith citation observations for arXiv:2411.18077.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18077 v3

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:37:07.697300Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:38:47.483457Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T20:59:01.556666Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e7ad4eaf-d1d5-4d47-a479-334cd7017ad5 · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.114579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.544023Z digest=sha256:8a5d4c842744b7d4daa0fa10a5f6434d520de2895b36d34755be1b83c643b08f

Observation 5af67f7f-5cea-4d68-bef3-4751fcbf1add · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.548481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.548481Z digest=sha256:20fa35acab0c90baed9180625e6cb2ab3c1c9f7d695435957467092275872846

Observation 592c3a29-2acc-4656-848d-50800257f37c · outbound

This paper cites Reducing Transformer Key-Value Cache Size with Cross-Layer Attention.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.553067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.553067Z digest=sha256:9d9c7f4030cd8173ae0555abfd2f3188328163dba907b83096c0ee4456b207d0

Observation 7d456d26-f44c-45d2-870e-b89f54a6f0c5 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.557293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.557293Z digest=sha256:c9b0488d2ade561eed44b07561cf7a1b24fd662659191dbe02a58092c8f82eb1

Observation 03199717-1dcd-425a-9376-503a74745195 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R \' e.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Fu, Stefano Ermon, Atri Rudra, and Christopher R \' e

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.561541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.561541Z digest=sha256:e21a391b73db990006c6aa16da562039e2f8c32eb21f7594edbfb358bd3904f1

Observation 199d81b7-0328-4935-ada3-998fede18f89 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.565472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.565472Z digest=sha256:677233fedd6351f5637781499ad6e16c14c9634538b29c3c99c0d0c4898457c1

Observation e8433fc1-ed3c-4e87-9e95-4dd9418229ce · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.570007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.570007Z digest=sha256:e6bb35dc8735f6feb3b8cf284be37907efe49e5e129b6ba777c4be15d43207e1

Observation 32394cf1-ad8c-4554-8e12-3cdc54a95116 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.573793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.573793Z digest=sha256:acaa1379b822701d743c42c470453df56fd53ba43206af25c798512c0bed7663

Observation 4cf2ff58-02f5-4c4c-91b4-2c10cc73130d · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.577744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.577744Z digest=sha256:8a4a93b6c9978daa6f9d6d172a1b25239c791dcffdfbefa190bb14765460f4ce

Observation c6ce3c48-7d7e-4ca0-840a-3a7eb8132d43 · outbound

This paper cites Mistral 7B.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Mistral 7B

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.581390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.581390Z digest=sha256:06e72d576b29ab220f264cff6dfaf0579206d5af76ca1c7f8d4d1cec0a9c82f7

Observation 95b0495f-cc05-4fd9-b454-a6d9643c4add · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.096814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.585359Z digest=sha256:7f24cddc8b8667ea65f2d972830d73a6da35ba42746c020bcbafffcb750e93e5

Observation 904d9b76-48d9-40e5-8e95-fd183426b773 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache SnapKV: LLM Knows What You are Looking for Before Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.588847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.588847Z digest=sha256:dabdb3e6fd4e77cbe558f6a0443d971f05522d76219d2ecb3365439479eadfe0

Observation adbad6c2-b9d1-468a-b410-838668088066 · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.592449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.592449Z digest=sha256:bd08f1483858672de40575b0b25e7e348b7e9d94b3b1df131d0190c4305ab282

Observation 39ffc95a-dca4-42ef-829d-d6bc784359bd · outbound

This paper cites MiniCache: KV Cache Compression in Depth Dimension for Large Language Models.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.595756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.595756Z digest=sha256:bb60fb0be0ff058b23f644b53f6e15c581e8d13214df0ed668a9936a43a136e7

Observation 4f7676d9-5061-4d50-a2a4-ea539a3a263e · outbound

This paper cites Multi-head or Single-head? An Empirical Comparison for Transformer Training.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Multi-head or Single-head? An Empirical Comparison for Transformer Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.599724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.599724Z digest=sha256:b405b561d666e72de20483b232d50235666c8987a88d70f861deaf7652da6cae

Observation e1734be7-5d10-467c-9a4d-755149e86b5a · outbound

This paper cites LLM-QAT: Data-Free Quantization Aware Training for Large Language Models.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache LLM-QAT: Data-Free Quantization Aware Training for Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.603254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.603254Z digest=sha256:526d17d7ab58ee40fa83f90d8ba5ee1097cb73118fa2bfb18beb348430b3b9ea

Observation 7ae1c708-e961-4a3a-b887-930d580ce0d4 · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.078454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.607108Z digest=sha256:dde46e51972aa66ad0b8fa937a3c9ee36a42d5ba5e8f7ad88c07f5e0d405701a

Observation 1f830ab9-1b00-4ec0-95da-5b50c75c7476 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.610637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.610637Z digest=sha256:a38d1e5c55ca916fe3a7cee8c058264d1fdd48daf03251d45c09e52ee57acc15

Observation e45af902-9bf3-44bc-bf9b-cd7f903b4e55 · outbound

This paper cites Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.614689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.614689Z digest=sha256:e2d392e401d555312bace0c1d5e986dfb8c2b01d1fc28caf2d79a03265ddda27

Observation a8b8127c-4618-4611-bc2b-9dbea5525753 · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.067490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.618605Z digest=sha256:3f8d8aec6e87ac07695eebb81d22ad603a66b5f823961a54fff63a56ef50dce7

Observation 7057606f-fac1-4ad3-aeb0-15f972b7e1ae · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.622246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.622246Z digest=sha256:bf24a576a316123f7046e477001e43e0d53777a103460b3110dc4329ad644712

Observation 57c805ba-a686-4df7-9460-27b12a7190c3 · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.049427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.625766Z digest=sha256:3e60aedd41d53e8d33aac6e3972d78fc06dad007fe79020e20efeadb767c2925

Observation 3fa7a7c7-6b28-4d72-bb1f-86ea6ca9b1c5 · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.629296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.629296Z digest=sha256:b297e912bbaeb89ec48f4282170745cbbc76645f20b4e51e5a19bda062a175fc

Observation bd394d09-30f7-4f6c-bfe9-99f30fa42aff · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.037753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.633190Z digest=sha256:a3809b7db964ba437c27d307d898045179eba568321378be0afc24e6d66bdbe7

Observation 80ae3419-7cf7-48bd-860e-22fbb5e45893 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.636669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.636669Z digest=sha256:87ca709eb544c2eaf760055667dceab420a68b88310bdcf637c61a9be3e6414e

Observation 3788dc32-33fe-4100-9964-5c88c6e2a9c1 · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.026080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.640473Z digest=sha256:82af361f87f2cc6f719bd8d7f16f54b6fb03d6807437b9e017ed325812e0735c

Observation 97a6cc9b-743d-4e2d-8581-1a9877b43c1f · outbound

This paper cites Do Large Language Model Benchmarks Test Reliability?.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Do Large Language Model Benchmarks Test Reliability?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.644899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.644899Z digest=sha256:d3f5bb311921471b205a3f1f3dcea2e587782b19a2e7cfa280d0074c2a2ebc8e

Observation 7e18dec0-c176-4357-b125-3a597f77447c · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.014635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.648882Z digest=sha256:ef364a16e87c0b451a5e919a8c5b6a15228ca70c6129167dcf6a46fcb9021a59

Observation 04249263-2fab-41bd-a950-dabe643a7754 · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:08.003790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.652446Z digest=sha256:55bd68977768d541487cb191cf64d046cd91a5cd9e98ef35a3323e4ac1ebf4d9

Observation c5ca8fff-152a-451d-aea9-2bebe4f1518c · outbound

This paper cites D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.655879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.655879Z digest=sha256:2459e0d245584a2b87b486e4a662bd739227e9b0fb167024004342a9cfb51208

Observation e6564002-5de1-46ec-94b7-1698ddbae20a · outbound

This paper cites Retrieval Head Mechanistically Explains Long-Context Factuality.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Retrieval Head Mechanistically Explains Long-Context Factuality

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.659393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.659393Z digest=sha256:96c142b90a08ef276fe3b3becad195e54854e2aae4c32817de9bbc5b87b14382

Observation cd0916af-31f0-4431-ac51-2fc67cc52b04 · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:37:07.992453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T11:37:07.663004Z digest=sha256:1d1525f823783ef99e4cd23e9daf1e81d444f058022519796a5bf582c17316e1

Observation 20e63b59-4d5b-4e35-bafc-1cbcebf0873c · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.666624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.666624Z digest=sha256:c28589fdc451077f18aeb109b363a5b0a7fbb70b9fd3bcd1da4904787ca0edc5

Observation 6c22d63e-a693-4633-bfd8-eb0c8e03e13b · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Efficient Streaming Language Models with Attention Sinks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.670319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.670319Z digest=sha256:be12b44ca1a5ce575b6a6c7a4894a112dece1bce797d06bab339f471b21f5c29

Observation 7ced9945-c15e-4e04-939c-711b8dbb8902 · outbound

This paper cites PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.674267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.674267Z digest=sha256:749cad9175add7db3a1d9db44e114ed7b6ef8e7746930983fd54fffe6fd3ea08

Observation c9d88fdc-0c2f-4b57-87f3-7a4dfaadb303 · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.678061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.678061Z digest=sha256:4cbe9b10a76f870d8a6cf84853dde9424a970336db0cdaab1ade7d8f02b70c73

Observation 7267b29a-8ed7-424e-b943-25a38825d5d3 · outbound

This paper cites $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.681788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.681788Z digest=sha256:d602c954ab1f413057e655110f7c2d4222b23789d3657ef173e53712e5ff596c

Observation 9340b39b-65a6-4abf-adff-3f0066468f2c · outbound

This paper cites an unresolved cited work.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.685940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.685940Z digest=sha256:7f2e9a57a689fecec9374cc450e4781138f4bd09cb1c716dc08b55a67f9d3745

Observation 248e51a3-32a2-4a0c-b655-ab701a21bfa8 · outbound

This paper cites Barrett, Zhangyang Wang, and Beidi Chen.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache Barrett, Zhangyang Wang, and Beidi Chen

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.689428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.689428Z digest=sha256:f4e71010317300fa163ff16b6ce8f1085107e91a3f51f720eb9b077726808765

Observation 93f5e1f6-0042-405c-9339-19fa85f70aa3 · outbound

This paper cites online" 'onlinestring :=.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache online" 'onlinestring :=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.693059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.693059Z digest=sha256:4cd119e90dc43f4126c0ae161ad0e7a6057397101f144d7186b04a27b33e72f7

Observation b66d035f-d673-4e4c-92c0-6329e84a19c7 · outbound

This paper cites write newline.

MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache write newline

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:37:07.697300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:37:07.697300Z digest=sha256:8651eff549e9de56bd94ffbf58e00c7830e8eaf47ac574c02ab8a5efa56a489f

Pith citing papers

Observation b1956160-20f3-4901-b2c6-5b51d48b213a · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:47.483457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:47.483457Z digest=sha256:c13be5166b3abebe44552c39dc80f44f41b1cad96e4a5d04047668df0c3e1a61

Observation c752e517-becb-4267-8201-5ca35c10a045 · inbound

Minimal-Intervention KV Retention via Set-Conditioned Diversity cites this paper.

Minimal-Intervention KV Retention via Set-Conditioned Diversity MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:59:01.558729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-20T20:58:07.403912Z digest=sha256:72b8e82936109407852b986ed5f574b4e3edf7c68e394336f4766e666388089e

Observation bbcf9321-28d6-400b-946f-c5dddb42a72e · inbound

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization cites this paper.

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:18:16.562435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T12:16:07.797702Z digest=sha256:bcde986e7bb3970dba00c355207af8b220fe1626806e9b91f85999e3df7f7d15