Pith. sign in

Paper Citation Record · LEDGER

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference

As of 14 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2412.05896.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05896 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:18:18.542070Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c055c66-937b-4714-b416-2a93fc7c64c3 · outbound

This paper cites GPT-4 Technical Report.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:17.512660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:17.512660Z digest=sha256:cede7586023dd10c43b801f367d35b074d28a80e82dab250a6fcb70082ee6a50

Observation 07b9240e-b607-4c10-96c7-951783dffdde · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:17.574028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:17.574028Z digest=sha256:3cfaae1e80ea0231fec88433e4995d375ba4e95f7c0f2d3f32b537aefd3c9f91

Observation 9cb2e7ac-7654-44df-bd4e-50c5a5bd1c4f · outbound

This paper cites Code Llama: Open Foundation Models for Code.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Code Llama: Open Foundation Models for Code

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:17.628259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:17.628259Z digest=sha256:14d87b46f456e925252b4d2969988666c2636e3d80902ee404ac8c16bbe371cc

Observation 09a7a6b8-2d7e-4dd0-85b2-208b50436ba1 · outbound

This paper cites Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:17.661932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:17.661932Z digest=sha256:146473ec2643ca057470deb6df6cb84409f5ed68149f8402020d33af986e7cfa

Observation 5e0f265e-8a7e-4b8f-9fc3-b56ab7a8531e · outbound

This paper cites BigTranslate: Augmenting Large Language Models with Multilingual Translation Capability over 100 Languages.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference BigTranslate: Augmenting Large Language Models with Multilingual Translation Capability over 100 Languages

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:17.694269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:17.694269Z digest=sha256:710947e837929a5bba8d0dabe308b0585a4d7f506bc6223c0c3a24116959cfcc

Observation 6a4c3da6-16b0-4a75-8e3c-aefbc8fcd193 · outbound

This paper cites Attention is all you need,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Attention is all you need,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:17.726994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:17.726994Z digest=sha256:931d5d01bec3ed1b914f934fb8edd5357fa1335dea66ba30c6fbd8bbab695fcd

Observation 09f9e62c-f1e5-42cd-a751-790c9f667659 · outbound

This paper cites Language models are unsupervised multitask learners,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Language models are unsupervised multitask learners,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:17.745916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:17.745916Z digest=sha256:77ceba1a43dfd69ef02a763042470ef87fb095ea88c27762dbc8c835a524a007

Observation bd8d2ff1-9305-44d5-8d28-3e113cc26bf4 · outbound

This paper cites The Llama 3 Herd of Models.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:17.749958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:17.749958Z digest=sha256:d5d66870449544d9c3e92d9e9b7a432745b97ee4ed6304489db2ad36ce198680

Observation e865ae17-693f-497c-bc1c-2f350d3c17b2 · outbound

This paper cites DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:18:19.663061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:18:17.754162Z digest=sha256:0c80bb63e4fcec95f9191e72681799ac9bf881260860c38d1fdec93cbc5421f2

Observation 635d9776-ddcd-4f9a-abd7-b52a8ce58373 · outbound

This paper cites Taming {Throughput-Latency} tradeoff in {LLM} inference with {Sarathi-Serve},.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Taming {Throughput-Latency} tradeoff in {LLM} inference with {Sarathi-Serve},

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:17.758144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:17.758144Z digest=sha256:12e6a99386940c451b8fd4a8e0b0a0f69e8b924a8cade5ffc4231d1da5284654

Observation 4bc970ca-837e-4ab2-907a-dc13a6579332 · outbound

This paper cites Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:17.762149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:17.762149Z digest=sha256:cb2c4315f29f35151374f64ac22f49ad44ce7ca793b056f8b26e01ff608aeb74

Observation f949eb0b-9063-46f8-8d69-5c8ca568bc7b · outbound

This paper cites Cam: Cache merging for memory-efficient LLMs inference,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Cam: Cache merging for memory-efficient LLMs inference,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:18:19.617635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:18:17.766660Z digest=sha256:d9e4179efadda94f49a9de937bbd385238326f55ee6908563556f9491abd2207

Observation 4403e995-db45-424f-b38b-3cc8f5c321be · outbound

This paper cites Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:17.793321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:17.793321Z digest=sha256:3361fb4bda376603399fa6012b2d93c0614aa0bf4d2592fbbe674f556ee0f516

Observation 01665fc2-6ef2-4415-bf6a-6d14cac1887e · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:17.868031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:17.868031Z digest=sha256:e87322e0fb7dfb5411f759f5e4ee7a729c4b23c5c5ad5e673dfe0c9282ea0f25

Observation 78aba339-3138-4ca1-84c4-ff1e468612f8 · outbound

This paper cites PQCache: Product Quantization-based KVCache for Long Context LLM Inference.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:17.944588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:17.944588Z digest=sha256:ac7b54ce4f3a5cd778fdbc5be262885f99121072aded97dd61c5b318cc184279

Observation 9ad1737c-78b2-4c46-a9ee-7431ac55c420 · outbound

This paper cites Efficient streaming language models with attention sinks,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Efficient streaming language models with attention sinks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:18:19.568865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:18:18.020570Z digest=sha256:fe47259b51e4349123afe1d85b5d4b8c5b6a25340dc34f4f1ebfe372c0e035d1

Observation 42f33cf2-bfd5-4168-9fb0-4c80c50091a4 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference H2o: Heavy-hitter oracle for efficient generative inference of large language models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:18.062326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:18.062326Z digest=sha256:03a708c34fddf738a51ec9dcecf4feb80d2cb43a3ad84595a8a45c75a40c1ec9

Observation 5da5dd4a-5a53-4efc-b9ac-9d29aff8cf88 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference SnapKV: LLM Knows What You are Looking for Before Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:18.109882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:18.109882Z digest=sha256:10aa33d8b9c4c87779a8a47eb36ff2c94ac80af4ce9c0ed59b1444f16e5d6e21

Observation 33ccda45-dabe-436d-ace9-c95026b75da0 · outbound

This paper cites PyramidInfer: Pyramid KV cache compression for high-throughput LLM inference,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference PyramidInfer: Pyramid KV cache compression for high-throughput LLM inference,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:18:19.490479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:18:18.160144Z digest=sha256:5dd70c05e154dc23d02be8136d973ec0d6aaf2bce59f1e1bca169d9ab6410a23

Observation fe9687a2-803f-4f7a-96c7-78bec9e8fd1d · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:18:19.375500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:18:18.177088Z digest=sha256:d18ef948c748292e723dac18e038889eb00b1d79985123bd882afc9512b589fb

Observation 63227244-9140-4c27-846e-1fd5a6032cb9 · outbound

This paper cites Model tells you what to discard: Adaptive KV cache compression for LLMs,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Model tells you what to discard: Adaptive KV cache compression for LLMs,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:18:19.344417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:18:18.186155Z digest=sha256:78a91eee4f027dbd99fcfd6bb69115e3a81ecd190c8950510145a344f31d75ce

Observation 9e3e2cf7-8c4b-4ae0-95cb-10a70f904c40 · outbound

This paper cites LooGLE: Can Long-Context Language Models Understand Long Contexts?.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference LooGLE: Can Long-Context Language Models Understand Long Contexts?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:18.203912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:18.203912Z digest=sha256:4354c434fcf95936f2d1274d964c940eff430bc9ca42e2cbb00b82b25b58868e

Observation e5a6bf00-40b7-4191-9f24-2baea217dc2f · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:18.218270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:18.218270Z digest=sha256:a034de78f161572acbf00cb378605fcfea5619efc646f22e68f8c14f7a8192d9

Observation 7d453edd-8979-4d3e-82b0-9f9cf16794cf · outbound

This paper cites Tsplit: Fine-grained gpu memory management for efficient dnn training via tensor splitting,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Tsplit: Fine-grained gpu memory management for efficient dnn training via tensor splitting,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:18:19.270618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:18:18.240316Z digest=sha256:f891da3d007965e7b39996e2b3cc8c430ca176089d53a127e9fd3e9724d81bdc

Observation c231ff2d-aca3-4261-9acb-1ec24f10313d · outbound

This paper cites Het-gmp: A graph-based system approach to scaling large embedding model training,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Het-gmp: A graph-based system approach to scaling large embedding model training,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:18:19.176927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:18:18.250080Z digest=sha256:09adf97180f8f5144b524336c790d47f587f692066d875895f6d8af55f908f38

Observation d6c90f88-3dd7-4017-902f-92f4384eae7b · outbound

This paper cites Platod2gl: An efficient dynamic deep graph learning system for graph neural net- work training on billion-scale graphs,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Platod2gl: An efficient dynamic deep graph learning system for graph neural net- work training on billion-scale graphs,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:18:19.130730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:18:18.254703Z digest=sha256:9f42a4d40fc8f7bd6202f30996747babedeec2207f00cc37cb316218735d066a

Observation 726d85d1-ff7f-45fb-b65e-30e9558d788f · outbound

This paper cites Optimizing tensor programs on flexible storage,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Optimizing tensor programs on flexible storage,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:18:19.050351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:18:18.259394Z digest=sha256:487ef47ae63579bb3ec050d8bcdc5ed2f51174de17da304a24911f8bbbcdb5b9

Observation 46a7c86d-7356-41e5-a6ff-ba62edc91568 · outbound

This paper cites You Only Cache Once: Decoder-Decoder Architectures for Language Models.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:18.263580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:18.263580Z digest=sha256:697b86a1b8cb26a0c7a30fb3b6b8fce569b76d4edb9cf76e74bf0dd2146023a4

Observation f1f659a8-005c-48d9-b736-fb859617fdb4 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:18.267468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:18.267468Z digest=sha256:da81f7c0321e5def6817300b150f2c922790ccfa7f25bf6add26e531bf373ccb

Observation 28cfb6ec-bace-40b4-bb9a-039f587b4091 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:18.271340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:18.271340Z digest=sha256:2486255c6f65341fd1ba02f5dc8e3f243d322a56187ccf182cc79bf2db3e6eed

Observation 8f9eca48-fff4-4f70-95ef-c7730d1660cf · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Flexgen: High-throughput generative inference of large language models with a single gpu,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:18.292763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:18.292763Z digest=sha256:b4a92c788b1b2739364397914c1f736ec07f460d816d40d9301bada950fbdad8

Observation ffdf17bf-0f5e-4a83-8e81-3bfc63d2eb51 · outbound

This paper cites Efficient sparse attention needs adaptive token release,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Efficient sparse attention needs adaptive token release,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:18:18.987661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:18:18.349173Z digest=sha256:c0826a6907bba8b7d6b802ff813ddd0303fa73e581466bd1305927ebc990aeaa

Observation 24931133-797e-4db4-92b1-5103e501c2a3 · outbound

This paper cites ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:18.404941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:18.404941Z digest=sha256:2c8f5f4e902fc27d4aa85aa4fa5fe5e9d4ab63901030954ba550cc18ef78c245

Observation 88d4850b-4452-45df-91fa-515eae97b1f6 · outbound

This paper cites Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:18.448561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:18.448561Z digest=sha256:afc6e39efe41f8b4e8f330adc13bed9c828a1ccf890a2c5473cba22121d9dd36

Observation a767261c-f5ef-4981-b34d-80cdd15589ab · outbound

This paper cites Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:18.491299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:18.491299Z digest=sha256:b679485b0c4b9a04a653ad5492a862474887a211d133ea733cf6cd91429136be

Observation 2098b003-d575-429f-a8aa-8c95cf730341 · outbound

This paper cites Lost in the middle: How language models use long contexts,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Lost in the middle: How language models use long contexts,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T20:18:18.514907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:18:18.514907Z digest=sha256:0b8aa86a316b56e4698a09f06f3abd349286e87fb07688b8101b62fc5de841bb

Observation 3bcb1c91-78dc-451c-a2f1-327f1cdbc760 · outbound

This paper cites LongBench: A bilingual, multitask benchmark for long context understanding,.

XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference LongBench: A bilingual, multitask benchmark for long context understanding,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:18:18.894224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:18:18.542070Z digest=sha256:a726aadd4490500c11d9c203d5847d1d5b1c71b1514c73c1f47ae9e6123a9230

Pith citing papers

No inbound Pith citation observations are available.