Pith. sign in

Paper Citation Record · LEDGER

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization

As of 10 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 2 inbound Pith citation observations for arXiv:2505.19586.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19586 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:16:42.984625Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:45:12.553567Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T17:55:13.328095Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0009c9d3-4a69-441a-84cb-52e148f9e73e · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:46.286667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:16:38.963314Z digest=sha256:2522c8428a1c3a6a674db6160a8c3c40b56f486bdb1480d344decccdea161755

Observation e61ec203-ee94-4a34-acf9-b5da69754873 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:46.145966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:16:39.013330Z digest=sha256:da30b7e101e73d48f48d611e4490f292bb9dedb7271b002a1fa372a6785a24b3

Observation 56da942a-5538-4a0f-801a-95e2e1113a05 · outbound

This paper cites GPT-4 Technical Report.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:39.105730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:39.105730Z digest=sha256:50f8f844f65474d74122d4d53e3e4815d25d8f0938b2851696e335d113f370a9

Observation bff1fdef-e70e-4e89-ae0f-f8f661f4b18d · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:39.190085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:39.190085Z digest=sha256:ec0fcba2259e7bd1f608a211afee3963039b33990070031874f24e6e57be5e60

Observation 3da24f0b-5d7d-4529-b2cc-5e6e828b937e · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:39.266790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:39.266790Z digest=sha256:c3018c882393cac988a3242ff2207058ef5115a97234477702436ba2019451c8

Observation b0700a1d-b0be-4d48-a286-c8dc4035002d · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:45.978753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:16:39.358085Z digest=sha256:7f4d7264ad00637cfff1aa948835f784db43f5e03bdc3b10f962c2c7cae974fb

Observation c358e741-9b03-4389-adaa-4897f3af5a9c · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:39.438404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:39.438404Z digest=sha256:78d7c38dae09b8bbabde87fdacf0149b11600dee7e88399f20c26bbb084d858b

Observation 7917efbd-44f6-4c9b-adca-9457f47d1654 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:39.530091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:39.530091Z digest=sha256:72920ec38b295c2c66aa3fbeea33149627258ec9d6c686bafa7bf4dc1ae47b3d

Observation 050c77f7-e884-40c9-8c1c-5b1a9e8b8769 · outbound

This paper cites The Llama 3 Herd of Models.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:39.619435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:39.619435Z digest=sha256:26a7d9ffc29bd67f73de27b3b90e10c716d538e9f646db2cd3a8956855ebd10c

Observation be925673-3d0b-4f96-ba40-04959777713c · outbound

This paper cites Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:39.736272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:39.736272Z digest=sha256:6f024528e305f0ebad1a9813aed61841487ad26b7030a6dd4156f9f25023bc84

Observation e5b8f505-a5b4-486f-add0-f3ca2a95bc1c · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:45.756652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:16:39.882527Z digest=sha256:f31a9c4e0f271ab69bf4618a7030c00a8f4432ce08443df3a4265cc1ca534d8e

Observation 75d60ba6-1429-42b5-9726-e7d47c12acd8 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:39.957230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:39.957230Z digest=sha256:f49ce8458b9925f7ca48e8482915710011a3fab85465ac60e2fabde59177b796

Observation 1d546562-dc02-475e-9b47-73266fb0ab03 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.059217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.059217Z digest=sha256:dda3f41336e198edc60dddfbd14eb763f52a8622ea9348bc0ff6a7c480c33f7f

Observation 4d09a508-089b-47bb-a5f2-97700cf23728 · outbound

This paper cites Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.185668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.185668Z digest=sha256:b25445b464e18e82dc2598a77d91072e3c374fdfac860009cdd253afb2eede47

Observation 7e68ccd2-51c8-47b1-a411-fa6b04f267a7 · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.300614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.300614Z digest=sha256:38e6e6dbe6b48ddcbf355ecdb2c6cdb403d2d4c05a12312a9520a5455c29e570

Observation 83393d5b-049a-4ca5-bb72-3ce241205685 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.399838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.399838Z digest=sha256:69527d96df3bdeeb8ba8dceab8633090a935b9e00b9176e28adbc015635b4c26

Observation 0ba4e595-50ed-412f-a995-e05f8a7aeaea · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.533731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.533731Z digest=sha256:06930f654d2ac90cc07556ad7389c57a5a4db6bc4e4e9ba7f6620a52634891f2

Observation 9c533e52-2d32-4895-af8c-64dc29d1e3eb · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.602920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.602920Z digest=sha256:4d7a738c9e877c9995e7814feb5726182597911512664e2ec6b107fc4f8ed14d

Observation 2a066f05-082b-49cb-b01b-6ee0622ab53b · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.733388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.733388Z digest=sha256:8f619d755fc09dde1f05d2d57c43ae07a59fc600a86d87cf7afffef55536fc2d

Observation fd7275c9-816b-45ff-96c9-3afbc57e709d · outbound

This paper cites DeepSeek-V3 Technical Report.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization DeepSeek-V3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.840599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.840599Z digest=sha256:3033094dacc2787ef623726ac61c64d58933d75058be947d76e9c79556e79f1b

Observation 9a04388b-1d84-45e3-b2ab-1a747ca1e589 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:45.426399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:16:40.929427Z digest=sha256:0ff80d401dec75635b721f88022532c5f414623f016247351cd13f04bda097e7

Observation 8cd1bf44-9bf8-486c-96b6-084e41673e6f · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:40.998220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:40.998220Z digest=sha256:08364fe623149c0f2387a5d11866128246e1a47e8290805b6b7e8c1ef7e728e1

Observation f66f4c85-15b0-439d-8ff4-9ce75c957d13 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:45.178131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:16:41.097932Z digest=sha256:e30851b11f0b06c713fa55c0dbc1eaf02bc97a3d3320f9fe0c8675aac4822b24

Observation 663a592d-2f60-412c-b8b6-82a99ede768d · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:44.909157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:16:41.196983Z digest=sha256:5326752f08e68887a394544a9b3ea662bbe9339051f843e9ad2cd5799f273932

Observation a079a4d3-391c-45f9-b92e-54d21c51123b · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:44.754722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:16:41.296524Z digest=sha256:523623af682375135dac652fd217bb97fee8be896f48d6b676d0b4374068c570

Observation 1a8fe988-6039-4b96-bfdf-7a2269b3c748 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:44.592866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:16:41.372786Z digest=sha256:13d688d9e91185eccdb4401a09e372282aa73323a4a09b8d611385591cecc516

Observation 52e8d0bc-9369-454a-a3b1-c7267ef526cd · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:44.434457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:16:41.464856Z digest=sha256:92937cd2502397845331d30756ffa396d59c9f6e167ed279b2e28e21529a8887

Observation 83f3bebf-fe43-41b5-ae85-56d670b907ed · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:44.304156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:16:41.553509Z digest=sha256:57dee44f8195f94d129bb2b1dcde6711100f584921adde9f7662a73fed512d24

Observation b100b888-b2b6-4049-839b-082a781e0eac · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization LLaMA: Open and Efficient Foundation Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:41.649120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:41.649120Z digest=sha256:de673063632c436c3f41cd53ea7b9299848db7f543779dac3b4955e267c62743

Observation 868b6744-535c-4cc2-8824-f3fd5b1dad29 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:44.083353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:16:41.769077Z digest=sha256:bbdf639c3075c9aac6498145f7694cc0a047dc45a7e53ef0768284f1088ac164

Observation c4e447d9-8d09-4019-8676-3864b2e16da2 · outbound

This paper cites Deep Graph Library: A Graph-Centric, Highly-Performant Package for Graph Neural Networks.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Deep Graph Library: A Graph-Centric, Highly-Performant Package for Graph Neural Networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:41.882697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:41.882697Z digest=sha256:8374bf542e2ba2f87faf7702c7fa71bb879c59b18c8b18827fe54514242b2e0e

Observation 7924559a-cd26-4bca-b90b-b394ea4b6781 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:43.891966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:16:41.976279Z digest=sha256:ebe8b21b9f6abb6c13960cc65bea1647f07e7a606ed5ef5f258b4dfe6e41e0b7

Observation 0fa3dd3c-f37c-4e1f-a563-75bc8879d05f · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:43.696181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:16:42.074404Z digest=sha256:ced9661df9ecb98f2442f1abbc715eeb794e8d875564e0b70e82b40266894956

Observation e3a1e404-48ed-442b-abc8-6981acdd4d24 · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.177025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.177025Z digest=sha256:4f776a30e8e73a1395d36a7d63f21f48bc18bccfaf900bea45b167120a4313ff

Observation 8bae7c40-8cad-4de8-b326-6a4e6348b460 · outbound

This paper cites Post-Training Sparse Attention with Double Sparsity.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Post-Training Sparse Attention with Double Sparsity

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.302352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.302352Z digest=sha256:673364d93648e52d48a1864acda02e411b3dfde503be15a098a5ae1e195a21dc

Observation d06da1e1-b521-4fa2-8675-91a428ba9431 · outbound

This paper cites PQCache: Product Quantization-based KVCache for Long Context LLM Inference.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.391215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.391215Z digest=sha256:9acc5b4d2badfe2662dd61783aa21d8f20b4772329030768080019f17a8eeee2

Observation 32ac6e75-29cb-4ae5-8a7b-7d439bd44850 · outbound

This paper cites Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inference.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inference

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.476185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.476185Z digest=sha256:10e4ada38d9cd4121246f8f1ee1e989ca69a705e826090d812982a51a1f214d3

Observation 11eb4041-968b-4b28-860f-0d80e8b8c153 · outbound

This paper cites an unresolved cited work.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.573762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.573762Z digest=sha256:914c22b1ab5d9813b78b8701412130a1df9c3ca7ff2ec4ad7f43bfbde05f8f8d

Observation ceee1b52-8eca-48c7-8c1e-6345ac942997 · outbound

This paper cites LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.683728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.683728Z digest=sha256:cec0f2c06457d127bb4707217036e4c1c39c71c2e0401a51e093e24280c73c01

Observation ef746939-1c9f-47a3-9895-200ca4c889e3 · outbound

This paper cites Barrett, Zhangyang Wang, and Beidi Chen.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization Barrett, Zhangyang Wang, and Beidi Chen

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:43.494267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:16:42.793624Z digest=sha256:b283a27876abcbbf065c158bc46586b61385c9ed11ba3a7be069934c94e36ee9

Observation ba9f6c0b-ec83-4895-932f-6a81fdad01d7 · outbound

This paper cites online" 'onlinestring :=.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization online" 'onlinestring :=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.896424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.896424Z digest=sha256:2f143a9daedcbe55bc59de3b319c0a120473c3d1c60634d945fea54992e4527d

Observation 5a44742e-afbb-46f6-8f28-6abae5b91aea · outbound

This paper cites write newline.

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization write newline

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:42.984625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:16:42.984625Z digest=sha256:d6f21f192084b39b93b3af1068cb052c3b88a8585ed14a42e92e1ef921a8b63e

Pith citing papers

Observation c81e1b24-b203-471f-98e7-755364f2c47d · inbound

LOOM-Scope: a comprehensive and efficient LOng-cOntext Model evaluation framework cites this paper.

LOOM-Scope: a comprehensive and efficient LOng-cOntext Model evaluation framework TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:12.553567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:45:12.553567Z digest=sha256:8decdbac1d39e8fd9cde186c70053b5a6362d8b6339f0df725608bf16c581b79

Observation 06792c92-3127-4413-b66c-d107532327c0 · inbound

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference cites this paper.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:55:13.365091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:55:10.372511Z digest=sha256:f7cdb7d94ec8b63c3ab0f9ae9cba0c32e00623381fe3fa942899cc01eb3eb3bb