Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-09T16:06:26.450483Z
Paper Citation Record · LEDGER
As of 3 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2605.02262.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-09T16:06:26.450483Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T02:23:36.996872Z
A source-named dated measurement, never combined with another source.
Source: cited_works
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3cc82a3f-a3c7-4333-9d93-0fbbc2db7041 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 37e8af78-8e33-4093-8fc4-bf7954f83427 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Matryoshka Multimodal Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2801f0d3-c11c-4d8f-a085-9b0edaafa0f5 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bf729992-932b-477a-8aa5-4dd3338c15b9 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 9469f6a4-1c2f-4707-b12c-1283696cb9e4 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization In13th USENIX symposium on operating systems design and implementation (OSDI 18)
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 79873381-857b-461f-b6fb-ce368b428d54 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation c7d240ed-94b7-4519-8a6b-4982a3372d46 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 246cf9c0-71fb-481a-92aa-aa04dfc68079 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 6c7117ef-d262-4868-a47d-549f34a798ce · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b8c02504-2085-4f61-9442-fac2ff25199f · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a31d2bd8-75ae-4446-84fa-f321c3bdd100 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 0ced4198-9eb5-40f7-b567-d41f36521502 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization QAQ: Quality Adaptive Quantization for LLM KV Cache
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f0a8787f-f301-4d6d-ad0b-aba83ec3c8d5 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e468c763-14bb-44d2-81c5-a68e541c5363 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 877ab0cb-2e9c-4fd1-b649-cb6c0967a29c · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8d3568d7-5429-4c68-8b5c-e778931cea00 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation c7fc36aa-0ad4-44f9-93ab-ccb34322ac62 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 99e832e7-3ec7-4588-b1ac-415de378cc29 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization SqueezeLLM: Dense-and-Sparse Quantization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 86068542-c261-46a4-991c-eef1232bbb50 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bc9b8e46-5a4b-4734-9921-5862e941df57 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 533e1580-9583-4ec7-adbc-a54c3be68e07 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization LLaVA-OneVision: Easy Visual Task Transfer
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 084dd3fc-267a-4959-be44-6851cc5c383e · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a4207f54-b322-420e-a23d-f718b8e77c58 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 599a310e-c1e9-4c0f-91c8-c3f1491d1c66 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d5e85bf2-2d21-4c4f-bb89-79d5993cbf97 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Manuscript submitted to ACM WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization 25
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 4b62f573-e195-4e30-b612-c1e729d42fee · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation acde94bc-2b0a-4488-9bb1-c541197753db · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d167c778-4b1e-48bd-a622-532b92867be4 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bbbf16e3-2b8e-420a-9988-0793452777e2 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7498c618-cab7-4469-9ea9-993aed986837 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f5b11e22-c0c1-4e0a-a04b-3ff371b580f7 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation cf004705-ee7c-4880-a1fd-06a3ee56daed · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 25f471b0-a9c3-4e60-9159-9bc9c90cebd4 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7ab13a72-035b-4937-a7a4-e94e0ac48e64 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bbac7741-6968-45af-9c5e-35fc5092ab46 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2e46c165-d1ad-417a-bbfa-54f63fba6662 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Fast Transformer Decoding: One Write-Head is All You Need
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 0942019a-774f-4b29-9e98-9f853ee815a7 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 3bf906e0-6e71-48f3-98aa-5f0ec809dcd8 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization MileBench: Benchmarking MLLMs in Long Context
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 201d4e8a-699f-44c6-9876-8e55a9219ca0 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 66c56521-712f-4369-a62d-1b327c02f325 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation df4a4cbd-303b-41ce-92c1-e4d7b4827f3d · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 12b07bf8-6a60-43c0-8d95-f0aefed8219f · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 18fd78d6-b215-4251-97e1-d6af1a971905 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation c50bfd0d-a49f-4899-80ec-f2b49a547df6 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation ddf73fab-8b44-4c70-a080-0e82e1c731ad · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Efficient Streaming Language Models with Attention Sinks
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a37c20aa-2276-4491-ba22-a13e2062233a · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8ed0375a-ec45-4b75-ba15-d609e1622694 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation eb21f3d7-a08a-4c22-b42f-56b6c23b5bcc · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 4a33b995-8f1c-4cea-bc60-22d1de9a323d · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 266804d5-3649-44ae-90de-c440a4082888 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2cd59fa8-809a-4504-bb82-7deb96693212 · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation dd333923-b297-4ee8-9113-4b6a73a2c09a · outbound
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 3fc3fefc-78ac-46d8-9d2e-299633755a4d · inbound
Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.