Pith. sign in

Paper Citation Record · LEDGER

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization

As of 3 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2605.02262.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.02262 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T16:06:26.450483Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T02:23:36.996872Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact29
  • verified fuzzy2
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3cc82a3f-a3c7-4333-9d93-0fbbc2db7041 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:36:08.068151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:90d2edde99f89251229998632bb9c0275fca2ee0dbe6cee8db2ff14fb8ffa8fa

Observation 37e8af78-8e33-4093-8fc4-bf7954f83427 · outbound

This paper cites Matryoshka Multimodal Models.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Matryoshka Multimodal Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.098510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:4d6d5090b15e0efc920019b69ed3f2f7a84d3291c6615478bb8cf52f166324ca

Observation 2801f0d3-c11c-4d8f-a085-9b0edaafa0f5 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:36:18.892007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:fabefc08657bcf5d10839e01d852ce752834c33904ecd275c457b62897b11ac1

Observation bf729992-932b-477a-8aa5-4dd3338c15b9 · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.096289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:7f2fc9467790ffaa9026ff2f1538cbce8b403af1942b5268bd09077f415c8613

Observation 9469f6a4-1c2f-4707-b12c-1283696cb9e4 · outbound

This paper cites In13th USENIX symposium on operating systems design and implementation (OSDI 18).

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization In13th USENIX symposium on operating systems design and implementation (OSDI 18)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T00:56:26.147564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:95986cc2732a94aae3ef611598cf20aba54e3e699b2f71717644d8ad72ea9350

Observation 79873381-857b-461f-b6fb-ce368b428d54 · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.138422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:58b63d17f68fee76683f245f4f99374972ccd8889d970cf16f7445f394194f0c

Observation c7d240ed-94b7-4519-8a6b-4982a3372d46 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:36:08.060568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:6f1bd1d92bdffd4ed084add452a5f9c3dc20121b21a6546573c5ae7e7b35e939

Observation 246cf9c0-71fb-481a-92aa-aa04dfc68079 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:36:08.013589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:633dbd0b1a2991e95c75b2273c554afd037441820001b843762b4ea3bbccb8b6

Observation 6c7117ef-d262-4868-a47d-549f34a798ce · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.091991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:944fc910b771ed58f9440343e70fec16dba4313c7928238bce51b63a5c67bc4e

Observation b8c02504-2085-4f61-9442-fac2ff25199f · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.179452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:b8f44d6606db6d5b54fbb1bca86888181f06ebc87a6eee1c5f46a25cde4486e1

Observation a31d2bd8-75ae-4446-84fa-f321c3bdd100 · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.092490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:3b2ab5b43a1d8f15aa8d1154a5445667a658b1ec14e3da5822dd2a48f9f93016

Observation 0ced4198-9eb5-40f7-b567-d41f36521502 · outbound

This paper cites QAQ: Quality Adaptive Quantization for LLM KV Cache.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization QAQ: Quality Adaptive Quantization for LLM KV Cache

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:07.921645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:c9e4b94bde06cadbdf031d2cd221053a23ddd66457290a1b2455730caee5c475

Observation f0a8787f-f301-4d6d-ad0b-aba83ec3c8d5 · outbound

This paper cites SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:07.942065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:c42f0dd06783dbf8b57e9b8c2c8b3703b35cf825c8a2fd499cef9c56a2568d16

Observation e468c763-14bb-44d2-81c5-a68e541c5363 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:36:07.981486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:a35b8911c9cbf2c30707759c8b20798d7b430ff7de268859ed35454e545e1681

Observation 877ab0cb-2e9c-4fd1-b649-cb6c0967a29c · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:11:21.895994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:97bda5fc81d42ab1a0df7d1dea07ece02bee8907f3d48267cace5d223cc6f33f

Observation 8d3568d7-5429-4c68-8b5c-e778931cea00 · outbound

This paper cites LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.008421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:2aead640e5f91ef4308aa83a34dbc090ae9870e13649937c9ec83b920430f0d9

Observation c7fc36aa-0ad4-44f9-93ab-ccb34322ac62 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.002040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:8aeaf06d16b4653d60509be5980dd6390c07edf32f305bf95acbf52b69286642

Observation 99e832e7-3ec7-4588-b1ac-415de378cc29 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization SqueezeLLM: Dense-and-Sparse Quantization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.029916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:074ffe97557f6a8a5dd940b8e608daa8f7e7f391f0f9a25b6124fd32556b5ef3

Observation 86068542-c261-46a4-991c-eef1232bbb50 · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.171537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:710e068c8d1009d033c6796582ca63079f3af400f36bdf71295f4af9a5c1a827

Observation bc9b8e46-5a4b-4734-9921-5862e941df57 · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.167827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:846cd64d89d7f87abf352e2c42ac649f9c3d3f51b0ab14076a146d54ea20022a

Observation 533e1580-9583-4ec7-adbc-a54c3be68e07 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization LLaVA-OneVision: Easy Visual Task Transfer

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:36:08.109613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:625e6641910d7c9e68b0b3d529dc64c29ff884dc4baac1265145ad8aeecc9896

Observation 084dd3fc-267a-4959-be44-6851cc5c383e · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.104973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:32e4fc41f6c24c3d5cc32ccf103467b4630190a1f3720bbfb849612b480d0946

Observation a4207f54-b322-420e-a23d-f718b8e77c58 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:08:01.613687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:b68827714494df666e161d5a527a77c3a42f8a651f79c09eba826df6456f0bc0

Observation 599a310e-c1e9-4c0f-91c8-c3f1491d1c66 · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.187156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:08b89113522e16718c78024ba3d50b01b22d3db4ca39e102dcfdd731fc1a1efd

Observation d5e85bf2-2d21-4c4f-bb89-79d5993cbf97 · outbound

This paper cites Manuscript submitted to ACM WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization 25.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Manuscript submitted to ACM WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization 25

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T00:56:26.109518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:b4fc9779e122dbe84df23c1946cc989f05d78383287f1c97b43873008e2bf9a3

Observation 4b62f573-e195-4e30-b612-c1e729d42fee · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:07.899306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:dbfb0007230755882ce733c2a3a1175629b4a5124adc5b7b801e98a9a824b773

Observation acde94bc-2b0a-4488-9bb1-c541197753db · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.116348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:2fefce9e444cca7393c8f8b0804e87b1e9a638d8a2e4b1e56749a916540fda4d

Observation d167c778-4b1e-48bd-a622-532b92867be4 · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.080552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:1142fe5ab4b9c2af9486303714c98d0ff0dd04313c40108cc6813660ae9db62b

Observation bbbf16e3-2b8e-420a-9988-0793452777e2 · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.159840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:ccc67f3bb7e956de3d65b2a4de982daff7ceb1fa3ad92133edead9327173e534

Observation 7498c618-cab7-4469-9ea9-993aed986837 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:53:12.376678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:fc95a8396118422ba96f4e06867934dbd6c1b82cfe7fcd39f0c67b994d836c0b

Observation f5b11e22-c0c1-4e0a-a04b-3ff371b580f7 · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.076700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:4f05f4cbba458eb7c71a51463c929cf1faf5659abdd01aa1daa5ec3f25285d28

Observation cf004705-ee7c-4880-a1fd-06a3ee56daed · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:36:18.685087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:7b6b85ee66fca54284d78051b66d9bad42b07aa37675973f4080fc90d2765efe

Observation 25f471b0-a9c3-4e60-9159-9bc9c90cebd4 · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.100009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:c01a5832ec31557223d1e0f79f6de585b04f03e324cb3ef197c9d35f57c6d0fe

Observation 7ab13a72-035b-4937-a7a4-e94e0ac48e64 · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.088052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:fd2196cff845c8355277ae49034da3778c19905cfc5528c3c8f11a97e316091d

Observation bbac7741-6968-45af-9c5e-35fc5092ab46 · outbound

This paper cites TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:07.914366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:5781b22d6a620c8374e1b68e4ebc702583cbd1064fe7e1739233e8f1670d6ec5

Observation 2e46c165-d1ad-417a-bbfa-54f63fba6662 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Fast Transformer Decoding: One Write-Head is All You Need

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:36:08.114392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:05d2541ce7b2349c38d744c770ec804f2cae1c279ea9a842a089b29bb6ca481c

Observation 0942019a-774f-4b29-9e98-9f853ee815a7 · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.152173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:ee64c0d3f5e8387a0e3d40eb302bb64c11d194787ee96fe88787d445b577d886

Observation 3bf906e0-6e71-48f3-98aa-5f0ec809dcd8 · outbound

This paper cites MileBench: Benchmarking MLLMs in Long Context.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization MileBench: Benchmarking MLLMs in Long Context

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:07.948947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:78364f6ed32de51eb0eae328ec2b72459303179688db5e2befb2476e13535dd9

Observation 201d4e8a-699f-44c6-9876-8e55a9219ca0 · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.156078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:40b38bfb62eadad5145f0267a4e95caeb90da3674bde29be21f9c94b26418fcb

Observation 66c56521-712f-4369-a62d-1b327c02f325 · outbound

This paper cites AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.018815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:cf4cdaecea82aa153189e47a08a5546a676f9a548f0c0b6c7fc5087a0941f886

Observation df4a4cbd-303b-41ce-92c1-e4d7b4827f3d · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.143184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:e89a634f7cfae9df680a05628d45a596b1db0a10c2317c511ce8a3298a6a1f31

Observation 12b07bf8-6a60-43c0-8d95-f0aefed8219f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:36:07.972702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:8ff033ae7d55ff0793394e958079115c3883f7f43455fa5a6c5066212f8117bc

Observation 18fd78d6-b215-4251-97e1-d6af1a971905 · outbound

This paper cites SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.082176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:408982eb057936a8e4f25ee02d6b5a63921519218d80c9711365a612194f065f

Observation c50bfd0d-a49f-4899-80ec-f2b49a547df6 · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.084273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:57bfb8448922d4c12a4660432be451085245d8246ee93dcbe8c778fcc6aad915

Observation ddf73fab-8b44-4c70-a080-0e82e1c731ad · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Efficient Streaming Language Models with Attention Sinks

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:36:08.073518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:ab4ed0974f2d7b63c28dcebc4a5f7c16cfb234a0cf1a29d40459813ef040139e

Observation a37c20aa-2276-4491-ba22-a13e2062233a · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:21:58.165262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:10c0f5df6233ebc0d529c344a2808825b491a64da5daecf98dc2f6c535ef5871

Observation 8ed0375a-ec45-4b75-ba15-d609e1622694 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:07.887239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:d5dc3e1bfe44bafce9b6c8424b07616c1fb92ca930d30e341ec4b179e48b22d8

Observation eb21f3d7-a08a-4c22-b42f-56b6c23b5bcc · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:07.930541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:e3e5f99160234ebe1916ce225f6433e122b28d800fef5c70438ca10d1201de86

Observation 4a33b995-8f1c-4cea-bc60-22d1de9a323d · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:02:01.218694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:8559c40810a50c40aa5898dfabfc4508c02e9ead858b5c1f2f6f1cbb99e62642

Observation 266804d5-3649-44ae-90de-c440a4082888 · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.163569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:0eb8221f367552d05f9f34e1f34ce1629615cab608696320a791cc7bbf205462

Observation 2cd59fa8-809a-4504-bb82-7deb96693212 · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.175485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:74f550bcf47f524445a36685263963d97a6d45a022019bd9b5e7185231cebb03

Observation dd333923-b297-4ee8-9113-4b6a73a2c09a · outbound

This paper cites an unresolved cited work.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-26T00:56:26.183259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:be31b1645089c64e799891165c5e1d0394496cc0272fc614213eae61f9db18eb

Pith citing papers

Observation 3fc3fefc-78ac-46d8-9d2e-299633755a4d · inbound

Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns cites this paper.

Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T02:23:36.996872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:23:36.996872Z digest=sha256:8d8ec7017dde648c108a11a59ff66f20070ff2f0e034e1ce78c4b64035372513