Pith. sign in

Paper Citation Record · LEDGER

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models

As of 19 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 0 inbound Pith citation observations for arXiv:2607.22586.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.22586 v1

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:58:32.095426Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

84 of 84 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved84
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4c9fa41f-ceba-4ddf-bf19-fd591d23e4c7 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:19.894867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:19.894867Z digest=sha256:621e31bfde9feaae14bec1a4951c53906661877832d0df60ac20b29df36c3bf4

Observation 610b6963-6a14-426d-9b6a-0b323fa6235a · outbound

This paper cites Nikolopoulos, Hans Vandierendonck, Deepu John, and Bo Ji.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Nikolopoulos, Hans Vandierendonck, Deepu John, and Bo Ji

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:20.054978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:20.054978Z digest=sha256:e210ea58fd955e8d81cc3e0015bb4c9d4a05779d03dc1a375e72fe751f9cf84f

Observation d5560da7-6839-4c85-a888-927562fbf26b · outbound

This paper cites Qwen2.5-VL Technical Report.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:20.222339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:20.222339Z digest=sha256:bbca9018dfb72ef3a6653ee45afc2d6d8eefba4cbee441a6c994aa01e123b0bc

Observation 1403ce93-0877-46f5-8cd4-6d4cd083be51 · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:20.585480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:20.585480Z digest=sha256:a44c18d9e6b5dd47f38957c032c78e80e944dcd23ff03dfd3b445a020343d277

Observation 1ab748ed-05c8-4f8d-8d04-4e7a709034e8 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:20.913521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:20.913521Z digest=sha256:ae9fb033d62263f28ec9ac95478184adea089ca4e85c1d2894e47495f8a3af61

Observation 4ed5212b-863e-4cd8-b443-04f25aa6187c · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:21.026569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:21.026569Z digest=sha256:d3041bf2f5844198b71783a1f8b961952cef0c457f8f4d1d9d0f18977a55d0af

Observation 04a2b01f-4983-4b17-80b9-b58b4dbf6f4c · outbound

This paper cites See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video LLMs.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:21.450728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:21.450728Z digest=sha256:8e10e219accd9834b6a9c90df68b6c40858a39c28e45c7b6defabef564202262

Observation ad466646-dfd4-4c2e-bfbf-33543f930347 · outbound

This paper cites Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:21.569466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:21.569466Z digest=sha256:ef28bf1747e214cefef31bde71388e7368299db577bc30d34a8041f938ca578c

Observation ea860bbb-30e8-454f-b31c-de52ed70f3d1 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:21.770610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:21.770610Z digest=sha256:0860247ecd55133bfa699ff3df97ee49d85bae8e713acb9d4396e541d179afec

Observation 3df26648-73ab-4607-b379-44f33c18f53f · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:21.891492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:21.891492Z digest=sha256:5d2445899a9962abaf3b60f545ecdbb8ac0970c654d9aa5a3db328ce0779a447

Observation 672026b0-2428-4a96-852a-a19364078c1e · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:22.063778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:22.063778Z digest=sha256:c267f123fd51e26dae4d77fcc25419121f6dcfe299750186422ebb44c86a895f

Observation 40a1932f-c375-43cf-892d-0df718559623 · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:22.181795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:22.181795Z digest=sha256:56fdc9cd6adae2d94f5c7c4fe626e72e876c00a766aed3d095593bd857b659fb

Observation 713147ac-3da5-4830-9770-8dfa6e57943e · outbound

This paper cites Training-Free Activation Sparsity in Large Language Models.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Training-Free Activation Sparsity in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:22.293241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:22.293241Z digest=sha256:2ca3726ca4f41dcf76cc7a5741545377319411d0605f7f8ce5c6f7e3044f1917

Observation b2be34b0-341b-4c68-b7b5-37c4b0c39e49 · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:22.698273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:22.698273Z digest=sha256:0ac65ff8afbf994b77eec3ba917e5d0f0b52c23cae169e586918c20d6ce742cd

Observation 7fe856aa-0215-44ef-9a08-e0b5b3992028 · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:22.839506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:22.839506Z digest=sha256:126439bd3f16f93e745fb5f7254d6a9ce28066cc34734880430b5c2831312841

Observation 50b5f7a1-d2c3-4162-a318-e3c84d2dc512 · outbound

This paper cites Morse, Raghavv Goel, Mingu Lee, and Chris Lott.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Morse, Raghavv Goel, Mingu Lee, and Chris Lott

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:22.959713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:22.959713Z digest=sha256:93bb35ccddd2b27d0207d8d9939097c6384fd9bb41c911bd462ce79a4932cbf8

Observation b00a8e43-9164-48e3-8c1a-d9f65f3bad6d · outbound

This paper cites ANLS* -- A Universal Document Processing Metric for Generative Large Language Models.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models ANLS* -- A Universal Document Processing Metric for Generative Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:23.108518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:23.108518Z digest=sha256:083c73bac448ccdfa1bed19463e05d45a92ca6e5148c042dc2ef79f7c335255d

Observation c9c0e047-deeb-4640-adc3-d986e28eeb9e · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:23.423299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:23.423299Z digest=sha256:09694a02199a1093a60a3a71be1d27a3024cd2eaa74ef4b5c186446e0c0bd8f1

Observation b76d9ec2-e6a8-4c75-bb4b-7485e2118a70 · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:23.571629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:23.571629Z digest=sha256:2f28b3d2b84447e7c01ce204ae28bb14d3b8cef28bbe29eac77b35b837bc6bef

Observation 30e33678-916c-476e-9e8a-73626e926fc7 · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:23.810534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:23.810534Z digest=sha256:31716d65b23fe8ab38c8dcdb76ce8d10454e38da024ad75035152dac8cc661fe

Observation 20c69d2f-748e-4b1f-8a59-577e5f5e0335 · outbound

This paper cites SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:24.192980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:24.192980Z digest=sha256:449ee2fdac83ad3b25ad0bb43c34e78d0012c32bae25dd1a481ba9f82c8fcb84

Observation d24cf86a-d86d-4779-af46-ecac5407352a · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:24.335787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:24.335787Z digest=sha256:c7c41589f83d0d544973b69b82b215d4aa13ea57b705acad4e0aa03fac2042c2

Observation f2bd8cea-7ddb-4931-ba17-f052837b67fb · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Efficient Streaming Language Models with Attention Sinks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:24.448018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:24.448018Z digest=sha256:f201b348d64d2c080985f11fa8aacd3eb3ea8656292ebfc0c386e90dbb9ca264

Observation 3459f733-da4e-49d3-b874-c26b79a2e42f · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:24.611574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:24.611574Z digest=sha256:3472c798de762a91243d0c68a9e84c53a1a85b55c680a980574fd2eb676bb40d

Observation 7c186425-9c29-498c-8f4c-603e8ba516ab · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:24.737749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:24.737749Z digest=sha256:1c56047f824ae8c30f2c3dc1e19a2f322231b08bdfe06f930233407459079d6d

Observation 90b9b16c-7779-4e73-94e4-1064694c1459 · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:24.851839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:24.851839Z digest=sha256:64fd8a4d8875b306224cbdbfeda1448323383cad1e92a522037417eff8df2027

Observation 5dda8ded-2289-44d8-a403-4508c5cb6fa6 · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:24.963171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:24.963171Z digest=sha256:66a04eaa93c045a727c032b4599624a58212d22fa8e821223d12d1c2fe2ab4c2

Observation f811e027-85bb-4c6a-be33-3662a4401d3c · outbound

This paper cites HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model Inference.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model Inference

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:25.085915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:25.085915Z digest=sha256:219c4fa7614e7ae8de792426768875d0a298cb025ef0818b0ecc7f491ce96989

Observation 82653632-2db3-4d9f-a680-d6a8e1c94ad2 · outbound

This paper cites Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:25.244892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:25.244892Z digest=sha256:80b2cec8996a1791028420a47e2f0732fae1507152e925999258adce6b852cd3

Observation 36df42ca-8cab-4db5-bc92-6a6657fc4cb7 · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:25.477725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:25.477725Z digest=sha256:b7d03a4289b44246f73f0621112f22a0c498395a8ed20ec66966fb5430a1ce31

Observation 6922c32b-0f07-417d-a96a-34fa06eff11b · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:25.572692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:25.572692Z digest=sha256:2902c22b4e2fa722dc40d2cf23f71eadd6b56bbc7a12ce1afd809a9c46743266

Observation fc4a7285-a518-4125-97db-8435ee05097a · outbound

This paper cites an unresolved cited work.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:25.702041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:25.702041Z digest=sha256:57368c1b99c6252f4f3ff5570f538c7b12d579242f0155671b0f47e08600769c

Observation 5d211f02-76b7-4335-b5a0-4b146a7d472a · outbound

This paper cites 2025 , eprint =.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models 2025 , eprint =

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:25.791801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:25.791801Z digest=sha256:697fc96276ce82b92ef02a31cb9708fa99fdec6d8fdf8034ba2f867d4e3df220

Observation 4e6a33a1-3ea4-49e2-9cc8-d27a4e75c50e · outbound

This paper cites 2025 , eprint =.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models 2025 , eprint =

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:25.922021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:25.922021Z digest=sha256:3166c21c5a10afd678c46b5838087411a023823b713c33f51dce919bd4f2f1ad

Observation a6d8d12c-f06b-40f5-bcf6-2c953812cad8 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Advances in Neural Information Processing Systems , year =

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:26.082011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:26.082011Z digest=sha256:20c41a1938dc2211b9fbc63d892d71739850c7d7ec45c3e07e6a244181263d84

Observation f830f334-e753-4852-be3c-3b67bbe3877e · outbound

This paper cites 2025 , eprint =.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models 2025 , eprint =

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:26.168724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:26.168724Z digest=sha256:2292a074435eb5af244c91e43db6ac93077f24a1c6cb25e2d690ef972c0d6da7

Observation 246fab55-4d51-422b-b25b-f32212782cfa · outbound

This paper cites 2023 , eprint =.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models 2023 , eprint =

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:26.270260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:26.270260Z digest=sha256:229cb6340d1eda8e3539427c145807fa8728508662b7ad35f606162d55591040

Observation f0cb8b69-e774-4754-89e0-325ae7162404 · outbound

This paper cites 2023 , eprint =.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models 2023 , eprint =

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:26.385348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:26.385348Z digest=sha256:33bea8a6a6e073bdcd408f9b6300e6bfa2430b90dbf39feb925bdc22e0424841

Observation ad1ef3bc-18b5-42d1-98d5-ea8223b8c366 · outbound

This paper cites 2025 , eprint =.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models 2025 , eprint =

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:26.512769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:26.512769Z digest=sha256:921c4cd542049880878e6a46b96e9d70d3f390d8f8933683fe4878accb92d3bd

Observation 973b3d56-208c-4b95-ae63-42e351e05ff4 · outbound

This paper cites 2024 , eprint =.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models 2024 , eprint =

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:26.635446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:26.635446Z digest=sha256:adbf55cf7a13acd0cdb205c1b99b1f7d6508ab1b0f8e54c7a4185004044a02c4

Observation ba523d12-1836-4ff9-8495-eec731bfa467 · outbound

This paper cites 2024 , eprint =.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models 2024 , eprint =

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:26.768096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:26.768096Z digest=sha256:3cb3ca8008f357dfe7e78de59c770b8523b74e2ffb85968fc362401e2426339d

Observation cfb5dd04-0ff9-4297-ad0c-19a70376f45d · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:26.902599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:26.902599Z digest=sha256:fa9c390c3434017094aba0c41be25ddb246a84d0ce01e2d334281bd1a8969257

Observation 70e8483f-bc4e-4b6e-afd4-007a2ece93d4 · outbound

This paper cites TextCaps : A Dataset for Image Captioning with Reading Comprehension.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models TextCaps : A Dataset for Image Captioning with Reading Comprehension

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:27.039296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:27.039296Z digest=sha256:2afd3e5a84e948bd54a5f588de673142e32023323674f19ee2cf46a8adf68060

Observation 1029a303-26d1-49e1-a0cf-5863cb5e4817 · outbound

This paper cites MMMU : A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models MMMU : A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:27.150104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:27.150104Z digest=sha256:bbf9c6b94fa95d2404a9f9972377b2c329a5f12a691d7cd25f1bc76763a10c4d

Observation 956b0500-334f-4eb6-995b-dd6c947af03b · outbound

This paper cites Findings of the Association for Computational Linguistics:. ChartQA : A Benchmark for Question Answering about Charts with Visual and Logical Reasoning. 2022 , address =.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Findings of the Association for Computational Linguistics:. ChartQA : A Benchmark for Question Answering about Charts with Visual and Logical Reasoning. 2022 , address =

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:27.289634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:27.289634Z digest=sha256:b9ff3262df9efce2b8e132f93909a630e86348423a71a9dbcc1d91e0bbe9f330

Observation 636132c9-4b37-4837-8a89-f7b94ec292b3 · outbound

This paper cites Towards VQA Models That Can Read.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Towards VQA Models That Can Read

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:27.422148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:27.422148Z digest=sha256:3f3a0230b216564d437a4a1252265e6851cdb87be682c54f89eb5a5b4ea8aba2

Observation bb64e5c9-fa21-4552-bf88-666ffc84c5f9 · outbound

This paper cites , booktitle = "Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models , booktitle = "Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:27.549270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:27.549270Z digest=sha256:9133579cdf7bc7c5c89976c5dd05f5a10b4fef026caf31d10e1df1667cf93bfc

Observation 8a43d895-8e4e-49a8-a5fd-6d3a8be1cee5 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:27.683696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:27.683696Z digest=sha256:9b774e6c40b7891a5cb4871045395fb7360da66af6491640c9f965dddba65975

Observation d18f3edf-a0dc-4c1a-bff2-65ae682dd6fb · outbound

This paper cites AWQ : Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models AWQ : Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:27.842372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:27.842372Z digest=sha256:73f6659e519cc3d321a8c5eb13c22852eec60bdb3cc094a9605ebb2304bc23b4

Observation 1a51ac52-0af4-4c70-a11a-e24e0bd6f9c3 · outbound

This paper cites Computer Vision --. An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models. 2025 , publisher =.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Computer Vision --. An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models. 2025 , publisher =

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:27.962962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:27.962962Z digest=sha256:8c8a73ce4ba3dfe95e8f1386c0cc5a79f75932883ae4ee3afaec61b442a537b5

Observation 618513c7-b8a9-49ed-a6c3-48315be267e7 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:28.105753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:28.105753Z digest=sha256:6a45bbd5ac3f6b6964eba3d1189e50bb16a1802fd3141f2a3b68932e4ca59237

Observation 94e7f057-f949-40b9-897e-1162f22cc9e3 · outbound

This paper cites Association for Computing Machinery.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Association for Computing Machinery

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:28.185604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:28.185604Z digest=sha256:e47601c9e32b9b0dd021f93ad3d7504c391f5a6f5575393bb242d3c402328e70

Observation 9027bd55-13a7-4f06-9b86-e558141dad8e · outbound

This paper cites InfLLM : Training-free Long-context Extrapolation for LLMs with an Efficient Context Memory.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models InfLLM : Training-free Long-context Extrapolation for LLMs with an Efficient Context Memory

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:28.334271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:28.334271Z digest=sha256:cf51850f5c4b46c341d53fbfbf39d51de829208a071ad66e67f7c76b60146b75

Observation ae356680-5cb6-4441-89e3-60a5f58b5277 · outbound

This paper cites H2O : Heavy-hitter Oracle for Efficient Generative Inference of Large Language Models.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models H2O : Heavy-hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:28.466982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:28.466982Z digest=sha256:36ea5c7349bf9ff0ee4db1fe2b31e853860ab19fc1771cf5b839ed4c6e9e19f0

Observation 15fbe337-d973-4b09-b47e-e986514b5c09 · outbound

This paper cites RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:28.560458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:28.560458Z digest=sha256:b79e0afda05cbb642aa82bb805c0f57c211a47372e0a60a752c9ec078eb0ee2f

Observation 3a451ae3-ee0e-4995-b1fe-ff143af95e43 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Generating Long Sequences with Sparse Transformers

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:28.674400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:28.674400Z digest=sha256:c4c25bc67d89e9719e0cc6c3bc71329ba159811c8ff6802a7c1534e357bf940f

Observation 1429e6f6-af11-466b-b455-4c7253c6d26b · outbound

This paper cites Training-free and Adaptive Sparse Attention for Efficient Long Video Generation.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Training-free and Adaptive Sparse Attention for Efficient Long Video Generation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:28.800958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:28.800958Z digest=sha256:cff16503359eea36318b36eab58b70d80e6f123545fd59e9d9597b396da27497

Observation 5d9ccd38-9620-42e3-8271-7f82ef51b2a7 · outbound

This paper cites A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:28.898360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:28.898360Z digest=sha256:326c90b9f13a3eb1e43f1748f6dab2997fcb458b7a764eba266cbe9318ee597c

Observation 8ff85135-0005-4cd1-a816-5f5b2afa60ba · outbound

This paper cites , journal =.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models , journal =

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:28.992900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:28.992900Z digest=sha256:d7bbb0d1ae3566b3ee9dbad7dfcdb429910608e499dc2b6429d988652ee46977

Observation 01fccef6-38ff-457f-ba8f-bfd7f0bb4477 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:29.142425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:29.142425Z digest=sha256:e0ce7f2f03947efaf513d804a8da120aabcd4d2f716c111e265710086abbe5f5

Observation f69adfc5-62c1-4796-8bab-c8adb5afc3a1 · outbound

This paper cites VisionZip : Longer is Better but Not Necessary in Vision-Language Models.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models VisionZip : Longer is Better but Not Necessary in Vision-Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:29.256668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:29.256668Z digest=sha256:d8ea8518c21c0904f22303762524b1a8decf6745106356ef9812aec729d19e44

Observation bac81452-c74b-4091-9a4b-3ade95250148 · outbound

This paper cites DyCoke : Dynamic Compression of Tokens for Fast Video Large Language Models.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models DyCoke : Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:29.361640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:29.361640Z digest=sha256:fe7e965df2b0c2917c88861d6e1d2be266f411933acaa6cb7b1171b7cd6d0b5e

Observation da5de966-860a-4fa6-8ea5-f2d30ba177a6 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:29.497400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:29.497400Z digest=sha256:181dcf8508050b17f6712176e0b1a633903eda8970ef71693f0229264a48bf69

Observation eedee17c-53af-4b77-81d6-c8aafc2a1471 · outbound

This paper cites TopV : Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision-Language Model.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models TopV : Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision-Language Model

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:29.609514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:29.609514Z digest=sha256:83894341e45388344a933eb03397f30f6653676287ace280fd61db2ae0133e3c

Observation f4c52329-4ebd-4876-957d-ec47b3ddc319 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:29.746228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:29.746228Z digest=sha256:f4a0a40cda065a1d0377502d189958f6514eb93a8062a94957e585c23be250a9

Observation 95ed4cdf-9601-4f9c-a19a-5b5fe7b61733 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:29.845572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:29.845572Z digest=sha256:8bbbb37da0b796c01cf9d2717d8ca1bdd1510914772ea04673330580bd5844c0

Observation 5348cdde-7e76-4659-8e7e-6d942dd84429 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:29.935195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:29.935195Z digest=sha256:0a290f7dcbb9f9e9cb558d3273f85cc56c0686ac4e900a3f51921a5c340fdf7f

Observation 0e0d93f9-2ca8-4c7a-8d0c-75e38cda5bc0 · outbound

This paper cites Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:30.056509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:30.056509Z digest=sha256:b425bf09e86f6e5100b27990f874930b0971361276243292561b6d42c5e1c474

Observation 9ddf2ab0-fcee-4955-aa19-8f1645883cd3 · outbound

This paper cites and Li, Dongsheng and Lin, Chin-Yew and Yang, Yuqing and Qiu, Lili , booktitle = "Advances in Neural Information Processing Systems (.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models and Li, Dongsheng and Lin, Chin-Yew and Yang, Yuqing and Qiu, Lili , booktitle = "Advances in Neural Information Processing Systems (

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:30.166806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:30.166806Z digest=sha256:2a27311b886cc571e2241141db652e393a6bd46fbe79541806315a649d618875

Observation c6446cc8-25ce-48e3-b393-e3363a4e9e7e · outbound

This paper cites SparseVLM : Visual Token Sparsification for Efficient Vision-Language Models Inference.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models SparseVLM : Visual Token Sparsification for Efficient Vision-Language Models Inference

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:30.314903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:30.314903Z digest=sha256:01dca14178c85ee3612c140fbb42cadfe00a9ac4813f5697dfdff90585b5d497

Observation cf8bd247-429e-4259-a8fa-97fc5e1019eb · outbound

This paper cites LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:30.462516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:30.462516Z digest=sha256:dd69967df4befefda3a7664810da5f25d418d3a2f7b171f6374bd41e672a925a

Observation e7840632-cd37-4724-a617-bf44c1d6517b · outbound

This paper cites GQA : Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models GQA : Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:30.589649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:30.589649Z digest=sha256:2b4d30ebd4a07bf3158aa15e1c0a6d393553740c5d8eb31320f70e065a01cac1

Observation b7d7e243-73ff-4194-a052-a263e3d8d0a1 · outbound

This paper cites Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:30.699815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:30.699815Z digest=sha256:7ab9121c5efa91f0a3edbce1b2b21e385c36ca809e88616a7a38d187e9bf9908

Observation d4b9599c-978b-4803-9284-d698d4c5d12c · outbound

This paper cites Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:30.811138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:30.811138Z digest=sha256:09179ee0bb287f035a1a56e4f0bf66ef702ae9cc01698583e5e12ff16e0c75d5

Observation 41e47fe2-a1c4-48a4-b890-0c33ada6060d · outbound

This paper cites Proceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Proceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:30.936965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:30.936965Z digest=sha256:53f5a9b6ec446402e1e132bf95a00d9cff711aa9c3c247cf8fbac9998bf5d468

Observation 4ae0194c-339d-47ff-bdd1-b83036f7ba99 · outbound

This paper cites Multi-Layer Visual Feature Fusion in Multimodal LLMs : Methods, Analysis, and Best Practices.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Multi-Layer Visual Feature Fusion in Multimodal LLMs : Methods, Analysis, and Best Practices

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:31.077641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:31.077641Z digest=sha256:1ee6375636e2e1dfa8901d90ef38ea726529e971ed274fa5fc2693fe252d6daa

Observation 40bbf5e9-7530-4ceb-950b-f498f214554d · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:31.210800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:31.210800Z digest=sha256:6ea175f8e26d65b8723e495e56562756b4e92b5b263db62c2d5637f23a920769

Observation c389fac6-4e4d-4d9e-944c-cea45b386618 · outbound

This paper cites Exact Matching: Algorithms and Related Problems.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Exact Matching: Algorithms and Related Problems

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:31.358219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:31.358219Z digest=sha256:a00bab51783a21ff487fada49e376694f336742e6560ad4091e9b92b154d294a

Observation b7eee086-72ee-43e4-b1d7-6ec22f130b18 · outbound

This paper cites CIDEr: Consensus-based Image Description Evaluation.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models CIDEr: Consensus-based Image Description Evaluation

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:31.477570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:31.477570Z digest=sha256:4ca64f641dd6a8bb0d451b4895c1b1f0482e3271ab65108ff0d06adfa1dc7d38

Observation 71d68937-354d-4e03-811b-8792b63dca2a · outbound

This paper cites ANLS* -- A Universal Document Processing Metric for Generative Large Language Models.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models ANLS* -- A Universal Document Processing Metric for Generative Large Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:31.594074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:31.594074Z digest=sha256:04d1a35a04f6d5e3c24e91f9c8e5af58bb396e03983c144f628f3cd576de1904

Observation 6631fa7b-dd24-4a0c-9655-bf93d45333d1 · outbound

This paper cites 2026 , eprint=.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models 2026 , eprint=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:31.690860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:31.690860Z digest=sha256:a54b8491f9eefe1f317aa3752c019cebb84ff0391bbe4568b9f99fb1a672c135

Observation bcc1a2f5-baa6-4f6d-bce4-865a6f4ac7a9 · outbound

This paper cites 2026 , eprint=.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models 2026 , eprint=

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:31.854901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:31.854901Z digest=sha256:9ab73ece3642df06644f0cb156bc9909e2fa061cf5183f48b36ca9e24523f285

Observation adebac80-f026-4807-85a5-c17826c9a6a9 · outbound

This paper cites 2026 , eprint=.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models 2026 , eprint=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:31.995756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:31.995756Z digest=sha256:1218cceae93f213ad76c9443a0bd8d158dd8225000fdd84bd53a684318c84e62

Observation eade1ae3-54f1-4bb1-b807-0b5d172e1d6e · outbound

This paper cites Proceedings of the 42nd International Conference on Machine Learning , series =.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models Proceedings of the 42nd International Conference on Machine Learning , series =

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:32.095426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:32.095426Z digest=sha256:6e7167a47716ee1cea3274b8b2da5339369caecced3baae80a86aaa3da8e5e95

Pith citing papers

No inbound Pith citation observations are available.