Pith. sign in

Paper Citation Record · LEDGER

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin

As of 11 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 0 inbound Pith citation observations for arXiv:2608.06411.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06411 v1

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:30:14.535753Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 105 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved85
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3fefcc0-c1e5-49a9-8b02-7fb1be112d50 · outbound

This paper cites Qwen2.5-VL Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.184628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.184628Z digest=sha256:3387d40f1b33ce66c392889ea5da47e72b7ced8c2a684bf69f4ae214295dbb88

Observation 5bf4cec4-7fbf-4622-8715-6616263e904a · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.189388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.189388Z digest=sha256:ea94e4262407b1196b45128c36cc40aeaacc54b611835494b8281ccf2d8765e5

Observation 6613c666-b789-4a1d-ab34-393cc274f899 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.193471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.193471Z digest=sha256:ef2be877e876cf14f90f5b64e7fc8f8d189d08f57208c979fdb0a4a62ccb9ac9

Observation c552ff0a-8265-45dc-a260-ef63ec3ca58b · outbound

This paper cites ICML , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICML , year=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.197067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.197067Z digest=sha256:b85c77ecaa978bed811c5aeaac1b276d14a06ead80a608e4b9f786c563e57752

Observation 9216dc4a-716e-4300-a6a9-785c5b82ebb1 · outbound

This paper cites ICML , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICML , year=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.200763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.200763Z digest=sha256:a4f73572185f60ecc0dad9d3e3c1cf45385378f5e1d492fc32bd026042f7c6ce

Observation 7f200349-2f41-498c-bbb6-00669a2110fe · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.204559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.204559Z digest=sha256:692ce60e65a496f12449e98855baf84e048a4a5ab81275c72ce723bb1ddd72f3

Observation 655ff2fc-a489-418d-8f29-ccdc15ad5dab · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.208606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.208606Z digest=sha256:6a0da4e0458693ce8d101d35f4b1c4c8ed017bfbc247a8a297797072eb74603e

Observation 788c189a-8fd3-4e98-b920-0f4319ec14bc · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.211519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.211519Z digest=sha256:1913b8e0ba61db1316cd88884c32b3e8bd48c8324eff41951d44901aadc870db

Observation 36222e3f-701e-4743-b392-5722206979a5 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.214635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.214635Z digest=sha256:0a5ce302e72f6ab011664b7f81b4e2f8632d04e19075864d120d39f58f412415

Observation caa82ddf-0c37-421f-a9fc-9152b55e2ff1 · outbound

This paper cites an unresolved cited work.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.217581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.217581Z digest=sha256:143e6d9c20a18a48e339d0e9370da9d7fc7982aa77f3c4695a2a38abafc90b46

Observation 456cabf5-7ad3-4338-a54b-94098b316d52 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.220359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.220359Z digest=sha256:1d8cddc777d09b91df0ff380c894eba863baa10fd29ee8e3ad58c9e35b0a908b

Observation 4a5433c3-b826-4b3f-8dd8-2695fc491d56 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.223732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.223732Z digest=sha256:8927f2ef643b1e2424589e379160e1c8da5c27b206d99b359d32465496156ddd

Observation e3ed016b-e8d6-4909-8094-f1906e58ba7d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Gemini: A Family of Highly Capable Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.227340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.227340Z digest=sha256:9fc1fa7b34d6b0543bfe5198574dc65d01ce4e2861928e6862dc87ec77885e8d

Observation 498c2365-7b92-4d8d-a347-8476b14d3987 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.230915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.230915Z digest=sha256:473d4886e8677e9ce1054f8df97863841afb2679365642dc7ce99bda7b2d46a9

Observation 83c95443-0f89-4b8f-b099-9c4efeda70fe · outbound

This paper cites an unresolved cited work.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.234665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.234665Z digest=sha256:d6de0bf5d0ebb984453b3b9fdd9772d4c9bebbe6a9b9aa2e2b4ad7cf39784e76

Observation 04fcfba3-a3aa-4fc8-be22-c34dbc4866f3 · outbound

This paper cites Token Merging: Your ViT But Faster.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Token Merging: Your ViT But Faster

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.237915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.237915Z digest=sha256:3a6c9e0214658450714ac681ea865271a793fb145b099aee6719ada1b2ba5774

Observation daddbc57-a323-435e-9fb5-7d783b1be326 · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.241516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.241516Z digest=sha256:75d3accf8b06508dcf0646bd78451c490a1cc99b781a8cbddffcfd9d3e5036a9

Observation d0f3ba0b-a1e1-49bf-ac40-98a4e7670049 · outbound

This paper cites arXiv preprint arXiv:2403.15388 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2403.15388 , year=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.244923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.244923Z digest=sha256:368278afb05a867c752476abc43e36f8069027089565a724ff7f5525d4589182

Observation 4cd2cbfe-3851-4f2f-9a68-ff4bc8485b30 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.248316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.248316Z digest=sha256:8ae8b24fab5fb543720a6dd71907ea3cd4cfc9c57f8dca2f1432b13c13dd83c5

Observation c2530b53-5a83-4dd9-9946-b4b97b816b33 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.251891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.251891Z digest=sha256:e9e43a08586d867cc09deb45bf8498a2f184a5942a2016166000c46880ebf641

Observation 3c74d0f9-56b4-442e-91c8-51e2e8a3ec73 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.255625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.255625Z digest=sha256:373a518b01d20221f16687c74806a0e4b158cdec2f32d519b7f1280cb885dab9

Observation 5c9d12e3-3039-410b-b995-177db0a4e534 · outbound

This paper cites arXiv preprint arXiv:2411.17686 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2411.17686 , year=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.259308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.259308Z digest=sha256:c2cc33e6e80945a3ebfd23cd039a1f737bb9d0b1c67013d12558f5b9f8b2970e

Observation b05b1e3c-ad65-4f76-a39e-be00825a9b60 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.263104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.263104Z digest=sha256:cd57325a6d76601cb2651382c248f6cd3b33dcb623f9e0a2a91ce61fd55d4a5f

Observation 73fd6f05-f54a-4edf-a71b-92ff8ec78783 · outbound

This paper cites Qwen Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.266859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.266859Z digest=sha256:5d55045ea8105ca5f1df02f24468543296143790adebe665fd1afc37323b82c6

Observation 6ed976d5-52fa-4f87-b350-6d6483148a3a · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.270569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.270569Z digest=sha256:440219d5859572a14f66c22a4e85be481725ca1e493bc77acf086114c623179c

Observation 100f7f4a-5ab0-4897-aeac-f15ce15c3ddc · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.274367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.274367Z digest=sha256:20ea76e563402be8acffa67834171bda8f106ddff823b128e96426deb7ec5151

Observation 51982df7-5667-4d16-8a10-81e67ee8acb4 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.278391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.278391Z digest=sha256:bffe8600e9e9d28f2705487fc145073c4d86763e101e76c2f63136f7891c4fcd

Observation 5bc6395f-69e3-4fb0-b1e5-d9efa37ef0cf · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.281720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.281720Z digest=sha256:ac7479d173d5c028a4a426716bce042bd86f6c775862cf64deb7f900b72094af

Observation 6164b773-31a9-42e1-a6bc-7d14aa21dccc · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.285282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.285282Z digest=sha256:2578401c7eb1688d7717444b0dbfd326a607c926242512345b353d4330aeea84

Observation e270aa12-acc4-49a4-8290-05c05ff6845a · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.288837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.288837Z digest=sha256:d4f447d595d4b34b3259d4774aada6998a366be12e8609e2bf6668b31a852e30

Observation b888b749-faa7-4be3-b08e-3024b0e967d0 · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.292670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.292670Z digest=sha256:ff67da1813c88d10ff94f91fee653c599b82e9281e25eed5c7cbc898f6691bd1

Observation 5f172363-a22b-4c8d-b29a-583864f922bb · outbound

This paper cites ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.296106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.296106Z digest=sha256:498817596cd363c42f0bc4d26df8306fd123ce4198af9f141b0f275466a5bf39

Observation f7e6a82c-d1ad-4b30-9c0c-272f3de478b9 · outbound

This paper cites Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.299362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.299362Z digest=sha256:49001c3ae931844760cc420763147f93e52d1ad898496254bb5cabafc27192d4

Observation 608d509a-bded-40cb-8ee1-8cb0781c9a7d · outbound

This paper cites InternLM2 Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternLM2 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.302724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.302724Z digest=sha256:a2a9dc0d9355b7302d4d37d6530cfce7f464359fdb410b404b8348b4890d8fd1

Observation 92b9e655-b4f4-4e0a-9a03-f79c0a52089d · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.305742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.305742Z digest=sha256:ad249c515f879ac5e100af95f48b739f54b57ed417e830137642a9a4a04da0ba

Observation aaa4e496-0e8a-42c7-b286-2b822f85099c · outbound

This paper cites WACV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin WACV , year=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.308686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.308686Z digest=sha256:5e4a5f2f4308f0d66f42138f6a155548e40b98f98e055d0e5ef48a507cd7fd4f

Observation 5cb99119-9302-4632-ba03-ef4f6964d551 · outbound

This paper cites WACV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin WACV , year=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.312135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.312135Z digest=sha256:183a112e28ec4cbe1e5df5956a28d0000ce3f48926aaad0b0d3d8e2d1823cde8

Observation 678ccdbd-6575-4c75-9c72-c3410921ded0 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.315460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.315460Z digest=sha256:fad60dc52a5bd4057de1333c71f643a02b17cfc13b488c86e99e86580e59008b

Observation 9567cdd9-78a5-40e5-b775-3c853645e88b · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.319438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.319438Z digest=sha256:21d5369ddf633046e8615af5f599041569832605c38eba6311a59d85f965e403

Observation 8060f9d7-21bd-49d3-b775-04767abd0b3a · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.322726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.322726Z digest=sha256:df7731c4bf5735efb6f25b0902dac72cfb2ca467f611a4a8ff7c2baff4c25e5c

Observation c9bd4578-b93d-4a31-a85d-2ccab01855f4 · outbound

This paper cites EMNLP , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin EMNLP , year=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.326019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.326019Z digest=sha256:171e0b611dfb16996ad0264234e975f483e72b7558edad879c37371c29263450

Observation 27495724-e151-41ce-b800-bc8f69d17a57 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.329346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.329346Z digest=sha256:dbe8e3a794ecbdf0267c6c7b9a3b45a84cb80fc5518ed5dca525e35fb468e015

Observation 3c74206c-3482-49aa-93da-85f0681ee83f · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.332660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.332660Z digest=sha256:888ad9e74e8e9c65d70a5ea752ff36f52b6c42569ecd933518c91edf79ac92b5

Observation 8b16f436-314e-4cb7-879a-1c7375f87e85 · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.336324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.336324Z digest=sha256:ef0e2eb0a2ebc9b3f1fd7c8e95eb9f221d8f89b686b737dff7ae8b1d047f6a2a

Observation c8e980ba-49ac-46a2-9ef9-9724a899f10d · outbound

This paper cites arXiv e-prints , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv e-prints , year=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.339687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.339687Z digest=sha256:fb089260d84e94578848d04a7fa60e6bf52cc23cc21d565146518f43bc5a92b6

Observation cc095e69-dfcf-4ff8-96bd-0af283339d9e · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.343274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.343274Z digest=sha256:d9b32934c196ee39bed0a8ded8f013a814faf45133979e9d67ba7a79f7dc481b

Observation d44975fa-5d17-40c0-a7a4-11d836af12bf · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.347023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.347023Z digest=sha256:0026dd93a5999ecb056d12138f7f7e8ac5d567e953c46bc3264b1252088d066b

Observation 0a18acc9-376b-4343-b92c-c026baf70eed · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.350680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.350680Z digest=sha256:517e6ff9d82c18a97f22cc8a0888e7695d9c5fd90049b0fa542cd998cb5b3ff3

Observation 20e53230-899e-417f-92db-a8dfd19d0703 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.354396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.354396Z digest=sha256:2f3ca7b2edf78d72372f90abf1dfe8a20ca2501a2e5d7edbb26c4a11e84abed3

Observation d406f4c0-480d-4dd0-8802-74a953a56716 · outbound

This paper cites ECCV , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , pages=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.357716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.357716Z digest=sha256:fccc3b64eb24f4a4fbf27e023bc0d9584f68e7b212c9bc1652671169ecddad42

Observation ef15e32c-ce71-4f94-aa36-064ccacaabc5 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.360982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.360982Z digest=sha256:847aaee2a23fc59ff68fba1f46ee85ec0f59b872b2bc70598d971640ecb0e6f0

Observation 18ef600a-d58c-4cbe-8cdb-1238e205f75b · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.364525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.364525Z digest=sha256:abba49486a32059dfff4d5cb38faae70c35d59cd33197072c2beebf4934eb976

Observation 5bbe1032-10e0-43e0-96f6-cb7127e9fa02 · outbound

This paper cites Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.367925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.367925Z digest=sha256:4f5b7154061fa40429216adc6593bb854750bdc85ab6e432ae368ca5b937578a

Observation 57db35f6-fe3e-49e7-a9fd-989fbf4a5cde · outbound

This paper cites TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T04:30:15.167292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.371579Z digest=sha256:f9aa0b3403dcbd5262191df6e9914f890a5226cb1388fdb72ad56c469d44e615

Observation c4bfe39b-061b-4146-86c4-abc0b20949bd · outbound

This paper cites FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.375574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.375574Z digest=sha256:48e119e234052b2b7effca7f95b1595315d2d4ddcf369c91fa9584b3893d02ae

Observation 72812fb3-2681-4c41-b044-677deb2ad573 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.379313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.379313Z digest=sha256:b56d7532eea6a22bc16b5fba5c80b6d8c0ed37452555b7d87b0ecae9f2257907

Observation b7c065a8-5ccc-4a65-a1a6-4ab6c37f6b8f · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.382832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.382832Z digest=sha256:550c640ffabcc16e3bd7b708bbd54eb7176d338d0c2db482475d2a24dc7d666f

Observation 19a85498-d981-42e3-8a41-06ba36b4a12c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.386155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.386155Z digest=sha256:e86afc848762e90470391c7c9502af5189dcfd35494cd283bc9bf9f89935a99a

Observation 868f1a82-2656-424b-b0a5-40c2d9dde5d3 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.856172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.390028Z digest=sha256:12558a5e6666cb83a554e250e97f5466743b8eb498c25b649eacd0adcd021b9b

Observation 709f01b0-9b09-40de-bf1b-d9a8590fbc55 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.393401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.393401Z digest=sha256:68861dce3999f2341b6fcfe6676821a09028de088b96d34eeb14310ee7a980d4

Observation 444a931e-74ea-4114-b2aa-e45f19f5938f · outbound

This paper cites ICML , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICML , pages=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.845287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.396351Z digest=sha256:1bee3352714c9fe29629c1fe7b6f1ae11839cff2dd5ceef98f3e63f34462394c

Observation d3efd223-49a3-4e6d-b047-506a4b0bddac · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.399358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.399358Z digest=sha256:e0558845d76e27c34212ff01f5730c9e7adda82d165ef0497a1d44b1b317b07e

Observation 035b1fdb-dc96-40b6-aa2e-e4ddf1c81ea2 · outbound

This paper cites ACM Transactions on Database Systems (TODS) , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ACM Transactions on Database Systems (TODS) , pages=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.834253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.403709Z digest=sha256:ec1c3d04dd395535b58f2ee309c627e2e5586855b589b713f2d0cd5d800c0b14

Observation 6adf513c-42ee-4652-8ca4-90b192613326 · outbound

This paper cites GPT-4o System Card.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin GPT-4o System Card

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.406724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.406724Z digest=sha256:e019c21230413879ea992c8543d4a469b8f4d66b025444d5484d36e608687b8f

Observation b8db08b1-f786-44c5-981b-5861db570550 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.409858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.409858Z digest=sha256:b6db9245ca61f7859fdd9a7d37288c386ff9b67aec68d4514417012b6be40c1c

Observation 2d548e17-b98a-459e-b572-53109317b777 · outbound

This paper cites Proceedings of the 32nd ACM International Conference on Multimedia , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Proceedings of the 32nd ACM International Conference on Multimedia , pages=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.414203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.414203Z digest=sha256:0f6d474a1baf028a46b39c1a840e73bf85535ad833ef0992c3ecce2d2f30a72e

Observation 4717ca4d-4ce6-4747-b553-94e6f052838b · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.417640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.417640Z digest=sha256:bdecbaf6c2d95877e07a944add6a364ad916081f27a986092fa76cb59f480141

Observation 7d630164-4544-449e-8ca2-cd2eb7387854 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.421428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.421428Z digest=sha256:170a812f1b8079202bb48b3b19aadcbfb958ba492b95dcb64d3046bb6a9114aa

Observation 2bc0c6ba-b6c3-4fb4-a600-100a0fafb2da · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.424914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.424914Z digest=sha256:bc5875b2800575b8eb5c2f84953a5f9778b71689382e06b37dfdc143fb479758

Observation db8c182a-0f84-47e2-98b8-c2c81904cb62 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.800164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.428501Z digest=sha256:a413d2055c5e36b44a10c5ce1855dccfc6bb7a56343bcec2a148386d2acee061

Observation 457f16e9-e01a-42bd-a39e-d79d004b12fc · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.432345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.432345Z digest=sha256:287852a2ab9b9876388e677108d4b2dd7c8e283e6082e9dd494f29d93c57512e

Observation e148b1e9-b912-4325-bbd0-0e0adafa7149 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.436166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.436166Z digest=sha256:7719da8c2484b9084371e5a38a01e15dbbd886ceaf9e960462a8165f0aa98121

Observation 4f4ca747-1fac-483b-b8cc-40afb9b37933 · outbound

This paper cites ICCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICCV , year=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.788679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.440011Z digest=sha256:b1dbb3cb301033e8bafe578a9053c888ffddfc2a9dcb80eb56d896c951f98485

Observation 289f2c93-275c-42ca-99dd-849180301cb3 · outbound

This paper cites Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.443529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.443529Z digest=sha256:f96e8b197f76304f3e722b8684198399ac93ef5eef7bb8b8bdf019016658b9d0

Observation e7414c56-1099-4ff0-8fe4-97423e1612eb · outbound

This paper cites Qwen3-VL Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen3-VL Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.447236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.447236Z digest=sha256:a231a7db7d03fc4a5cf31db6c9371ca28da2f66ed4922e9cb60aaccdcd78a94e

Observation 34ee7d7b-0e4d-426e-92b7-98060254abd8 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-OneVision: Easy Visual Task Transfer

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.450759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.450759Z digest=sha256:207422ff176e36c1ecfc4fbf41f17ba7e984e660414c21c48c3442765cec3252

Observation 04c5c2e7-9d73-4d01-9880-688927c997e7 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.775868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.454179Z digest=sha256:90b66380806fb5a9aa1ae6678c42dd2ac82f305503f0298b27924cf3f59f2835

Observation b7787d07-583c-4a18-81f7-f54d3e940750 · outbound

This paper cites AAAI , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin AAAI , year=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.765249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.457768Z digest=sha256:349157d1c8f5bb2ea005be1e523c1a3670865c1bacbee0ca1d874c9caf7f29df

Observation 3ec38fa5-dacd-434a-b02d-69faf847e68e · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.754472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.461441Z digest=sha256:37c7d433d7c1e0d01d2e0d728484a6b66b3b44c37efffc97e7325c34cc2fe173

Observation 40a305b0-f104-48b5-866b-d91e4a27bc00 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.743459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.464845Z digest=sha256:5998ad57975b336cc30731d099fda9a7f46bfcd285ffbd7b59274f8fbd82ca39

Observation f66abced-ddcb-4954-8520-3e15de978fef · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.731968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.468161Z digest=sha256:76e3ca0b462e7a8fafaadbabc856ce235dd63dd2b5e5551f687846cd3ab2a4c0

Observation caa523ce-41c7-476e-a371-1ba415be5e77 · outbound

This paper cites Seed1.5-VL Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Seed1.5-VL Technical Report

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.471570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.471570Z digest=sha256:5b95e0b007263606d7cc2512ff7303ec166401240b951d92111296d2f7bc78eb

Observation 3e278458-c443-4484-aac3-85b23c675074 · outbound

This paper cites InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.475211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.475211Z digest=sha256:246cb507c73a456be5ab3563b4737b6ce9eaa33d37ee12b42d35131341c891a4

Observation f3da428c-2a7a-49f5-ad92-2ef1b917e3e8 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.478903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.478903Z digest=sha256:4ae21bcb9b7b22913fe49918be3196a043a83ade913f8b6a315e4eba57b803d7

Observation 33aaf4d2-def7-4a77-8995-12df540f570b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaMA: Open and Efficient Foundation Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.482592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.482592Z digest=sha256:9b7f8e080f7ee88692d59168a2486b5d6f66e49b6bf0711a02e516d36a9020f0

Observation 019cdcc6-3c79-4eb6-aa13-02a89bf8fe65 · outbound

This paper cites DeepSeek-V3 Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin DeepSeek-V3 Technical Report

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.486359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.486359Z digest=sha256:8de1116f91f2636bb22aaacaab67ec18800451c54b81d3e1b88315a25989e8c2

Observation a9fe5aea-29c6-4dcd-8398-07cc5861b1cb · outbound

This paper cites Qwen3 Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen3 Technical Report

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.490105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.490105Z digest=sha256:08d56367d6c0810d27322a9ac1e63036296ba888e800b1c48e74baf49364904d

Observation 4d21e384-dc0c-4128-b777-d704c90f14bb · outbound

This paper cites AAAI , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin AAAI , year=

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.712087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.493811Z digest=sha256:c6399afd4e6d66ed5c415502d68dcb498fa58a4acf6f6f8a115a0e1159d5b0cc

Observation 044e74ac-68f1-45b6-b553-57d2505b46b7 · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.497187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.497187Z digest=sha256:41a1feb336ab756b7c8e8b76f0ccaa65e8c5c9124de3bae26f15e3bd4cdd1d99

Observation f2c659c3-613d-410a-a1eb-09ec16f00ad9 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.500265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.500265Z digest=sha256:1bd87cf63c191b099750868df879ee0e155e8141f19fde05c0701f1328b9f5df

Observation c403b17c-df75-49e6-a8f4-2bb4ad3f6767 · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.503529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.503529Z digest=sha256:460b62fc5615e2fcd1300b65fb722011419b4f6c14fbef6e5a0d2d2d56aa273f

Observation 8ec82576-f1d4-4563-a1ce-99ee577f7eab · outbound

This paper cites SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.506806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.506806Z digest=sha256:79e574513018c1a533eaea99c49b30304b7592b66349e79cd3d0f4544db8327d

Observation 204158c3-e71a-4351-8031-9963114faf95 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.510251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.510251Z digest=sha256:976ae005ef80c7a5f9cc830e4a92a043df231b6dadafa2c6326e6983c92319b5

Observation 41a5184a-a68f-454e-a9cd-4422249b0d01 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.686593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.513961Z digest=sha256:18279809cccb8c2b07a50e10b06dc38c0825d077ca01bd2f9483103e6996fcdc

Observation 09aec212-b394-4a0a-ab5a-ec21b1b564b9 · outbound

This paper cites OneThinker: All-in-one Reasoning Model for Image and Video.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin OneThinker: All-in-one Reasoning Model for Image and Video

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.517712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.517712Z digest=sha256:37e54bc90fabd46b680cfec4a5ac79d8c6469bf89f2a87c99d24c50b7fa8f30c

Observation b8dbd646-f81e-45ae-9aed-dae7da1668b8 · outbound

This paper cites LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.521293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.521293Z digest=sha256:a4a2845d971c77cdfa3166966944ffd694ec39f76761e5538a597e89ee6ac787

Observation f5d1790a-cddf-444e-b41d-4f6533f24bba · outbound

This paper cites arXiv preprint arXiv:2601.22674 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2601.22674 , year=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.525086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.525086Z digest=sha256:600846c48f4dfa07a5afdd1ad0e3bbd84ba9b52a939d43a517f3e36c1c370e04

Observation e7b0d891-b32c-481b-b396-e3e10cc6872e · outbound

This paper cites AAAI , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin AAAI , pages=

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.675107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.528582Z digest=sha256:864f0752ace70aa4316453d38e7b4432b281406bc43651ae5424fc718deb58fc

Observation 13f1b919-3589-4789-9d9a-62e60ffdf5f7 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Findings of the Association for Computational Linguistics: NAACL 2025 , year=

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.662753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.531924Z digest=sha256:9962625df2896bee967a0e00f5120660dbe6bd86a10695ecad363b870c8f798e

Observation f250d59f-7c12-45d5-ac2b-c74a27ad96b2 · outbound

This paper cites arXiv preprint arXiv:2602.23699 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2602.23699 , year=

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.535753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.535753Z digest=sha256:9d5501571490441c5d97529bcf10a2b49fb9ac9a8b0b4392109588d53843f69c

Pith citing papers

No inbound Pith citation observations are available.