Pith. sign in

Paper Citation Record · LEDGER

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin

As of 10 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 0 inbound Pith citation observations for arXiv:2608.06411.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06411 v1

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:30:14.535753Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 105 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved85
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3fefcc0-c1e5-49a9-8b02-7fb1be112d50 · outbound

This paper cites Qwen2.5-VL Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.184628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.184628Z digest=sha256:ccd0e8fcf817ee24958579f69aea66b9be54f67cf709bd56ee867fe3c279da3c

Observation 5bf4cec4-7fbf-4622-8715-6616263e904a · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.189388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.189388Z digest=sha256:41e2d781dcb3b869de8ab5f89b8d6aa79dd5112984cb97f1b801bb8aceac40dc

Observation 6613c666-b789-4a1d-ab34-393cc274f899 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.193471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.193471Z digest=sha256:77f05c02bc3dbe59f9982cf0a4a3e8a246ef747f8c16401dbca5eb8f34cabb63

Observation c552ff0a-8265-45dc-a260-ef63ec3ca58b · outbound

This paper cites ICML , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICML , year=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.197067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.197067Z digest=sha256:6e00cd9d28823a8930f75a41e10e5626807fa9d94220c01e0eff224fbed829e1

Observation 9216dc4a-716e-4300-a6a9-785c5b82ebb1 · outbound

This paper cites ICML , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICML , year=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.200763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.200763Z digest=sha256:f852e04520d29a898e36ebb5540cd89051d1a5e55e58c402279cd044a17cb557

Observation 7f200349-2f41-498c-bbb6-00669a2110fe · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.204559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.204559Z digest=sha256:7e09e368d75351c0e93cc4539f32d142b13ccad1431bd7fb425e837527e7cd86

Observation 655ff2fc-a489-418d-8f29-ccdc15ad5dab · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.208606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.208606Z digest=sha256:21863704045f436018a0e415f39026e279fe2b3cb308321a20ef70b5be6fccda

Observation 788c189a-8fd3-4e98-b920-0f4319ec14bc · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.211519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.211519Z digest=sha256:4fbc0cbcebbc7fb765ba881512b7625740f73f45b406916f2692f748759862cf

Observation 36222e3f-701e-4743-b392-5722206979a5 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.214635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.214635Z digest=sha256:280ac61121c254c66240a345239b5b8da3c107f2b7a9d81625391b18cff57509

Observation caa82ddf-0c37-421f-a9fc-9152b55e2ff1 · outbound

This paper cites an unresolved cited work.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.217581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.217581Z digest=sha256:e6e515bb0387bd80ced4d360c80908c749d1f1b58b5fe8c575002a048a7552b5

Observation 456cabf5-7ad3-4338-a54b-94098b316d52 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.220359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.220359Z digest=sha256:5919b450a4b29e43534f8b9c9700b542a7e66bfe40b9f52b4d3a43799463231a

Observation 4a5433c3-b826-4b3f-8dd8-2695fc491d56 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.223732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.223732Z digest=sha256:d14c2a67070690724cf2075f66796c297a3c47e1ab72063a07381fa3ebbb50cb

Observation e3ed016b-e8d6-4909-8094-f1906e58ba7d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Gemini: A Family of Highly Capable Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.227340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.227340Z digest=sha256:55d64932fe76ed561326eb284ee472768152129890adbf85476cf2b2c76974bc

Observation 498c2365-7b92-4d8d-a347-8476b14d3987 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.230915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.230915Z digest=sha256:f1ee1dcd94c97b4ab5f9c276baf319e6536b82873df1eda343c45c9ccdcc9761

Observation 83c95443-0f89-4b8f-b099-9c4efeda70fe · outbound

This paper cites an unresolved cited work.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.234665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.234665Z digest=sha256:4b29e8068fefad16b2c10ae7a2c444586ebe8e9d8d3a0542a0411133891b7577

Observation 04fcfba3-a3aa-4fc8-be22-c34dbc4866f3 · outbound

This paper cites Token Merging: Your ViT But Faster.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Token Merging: Your ViT But Faster

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.237915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.237915Z digest=sha256:e460604b99bbdfb38a084536e2a532efbed5e2b2d6a11151006a8f13e1d324fe

Observation daddbc57-a323-435e-9fb5-7d783b1be326 · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.241516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.241516Z digest=sha256:4bd4872414142f38937d9207f332eeef66489380da2f41981169eeb1c91a986e

Observation d0f3ba0b-a1e1-49bf-ac40-98a4e7670049 · outbound

This paper cites arXiv preprint arXiv:2403.15388 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2403.15388 , year=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.244923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.244923Z digest=sha256:942fc07008fc7768a4a1acdfa0f21eb7d3fd2acf4cf9ecd003f6942a1aff5163

Observation 4cd2cbfe-3851-4f2f-9a68-ff4bc8485b30 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.248316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.248316Z digest=sha256:8c70fb97c3bf998949e88276c31547013f0e44713096ca7951951bf3c134f2a2

Observation c2530b53-5a83-4dd9-9946-b4b97b816b33 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.251891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.251891Z digest=sha256:bd5f1cbe30a43204a52712aeda84f9c63d34320a9269fa4d22a9c51e361fad49

Observation 3c74d0f9-56b4-442e-91c8-51e2e8a3ec73 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.255625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.255625Z digest=sha256:9679ef500fd658681784067bcebe867a1e9c9df9dc91965deba73e3d957fe948

Observation 5c9d12e3-3039-410b-b995-177db0a4e534 · outbound

This paper cites arXiv preprint arXiv:2411.17686 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2411.17686 , year=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.259308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.259308Z digest=sha256:588bfd20fd3d70293bd66ca77889f6201cee6eb3cc222f5a107cb4979235bd11

Observation b05b1e3c-ad65-4f76-a39e-be00825a9b60 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.263104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.263104Z digest=sha256:018384aa5ba548b3b3642df2ee7c4341a9f2d21e378ad1849c50cb848a169250

Observation 73fd6f05-f54a-4edf-a71b-92ff8ec78783 · outbound

This paper cites Qwen Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.266859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.266859Z digest=sha256:829a1b1e5e560b7fb6ec345c196c776dacc583e9e83da41c85cbf26f29c282f8

Observation 6ed976d5-52fa-4f87-b350-6d6483148a3a · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.270569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.270569Z digest=sha256:c584e52548b69b37f60631d03b2d1e200468007497aed87706ebeb2fdc9ffd15

Observation 100f7f4a-5ab0-4897-aeac-f15ce15c3ddc · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.274367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.274367Z digest=sha256:954f3e663b2a1a972da74d0d33cd19043a80c4bd6ec4881dfc1842ba90b4d812

Observation 51982df7-5667-4d16-8a10-81e67ee8acb4 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.278391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.278391Z digest=sha256:d2aa89a1047d3a5759cc24b90c98c3768016a05322383c258c9941e5cc9fba62

Observation 5bc6395f-69e3-4fb0-b1e5-d9efa37ef0cf · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.281720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.281720Z digest=sha256:162b99fcc0b601b76e3da3ee28d0ed11ace0aeba31c07dbfe9a37af6ed6365c9

Observation 6164b773-31a9-42e1-a6bc-7d14aa21dccc · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.285282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.285282Z digest=sha256:9269ae325c849ad571b9a40b377b2e9895ec210d66359266e75c316275b2eb22

Observation e270aa12-acc4-49a4-8290-05c05ff6845a · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.288837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.288837Z digest=sha256:07eb587f7287bda7d76709a9ab9fd369afbde3d64d8b16dd2dd0c4f3ec089169

Observation b888b749-faa7-4be3-b08e-3024b0e967d0 · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.292670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.292670Z digest=sha256:fca8fb108757c0e37334bd43051fcc54e3a6c1af9bdf343782e19344bbe9a423

Observation 5f172363-a22b-4c8d-b29a-583864f922bb · outbound

This paper cites ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.296106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.296106Z digest=sha256:2f4d4ef19b589a5354e0075411f4afc254d4c7c107e45d2f6110202ebfd94816

Observation f7e6a82c-d1ad-4b30-9c0c-272f3de478b9 · outbound

This paper cites Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.299362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.299362Z digest=sha256:74b0338fef308dcfda50c0deacfa26c010a0992ca21c248206e1420e269e0bb4

Observation 608d509a-bded-40cb-8ee1-8cb0781c9a7d · outbound

This paper cites InternLM2 Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternLM2 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.302724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.302724Z digest=sha256:b8fee90e3ffc56c3ac2ba3dd6e79369280da42b6e824d8b955c6e55e5a29eb38

Observation 92b9e655-b4f4-4e0a-9a03-f79c0a52089d · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.305742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.305742Z digest=sha256:37a0c15d94d090c48d0b96cb3a6cdbfbda771c5e19498fa62291e148d4185827

Observation aaa4e496-0e8a-42c7-b286-2b822f85099c · outbound

This paper cites WACV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin WACV , year=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.308686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.308686Z digest=sha256:aa4c361f2636ab4879fbe170a80b4e423f083388efec11b30154b5833fb159db

Observation 5cb99119-9302-4632-ba03-ef4f6964d551 · outbound

This paper cites WACV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin WACV , year=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.312135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.312135Z digest=sha256:80504ef3dc9559e68fcbf6fd919c37fc22c4fcefecbc7a122cb7102d0b2c8d21

Observation 678ccdbd-6575-4c75-9c72-c3410921ded0 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.315460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.315460Z digest=sha256:54bca663aea6c2d6190da4c09047fe0805cfae1c4295b4c2aca72df2fd8670f0

Observation 9567cdd9-78a5-40e5-b775-3c853645e88b · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.319438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.319438Z digest=sha256:76352e5da80bec653af1113bf1b00b76f2237210716f10665cb21896dd2f4c7f

Observation 8060f9d7-21bd-49d3-b775-04767abd0b3a · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.322726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.322726Z digest=sha256:ce0a33e364317638cd4a0dd08ba915660ec3a9cc25b02e590a415b92ce532233

Observation c9bd4578-b93d-4a31-a85d-2ccab01855f4 · outbound

This paper cites EMNLP , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin EMNLP , year=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.326019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.326019Z digest=sha256:ee3130a176180c3afef957b38fdd19364f3db687f81d371408444076acc15380

Observation 27495724-e151-41ce-b800-bc8f69d17a57 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.329346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.329346Z digest=sha256:16ce7aec9785fa7ab0a0f4a4db54678a59ef02a609f659307cf2e7fc4cd36b30

Observation 3c74206c-3482-49aa-93da-85f0681ee83f · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.332660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.332660Z digest=sha256:d268ca8207571b8d02ad709f45bc6a7303814d04f376b8fdafe3a42bad62ddca

Observation 8b16f436-314e-4cb7-879a-1c7375f87e85 · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.336324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.336324Z digest=sha256:0b34d8d8661c02d9f93118973b85ec4da109168d26304e7df6602257f9dbe454

Observation c8e980ba-49ac-46a2-9ef9-9724a899f10d · outbound

This paper cites arXiv e-prints , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv e-prints , year=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.339687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.339687Z digest=sha256:50b68d7bca7dc1fb3a70c0e0818b3381394fff8cfee51d88baacfd116c0afe31

Observation cc095e69-dfcf-4ff8-96bd-0af283339d9e · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.343274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.343274Z digest=sha256:66617e65201e63ac87046b8c79c1e0eeb5c762a4c8c50c62277505c8661a1b1c

Observation d44975fa-5d17-40c0-a7a4-11d836af12bf · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.347023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.347023Z digest=sha256:0bc2216159f839a1e4c384e7ec881267e6751de6c0adf566d75ffd19072460f5

Observation 0a18acc9-376b-4343-b92c-c026baf70eed · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.350680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.350680Z digest=sha256:d2de714bae24c65101c2e53205a47a414ff1d0e2246bbb0cbc21516129dfe92b

Observation 20e53230-899e-417f-92db-a8dfd19d0703 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.354396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.354396Z digest=sha256:971c01d1afafd349c2a26ad35e5984b7758e76ef3ce71bca48e2fd86a5f3dfcd

Observation d406f4c0-480d-4dd0-8802-74a953a56716 · outbound

This paper cites ECCV , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , pages=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.357716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.357716Z digest=sha256:a3aca0a85a3048f308047f5bffbeb1cb0f93a9d15b0fde3cd44b75a5b9b171d4

Observation ef15e32c-ce71-4f94-aa36-064ccacaabc5 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.360982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.360982Z digest=sha256:20899362d8aea26eea48c146c1ef3a1f7591ddefbf5b9fd9fa1a5c1d29fd320f

Observation 18ef600a-d58c-4cbe-8cdb-1238e205f75b · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.364525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.364525Z digest=sha256:4054952468ba0b5a1cd9df19747b1a6f0f698ef5ca76b29d4cf218a69acc555d

Observation 5bbe1032-10e0-43e0-96f6-cb7127e9fa02 · outbound

This paper cites Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.367925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.367925Z digest=sha256:aa0955f6a5bd1ecace80a19b2e46fd58c11241e9962b111c45e62b60a0c21a1d

Observation 57db35f6-fe3e-49e7-a9fd-989fbf4a5cde · outbound

This paper cites TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T04:30:15.167292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.371579Z digest=sha256:63da5412a1aec3332838ff577431075f0e465ca51acc847806db1708ecb73c95

Observation c4bfe39b-061b-4146-86c4-abc0b20949bd · outbound

This paper cites FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.375574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.375574Z digest=sha256:f1c65082042124580dd650e4a20b133fa62b94475512a765dccb6a8ed5bd4489

Observation 72812fb3-2681-4c41-b044-677deb2ad573 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.379313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.379313Z digest=sha256:b658e5791c10f1c58792577df5d82269c0dc4ca7145872e9d06bedec7552e286

Observation b7c065a8-5ccc-4a65-a1a6-4ab6c37f6b8f · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.382832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.382832Z digest=sha256:59dc527553fe4b8f9d898ba59fad863d579a24f41caa8ee4c958a60a985d8f37

Observation 19a85498-d981-42e3-8a41-06ba36b4a12c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.386155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.386155Z digest=sha256:445f44e735f8cdc05374ef888c1b1982937f59395126f58db81c58a7a5b55866

Observation 868f1a82-2656-424b-b0a5-40c2d9dde5d3 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.856172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.390028Z digest=sha256:6eec15e6ebcb7dcf4c6a74a9486b03da04f464f7f6d46124af90c6b76371c75b

Observation 709f01b0-9b09-40de-bf1b-d9a8590fbc55 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.393401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.393401Z digest=sha256:dfd48fa8eb09770c48a5aa785f31abcb02a9ba54ef4869b638191e35841ab9f2

Observation 444a931e-74ea-4114-b2aa-e45f19f5938f · outbound

This paper cites ICML , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICML , pages=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.845287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.396351Z digest=sha256:4122f888c903ed92dd40a35a7ce4d5e28e4a1a72ccb4380491cc8494a94288f5

Observation d3efd223-49a3-4e6d-b047-506a4b0bddac · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.399358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.399358Z digest=sha256:6b27e4ca3a25cd71b20f1c36d13d959f9fac28303a930d024d58b0921532cc01

Observation 035b1fdb-dc96-40b6-aa2e-e4ddf1c81ea2 · outbound

This paper cites ACM Transactions on Database Systems (TODS) , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ACM Transactions on Database Systems (TODS) , pages=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.834253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.403709Z digest=sha256:9e91729b80dcc1fc0bae098f63d79a28522a102f76f9f2783e40f95ec6961ad6

Observation 6adf513c-42ee-4652-8ca4-90b192613326 · outbound

This paper cites GPT-4o System Card.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin GPT-4o System Card

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.406724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.406724Z digest=sha256:ee7d3742e43481f74c16b8a6e0991483ea2e9cb36493f5507e53e9b2c59fcde0

Observation b8db08b1-f786-44c5-981b-5861db570550 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.409858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.409858Z digest=sha256:3fee41e6b736af7ce37256667a02d1eb4d3e60ce0cdf3c1ef096d895628f31b9

Observation 2d548e17-b98a-459e-b572-53109317b777 · outbound

This paper cites Proceedings of the 32nd ACM International Conference on Multimedia , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Proceedings of the 32nd ACM International Conference on Multimedia , pages=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.414203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.414203Z digest=sha256:25bcf004e908182d204ab128e3b0530e1c932f5b6c4841e878adccf145f848ff

Observation 4717ca4d-4ce6-4747-b553-94e6f052838b · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.417640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.417640Z digest=sha256:3b24fc68a5f1674780d1b8ce3a6681eb3697ce3b3267f09792475d157b85490c

Observation 7d630164-4544-449e-8ca2-cd2eb7387854 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.421428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.421428Z digest=sha256:c215c02b2b6ff2865634a87c516764f8fac93ca76c0ea332572f717bf5ecb1e2

Observation 2bc0c6ba-b6c3-4fb4-a600-100a0fafb2da · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.424914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.424914Z digest=sha256:c35f109731a5e94372e845949f25d3ad9b3074a79c3e72634c27ccc45c59e9b1

Observation db8c182a-0f84-47e2-98b8-c2c81904cb62 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.800164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.428501Z digest=sha256:389f688107176128437549cb6a77280868a82075475498c832e000cf43992b78

Observation 457f16e9-e01a-42bd-a39e-d79d004b12fc · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.432345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.432345Z digest=sha256:53e21d9a2e59ec93da05c75a517004ae479ae2d034e8ae99579aeb78a41cdde0

Observation e148b1e9-b912-4325-bbd0-0e0adafa7149 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.436166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.436166Z digest=sha256:b7ce93f027efcd6c8ea8ade6cfa767c5f12217f8365d31c35f3d011693f01e97

Observation 4f4ca747-1fac-483b-b8cc-40afb9b37933 · outbound

This paper cites ICCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICCV , year=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.788679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.440011Z digest=sha256:ade89da4b28ce194e1b6f2ee85ec926d841c46009d593ea2badff60094d91c74

Observation 289f2c93-275c-42ca-99dd-849180301cb3 · outbound

This paper cites Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.443529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.443529Z digest=sha256:9caca3ba637350ac9fc4c223269cd28e4c64d76deca4510474d0f13f00b3d441

Observation e7414c56-1099-4ff0-8fe4-97423e1612eb · outbound

This paper cites Qwen3-VL Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen3-VL Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.447236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.447236Z digest=sha256:dd38a7215eb344766aa0207e7d5a841fd59b5f92dd687ccbbcad7c0766d589fd

Observation 34ee7d7b-0e4d-426e-92b7-98060254abd8 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-OneVision: Easy Visual Task Transfer

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.450759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.450759Z digest=sha256:00d99faea128be66190e4f0aeeba4902218d3be0ae1ce6f9acbf16cbfa2a9af3

Observation 04c5c2e7-9d73-4d01-9880-688927c997e7 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.775868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.454179Z digest=sha256:865d1b249932f678f80d887769e29a8bf0a407e784a9191c9e45f80224f7fd37

Observation b7787d07-583c-4a18-81f7-f54d3e940750 · outbound

This paper cites AAAI , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin AAAI , year=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.765249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.457768Z digest=sha256:1783e4e73e19a162e5125d434699d1e516bb18ca54233757472cf10573de70ad

Observation 3ec38fa5-dacd-434a-b02d-69faf847e68e · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.754472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.461441Z digest=sha256:41763ec8de603399f2afb0f2421af42d6396b465c19f78fed123b7614fa2406f

Observation 40a305b0-f104-48b5-866b-d91e4a27bc00 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.743459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.464845Z digest=sha256:b7a74cdf20628768c392d4a8894f361b5d41265f3ee588388a6e91c5ffea7b87

Observation f66abced-ddcb-4954-8520-3e15de978fef · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.731968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.468161Z digest=sha256:4eeea5a39937e2d28d4032074a70ded37614fa6dd85a22c74ab32570db39bead

Observation caa523ce-41c7-476e-a371-1ba415be5e77 · outbound

This paper cites Seed1.5-VL Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Seed1.5-VL Technical Report

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.471570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.471570Z digest=sha256:c1c59719dfa5457e0277ef683d0c52071bae99fa3a2a3736da7f5e33b5b2e26b

Observation 3e278458-c443-4484-aac3-85b23c675074 · outbound

This paper cites InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.475211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.475211Z digest=sha256:9fea6b85e4358411f3f8228efe68697e7420e0a2b57f910b5b60ead49aa07a22

Observation f3da428c-2a7a-49f5-ad92-2ef1b917e3e8 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.478903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.478903Z digest=sha256:031c32e4039dbadd90fd7f4f78d6d1af3d5302c7da00d27217dedf96b1278f6b

Observation 33aaf4d2-def7-4a77-8995-12df540f570b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaMA: Open and Efficient Foundation Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.482592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.482592Z digest=sha256:7457ac30e4203a537a837a64f1f785b50fe3796b881de39712a573e958e73813

Observation 019cdcc6-3c79-4eb6-aa13-02a89bf8fe65 · outbound

This paper cites DeepSeek-V3 Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin DeepSeek-V3 Technical Report

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.486359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.486359Z digest=sha256:479570f2d3702224bc3cd72bbee7a326701019c69c72909e59f77ea407edc5d0

Observation a9fe5aea-29c6-4dcd-8398-07cc5861b1cb · outbound

This paper cites Qwen3 Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen3 Technical Report

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.490105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.490105Z digest=sha256:e8457c67c3a64d802c0e803f70a717237d2fb6a6465e5d3fc836e4f265d90576

Observation 4d21e384-dc0c-4128-b777-d704c90f14bb · outbound

This paper cites AAAI , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin AAAI , year=

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.712087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.493811Z digest=sha256:897f0de5048f9ba0a480fecccdc45cac2cbc45899ac11bf16ae8451ff7132de1

Observation 044e74ac-68f1-45b6-b553-57d2505b46b7 · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.497187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.497187Z digest=sha256:c2cb72b6f84c98a1e51dd97e11a0f8e7c67658d2a51ab3a96b6ca9f77626b910

Observation f2c659c3-613d-410a-a1eb-09ec16f00ad9 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.500265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.500265Z digest=sha256:3fee2fb231b66cabc7da35c9c7593711c2205981623879b4aebad5f8e42b2c1d

Observation c403b17c-df75-49e6-a8f4-2bb4ad3f6767 · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.503529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.503529Z digest=sha256:d000657b36fac27a28f1ad213a7e0bcce35536834b50cc56cbf95cd44cea97a9

Observation 8ec82576-f1d4-4563-a1ce-99ee577f7eab · outbound

This paper cites SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.506806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.506806Z digest=sha256:87672ded08102abac7397cc088436bb94c321e09817f34e473641bfb3c97b5e1

Observation 204158c3-e71a-4351-8031-9963114faf95 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.510251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.510251Z digest=sha256:c76b7a6fe64b57bc6b3b58d2296ce39b47b83affa444a36fd61547cc76809606

Observation 41a5184a-a68f-454e-a9cd-4422249b0d01 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.686593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.513961Z digest=sha256:ca1d3fe7f97a2c77f27aa0edf424e443aac2261b98a6370732ad8b56c648b7a4

Observation 09aec212-b394-4a0a-ab5a-ec21b1b564b9 · outbound

This paper cites OneThinker: All-in-one Reasoning Model for Image and Video.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin OneThinker: All-in-one Reasoning Model for Image and Video

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.517712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.517712Z digest=sha256:1a682eb1c9c53f56ba95e11716fa848e89dfd30d5bdfa6823a1066a3ffc15cdc

Observation b8dbd646-f81e-45ae-9aed-dae7da1668b8 · outbound

This paper cites LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.521293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.521293Z digest=sha256:d023a6933de4ff5d912048f297ae6435cfedbdf6efd2b965a6b537d33e23d49f

Observation f5d1790a-cddf-444e-b41d-4f6533f24bba · outbound

This paper cites arXiv preprint arXiv:2601.22674 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2601.22674 , year=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.525086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.525086Z digest=sha256:bd7ab7ad353ecab3feb636f7491dcbbe8b256a93ef35e467d1b5c13d5b930b61

Observation e7b0d891-b32c-481b-b396-e3e10cc6872e · outbound

This paper cites AAAI , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin AAAI , pages=

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.675107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.528582Z digest=sha256:d87e2817eac2ce351a6da9e08897b74a8d7fcaa5e1649789f544123363dbbc5a

Observation 13f1b919-3589-4789-9d9a-62e60ffdf5f7 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Findings of the Association for Computational Linguistics: NAACL 2025 , year=

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.662753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.531924Z digest=sha256:2f94bdeb9a3512c0e115e55d3d7585b538e104f6ea0952a7d3450e69e8ce6c9f

Observation f250d59f-7c12-45d5-ac2b-c74a27ad96b2 · outbound

This paper cites arXiv preprint arXiv:2602.23699 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2602.23699 , year=

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.535753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.535753Z digest=sha256:2f8c0bc5d0102c77f234c8f926a7a85b1902c7e9655adc7571d9d39cd28fe5a0

Pith citing papers

No inbound Pith citation observations are available.