Pith. sign in

Paper Citation Record · LEDGER

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin

As of 10 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 0 inbound Pith citation observations for arXiv:2608.06411.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06411 v1

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:30:14.535753Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 105 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved85
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3fefcc0-c1e5-49a9-8b02-7fb1be112d50 · outbound

This paper cites Qwen2.5-VL Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.184628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.184628Z digest=sha256:60807ee94fce6a9949bd636429df914e80123b45deded106611e1bb4e20a4a82

Observation 5bf4cec4-7fbf-4622-8715-6616263e904a · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.189388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.189388Z digest=sha256:d13258617b46149b8b7fd50ab183c05718907a68eff79d2b4af887bf607c8f65

Observation 6613c666-b789-4a1d-ab34-393cc274f899 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.193471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.193471Z digest=sha256:2446c6cadf78c3f6574aba7a6af4a7700d668d1a204e662ba4a50806078fd298

Observation c552ff0a-8265-45dc-a260-ef63ec3ca58b · outbound

This paper cites ICML , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICML , year=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.197067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.197067Z digest=sha256:328df8c356d38a85634df03a19689342a6ee1e0cc8bb6ddcbf782354ce8369e6

Observation 9216dc4a-716e-4300-a6a9-785c5b82ebb1 · outbound

This paper cites ICML , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICML , year=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.200763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.200763Z digest=sha256:80ad1897f3c330a68f1d917859f666de51b58b0ba09a2f53d36c41d3ada4c197

Observation 7f200349-2f41-498c-bbb6-00669a2110fe · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.204559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.204559Z digest=sha256:7a51436d087ac1d8a57c446c9398a4f21154ff90b20e1947d27244c9b79b3e23

Observation 655ff2fc-a489-418d-8f29-ccdc15ad5dab · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.208606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.208606Z digest=sha256:afbbd2274404f5ec0895c4e523885ff6ae555e56ef5581f23f4dacfaaebad3a2

Observation 788c189a-8fd3-4e98-b920-0f4319ec14bc · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.211519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.211519Z digest=sha256:0c927bbaede19a767c02a544f3d434cfa730f33e24f09dc35b49475aaef26d40

Observation 36222e3f-701e-4743-b392-5722206979a5 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.214635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.214635Z digest=sha256:f46740835676e9dc2a72071a6c3c35435a0fd11700ed8f9677402493ed16d3f5

Observation caa82ddf-0c37-421f-a9fc-9152b55e2ff1 · outbound

This paper cites an unresolved cited work.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.217581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.217581Z digest=sha256:1e560a6f264b2b647b33acf2e22797e2807cce8d5076fd8628363597d4c4d0fe

Observation 456cabf5-7ad3-4338-a54b-94098b316d52 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.220359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.220359Z digest=sha256:0a84d8ac716815de425783271ec4a8eebddb8a381f1b41043bd313da095b14fb

Observation 4a5433c3-b826-4b3f-8dd8-2695fc491d56 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.223732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.223732Z digest=sha256:a32242d4849ff79c637c3ac0f350504d98eb6d612accb3773a67420d46d9bf8e

Observation e3ed016b-e8d6-4909-8094-f1906e58ba7d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Gemini: A Family of Highly Capable Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.227340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.227340Z digest=sha256:3753a719af4f2113f90bce81090d22a1bc71b636e00cc804302ebb8b93110a2c

Observation 498c2365-7b92-4d8d-a347-8476b14d3987 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.230915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.230915Z digest=sha256:347bb4e2b7a4a312a37a3e7f499edd65804189d3e32960692d4a2d6b6e39dddf

Observation 83c95443-0f89-4b8f-b099-9c4efeda70fe · outbound

This paper cites an unresolved cited work.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.234665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.234665Z digest=sha256:469d48edb7be893479c15015b8b00addc3afdefd8ae50da10d607cb73bdbe33e

Observation 04fcfba3-a3aa-4fc8-be22-c34dbc4866f3 · outbound

This paper cites Token Merging: Your ViT But Faster.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Token Merging: Your ViT But Faster

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.237915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.237915Z digest=sha256:b2637ff8d3516ffb5341e017e70ccd4dec61564a1b2f67383ae65de16daeb465

Observation daddbc57-a323-435e-9fb5-7d783b1be326 · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.241516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.241516Z digest=sha256:fdfe7e9c806b52efc2e0df6351a24bca262842615b0a8a070445eb1a28340a86

Observation d0f3ba0b-a1e1-49bf-ac40-98a4e7670049 · outbound

This paper cites arXiv preprint arXiv:2403.15388 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2403.15388 , year=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.244923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.244923Z digest=sha256:7ddda661e82cfb492e89fb634e27b26b91e0335f0ed0c0f626d18fe8aee1f08a

Observation 4cd2cbfe-3851-4f2f-9a68-ff4bc8485b30 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.248316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.248316Z digest=sha256:27b5123d01c9a90031b45dc6c125a06ff11c7b2d4778b2737cba5673376859f4

Observation c2530b53-5a83-4dd9-9946-b4b97b816b33 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.251891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.251891Z digest=sha256:fdde257f24b318d4f4312092186fa6f8d107f6bac59d136194236dea80d929c2

Observation 3c74d0f9-56b4-442e-91c8-51e2e8a3ec73 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.255625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.255625Z digest=sha256:54f3ff34131199f72e2103818d86763f4592a079831474102bd806ded5d4eab4

Observation 5c9d12e3-3039-410b-b995-177db0a4e534 · outbound

This paper cites arXiv preprint arXiv:2411.17686 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2411.17686 , year=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.259308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.259308Z digest=sha256:ec2ea6cc89c02fa2fc743f23e8d4c73f173e7c15459bbb9d2416a73bd2a7c015

Observation b05b1e3c-ad65-4f76-a39e-be00825a9b60 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.263104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.263104Z digest=sha256:37518ddc9e7e354129f51acc4ad30b9cc0438240f2a2a785ba16a3ef36a5faab

Observation 73fd6f05-f54a-4edf-a71b-92ff8ec78783 · outbound

This paper cites Qwen Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.266859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.266859Z digest=sha256:cb5b1249dc10472b3ab330ad34c5a0cd967f1d03bfbee0c0d66f5ead5cf7b0ec

Observation 6ed976d5-52fa-4f87-b350-6d6483148a3a · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.270569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.270569Z digest=sha256:fa4ff70b3c105bd3bfdd53d399e946e81683ba9c71a58cb5795aea64d1cc71a6

Observation 100f7f4a-5ab0-4897-aeac-f15ce15c3ddc · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.274367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.274367Z digest=sha256:713c0ec5f6ff48cbe21282abd5fc06cd791678f2fc4fc05a6c919aca2f79beb2

Observation 51982df7-5667-4d16-8a10-81e67ee8acb4 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.278391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.278391Z digest=sha256:0c3a6023e44e253402f00b1f5a8f354f5ad423fd8c18c254d336890ad630b7d2

Observation 5bc6395f-69e3-4fb0-b1e5-d9efa37ef0cf · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.281720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.281720Z digest=sha256:33926d6c11f7d38c706c6464627234fb13ffb52b87fc26c4f8c002c197f2e8b2

Observation 6164b773-31a9-42e1-a6bc-7d14aa21dccc · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.285282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.285282Z digest=sha256:e4dc8b3de2a28dd392abeb6d7adc0c80523ad070b50a49ab6497e104a033a456

Observation e270aa12-acc4-49a4-8290-05c05ff6845a · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.288837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.288837Z digest=sha256:485ec4fb6f20e364dbb5c07a62f6f0919fd4368235289ce36240afa9fc67b1b5

Observation b888b749-faa7-4be3-b08e-3024b0e967d0 · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.292670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.292670Z digest=sha256:7003960f2959e68305c3b4510def528fd9a2c214396bef4a0c1afb3ba33b959e

Observation 5f172363-a22b-4c8d-b29a-583864f922bb · outbound

This paper cites ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.296106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.296106Z digest=sha256:0ec5a787f1580360d5d3db11cd0f10152b37cb021ef7806dcf3517d7252f9d57

Observation f7e6a82c-d1ad-4b30-9c0c-272f3de478b9 · outbound

This paper cites Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.299362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.299362Z digest=sha256:99be45194d699fa80b443d9782106892f10e5d26b201b17d7921825b93fb8b0d

Observation 608d509a-bded-40cb-8ee1-8cb0781c9a7d · outbound

This paper cites InternLM2 Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternLM2 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.302724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.302724Z digest=sha256:3a5347bf4efabd889d7fc44c74429737af14823e691fd449758769639d68dd3d

Observation 92b9e655-b4f4-4e0a-9a03-f79c0a52089d · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.305742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.305742Z digest=sha256:f84579e7dbe5ae9575ff3a09bc224475be29e075407c6c7575ace6f11a32c6ed

Observation aaa4e496-0e8a-42c7-b286-2b822f85099c · outbound

This paper cites WACV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin WACV , year=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.308686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.308686Z digest=sha256:a0f7685d4b315f8b2ca045505693cf52ed8c82bc1bf6592ad8101d6c09510e59

Observation 5cb99119-9302-4632-ba03-ef4f6964d551 · outbound

This paper cites WACV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin WACV , year=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.312135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.312135Z digest=sha256:fdc4ceedbcbfb1495ce3ff66b827812f12da0d9152313d1c3f0aae177356ce09

Observation 678ccdbd-6575-4c75-9c72-c3410921ded0 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.315460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.315460Z digest=sha256:fa35adc1fb5473bc799d18feee3faad2d4ab78986f01717f66a0cc690fe77b52

Observation 9567cdd9-78a5-40e5-b775-3c853645e88b · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.319438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.319438Z digest=sha256:078f132636b9e3fcf1b50d738abc3cf98c324effeb815e837beae48aab48826e

Observation 8060f9d7-21bd-49d3-b775-04767abd0b3a · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.322726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.322726Z digest=sha256:65e00c682b9b94d2384735195decd7adfd9fb12e4141141533f9994090ee0dbb

Observation c9bd4578-b93d-4a31-a85d-2ccab01855f4 · outbound

This paper cites EMNLP , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin EMNLP , year=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.326019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.326019Z digest=sha256:02288bd9cc47b0e15739e9f52d0bfc8d9aa1ad223aa02976143c8f063ec68fb4

Observation 27495724-e151-41ce-b800-bc8f69d17a57 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.329346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.329346Z digest=sha256:38b6e6e5f90b90494345e66ef9a009ce3bfbf2a88b0c8a55d14c2eb87b7ba421

Observation 3c74206c-3482-49aa-93da-85f0681ee83f · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.332660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.332660Z digest=sha256:242222722152c07e2bf417e97a08d2825dcfd0315151e2bacca70ac9b3217243

Observation 8b16f436-314e-4cb7-879a-1c7375f87e85 · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.336324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.336324Z digest=sha256:29b29b7df6577561b17c53eba84d579b2250f8f254e3f07971dbe88353114da2

Observation c8e980ba-49ac-46a2-9ef9-9724a899f10d · outbound

This paper cites arXiv e-prints , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv e-prints , year=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.339687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.339687Z digest=sha256:fffd422a78794beb949f58ca2f40c763d70245d0a6a25c21369358da62e0d73a

Observation cc095e69-dfcf-4ff8-96bd-0af283339d9e · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.343274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.343274Z digest=sha256:9495c8a952a3f4f3efdd70ac4d43c3538a67b24ad050afbc0b122ad86bd8c963

Observation d44975fa-5d17-40c0-a7a4-11d836af12bf · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.347023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.347023Z digest=sha256:8e83e0a6a24eb388b63a28a94adb917626657a39b5cd5ab17f8eba01a7ac060b

Observation 0a18acc9-376b-4343-b92c-c026baf70eed · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.350680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.350680Z digest=sha256:76b8ba7d33034d6ac68c798b851e433e1bbf6cd68f4b81318dc6e17e5d2b674b

Observation 20e53230-899e-417f-92db-a8dfd19d0703 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.354396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.354396Z digest=sha256:bfa6c5111fa36c16f7a6cc9edc9b4a652b0bbe42da2145941eb99273813c7cff

Observation d406f4c0-480d-4dd0-8802-74a953a56716 · outbound

This paper cites ECCV , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , pages=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.357716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.357716Z digest=sha256:078933fbccedcfc4053c509d88dcc22b8f4d4fc71b7309cf60c4f90ecabaac7d

Observation ef15e32c-ce71-4f94-aa36-064ccacaabc5 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.360982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.360982Z digest=sha256:56c3fc4d7fe8b0e4b94f4339e3f259f05f4695ccafd4eca73e08205e23407621

Observation 18ef600a-d58c-4cbe-8cdb-1238e205f75b · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.364525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.364525Z digest=sha256:5d95b9ed5ba4ca958ab2c778000c79fa7b512be41c3f565206d48faad355f7dd

Observation 5bbe1032-10e0-43e0-96f6-cb7127e9fa02 · outbound

This paper cites Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.367925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.367925Z digest=sha256:f48723f95081461815df2b3bbb0945fa7d0b0e374c6edf093aaecddbb0a0599c

Observation 57db35f6-fe3e-49e7-a9fd-989fbf4a5cde · outbound

This paper cites TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T04:30:15.167292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.371579Z digest=sha256:95df3b7d96aa81cfb96d67635073dcf330303a78bd4a1af97e1c63e91912e8ba

Observation c4bfe39b-061b-4146-86c4-abc0b20949bd · outbound

This paper cites FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.375574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.375574Z digest=sha256:6da0e4146ce335ee4349333ac9f476f89981bbc7e1cc7eb2c7fea3560120c1af

Observation 72812fb3-2681-4c41-b044-677deb2ad573 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.379313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.379313Z digest=sha256:f60dee7ebac4d55a9ff86ccddfd688c039f1417a98c7bcaacd610a72ab8771a8

Observation b7c065a8-5ccc-4a65-a1a6-4ab6c37f6b8f · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.382832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.382832Z digest=sha256:4fb2d98a46259546eb997465f60b2709824deab6ffed5057fabfa3c80cbfbd1d

Observation 19a85498-d981-42e3-8a41-06ba36b4a12c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.386155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.386155Z digest=sha256:e10be5dadad9f49942783a778e90e3b3d06f3fe435a0618328e9a9e46b468e24

Observation 868f1a82-2656-424b-b0a5-40c2d9dde5d3 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.856172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.390028Z digest=sha256:c691467ca5690cfad98635d71b8ec60c251e4afdd766d32fb5b381c961a40682

Observation 709f01b0-9b09-40de-bf1b-d9a8590fbc55 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.393401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.393401Z digest=sha256:4bbaed4e3980ca66e113241cbed7bcb6d3f360ac946e8aa058c181934a3c80a8

Observation 444a931e-74ea-4114-b2aa-e45f19f5938f · outbound

This paper cites ICML , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICML , pages=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.845287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.396351Z digest=sha256:4f83d48e5a0029f5b057db87e9c670b8153d05faf7aabdf8c4b1972fb97d2363

Observation d3efd223-49a3-4e6d-b047-506a4b0bddac · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.399358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.399358Z digest=sha256:a2003ca7f3da2bff6045f41db3b0f15bf72bfefcf5e010edd732cb8f9cf1ecdd

Observation 035b1fdb-dc96-40b6-aa2e-e4ddf1c81ea2 · outbound

This paper cites ACM Transactions on Database Systems (TODS) , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ACM Transactions on Database Systems (TODS) , pages=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.834253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.403709Z digest=sha256:e74418286ae55e5184ae11c2992ea72832765de56d52ca2c21d5527417de9644

Observation 6adf513c-42ee-4652-8ca4-90b192613326 · outbound

This paper cites GPT-4o System Card.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin GPT-4o System Card

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.406724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.406724Z digest=sha256:1680de9cd718c24f0dc268e7fceda2e723c6376ad0352234e5908cfccb4b496b

Observation b8db08b1-f786-44c5-981b-5861db570550 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.409858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.409858Z digest=sha256:f7594ed5c182faeefbad5a577a68fb2e9c2cae90a44e2697eca8e8a0390578a6

Observation 2d548e17-b98a-459e-b572-53109317b777 · outbound

This paper cites Proceedings of the 32nd ACM International Conference on Multimedia , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Proceedings of the 32nd ACM International Conference on Multimedia , pages=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.414203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.414203Z digest=sha256:5a1fd79a6ad0b2af341af0d1f882a590b419bf06ad418eed7b3d1fc254be7ef1

Observation 4717ca4d-4ce6-4747-b553-94e6f052838b · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.417640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.417640Z digest=sha256:1d407afe3552ca945a0acc0c9d2067afa4ad66abcd56baa9981a50bc488c3b0c

Observation 7d630164-4544-449e-8ca2-cd2eb7387854 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.421428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.421428Z digest=sha256:cf6a02f34aa19d54043a2a508c1c5292a470a70b1d3493ff29b0db61c5e0619c

Observation 2bc0c6ba-b6c3-4fb4-a600-100a0fafb2da · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.424914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.424914Z digest=sha256:046940560fa412ffff9a9ba36e63146c283f34f3a6b2ae0e858e44de01accbe9

Observation db8c182a-0f84-47e2-98b8-c2c81904cb62 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.800164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.428501Z digest=sha256:7875a0e69358afab29494ae12454c6ebd06af3ceadf0b23afe38895ab86b319a

Observation 457f16e9-e01a-42bd-a39e-d79d004b12fc · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.432345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.432345Z digest=sha256:6b35a5db154af87542e0930dd840629b990e59d303375859cf0284fac1d990ca

Observation e148b1e9-b912-4325-bbd0-0e0adafa7149 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.436166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.436166Z digest=sha256:fb6af3abf9c44fa2b627aad89134478b8be119a5604ca6adfda6270d600f2e88

Observation 4f4ca747-1fac-483b-b8cc-40afb9b37933 · outbound

This paper cites ICCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICCV , year=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.788679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.440011Z digest=sha256:2f088d0276b5d45a7eaa87408dd7aea442f96ef2150ebee20eeda746241f8a0d

Observation 289f2c93-275c-42ca-99dd-849180301cb3 · outbound

This paper cites Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.443529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.443529Z digest=sha256:f7fe2701b393bcfc98fc0ec1ce49e16ee8ba790579a7d421f609c577b94abad8

Observation e7414c56-1099-4ff0-8fe4-97423e1612eb · outbound

This paper cites Qwen3-VL Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen3-VL Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.447236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.447236Z digest=sha256:41c57b1c599cbd04fb61b38bc1073d3843b3d16a4ba227cf92b45a1f88d099b3

Observation 34ee7d7b-0e4d-426e-92b7-98060254abd8 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-OneVision: Easy Visual Task Transfer

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.450759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.450759Z digest=sha256:3e84dc48d9f39a97f6eef06bc3a9f1c4ca80766907588eca567acf4903355137

Observation 04c5c2e7-9d73-4d01-9880-688927c997e7 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.775868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.454179Z digest=sha256:b676a1ffc3620c3a5fdb32a19946034ce5738de7b1249fe824747ec8152f3ba9

Observation b7787d07-583c-4a18-81f7-f54d3e940750 · outbound

This paper cites AAAI , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin AAAI , year=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.765249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.457768Z digest=sha256:20d5803dc1d380805f12663811780d516b730ae76fd5579422e50d771e1a477f

Observation 3ec38fa5-dacd-434a-b02d-69faf847e68e · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.754472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.461441Z digest=sha256:6bd6013638a525f383c5669af9b2ea4df3891a1fb0585565d8df0cdf8c16ad78

Observation 40a305b0-f104-48b5-866b-d91e4a27bc00 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.743459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.464845Z digest=sha256:93b60834ca7d0c8456e6f16f10752851e61aea7ad2381596348955e2576b4bba

Observation f66abced-ddcb-4954-8520-3e15de978fef · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.731968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.468161Z digest=sha256:50f7fde7c430a0bcd5040d022106acc2694a659b9bb2b10641e1164296df0ef8

Observation caa523ce-41c7-476e-a371-1ba415be5e77 · outbound

This paper cites Seed1.5-VL Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Seed1.5-VL Technical Report

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.471570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.471570Z digest=sha256:88819f9b54ae7a5fab35e66c39b907bcf1df50f38443cefc7a6bcd956757f639

Observation 3e278458-c443-4484-aac3-85b23c675074 · outbound

This paper cites InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.475211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.475211Z digest=sha256:90ddeb94b6e8ecfbadba31743afb9585952844e929c312271f80db181cf622c2

Observation f3da428c-2a7a-49f5-ad92-2ef1b917e3e8 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.478903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.478903Z digest=sha256:bb66ce8419f6a2a328eb745783861e8ffa5f8482de6b61248ae8875da2074c1e

Observation 33aaf4d2-def7-4a77-8995-12df540f570b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaMA: Open and Efficient Foundation Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.482592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.482592Z digest=sha256:f505994e48dc4a980512963c2394a912851da88c1e060b8c1ebbec9f3f715541

Observation 019cdcc6-3c79-4eb6-aa13-02a89bf8fe65 · outbound

This paper cites DeepSeek-V3 Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin DeepSeek-V3 Technical Report

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.486359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.486359Z digest=sha256:6126b7d33745cfeab8f67c12d7096e00a7dd830591553009622b9961cee9d0a0

Observation a9fe5aea-29c6-4dcd-8398-07cc5861b1cb · outbound

This paper cites Qwen3 Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen3 Technical Report

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.490105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.490105Z digest=sha256:35eb640a6889dcc9a3dde2c38c5fec66b3c2cbd0ea3b7f9f9911f1baf1123465

Observation 4d21e384-dc0c-4128-b777-d704c90f14bb · outbound

This paper cites AAAI , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin AAAI , year=

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.712087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.493811Z digest=sha256:807f65b732810fa2eb6abcd9cfdfc01aab7ed953b3e1c7632509c41a4e4873b4

Observation 044e74ac-68f1-45b6-b553-57d2505b46b7 · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.497187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.497187Z digest=sha256:c63e60e2b2600a1bbb3727d113a5fe1c27ff306d141c6223964a3e8b3078f8f7

Observation f2c659c3-613d-410a-a1eb-09ec16f00ad9 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.500265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.500265Z digest=sha256:c1727974e22d5ef5639a85760892e8cdc68bcfb7db0fe3033096cee2a17bbeb9

Observation c403b17c-df75-49e6-a8f4-2bb4ad3f6767 · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.503529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.503529Z digest=sha256:a88da83973ac35fea411453a9635eb6c751c2d90767c92ac6a419b740a1d53bc

Observation 8ec82576-f1d4-4563-a1ce-99ee577f7eab · outbound

This paper cites SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.506806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.506806Z digest=sha256:ea18becf1494a8c090d31a586909e372af7b5cd20540dac061ea4fd4e93e56fd

Observation 204158c3-e71a-4351-8031-9963114faf95 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.510251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.510251Z digest=sha256:d5e3f16274538175e87ac84e8a9b1d792b5e4924111087bb3ceadb98c62c4232

Observation 41a5184a-a68f-454e-a9cd-4422249b0d01 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.686593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.513961Z digest=sha256:0bc52ea289e344c455f5b10c0bcabe1502f257c913d27a7b7f810174292ad754

Observation 09aec212-b394-4a0a-ab5a-ec21b1b564b9 · outbound

This paper cites OneThinker: All-in-one Reasoning Model for Image and Video.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin OneThinker: All-in-one Reasoning Model for Image and Video

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.517712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.517712Z digest=sha256:ed425f73d1fe7e57246bd2f0e595cdd9491ea34aa30e32d837f3bdacdfeec264

Observation b8dbd646-f81e-45ae-9aed-dae7da1668b8 · outbound

This paper cites LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.521293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.521293Z digest=sha256:607c9cf8bc0a91b0751939668e5a8280c416848f26b7ce753f03e124c4f72da7

Observation f5d1790a-cddf-444e-b41d-4f6533f24bba · outbound

This paper cites arXiv preprint arXiv:2601.22674 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2601.22674 , year=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.525086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.525086Z digest=sha256:7ede1279b017d546aedbd0a843cb7286d9b364bab098ac3a2b47172f5fea06cf

Observation e7b0d891-b32c-481b-b396-e3e10cc6872e · outbound

This paper cites AAAI , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin AAAI , pages=

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.675107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.528582Z digest=sha256:6927356201d5feb69ff6f5599e31f3cd240f2d65d3cdb1e8d4fd76dad8ef3304

Observation 13f1b919-3589-4789-9d9a-62e60ffdf5f7 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Findings of the Association for Computational Linguistics: NAACL 2025 , year=

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.662753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.531924Z digest=sha256:62537554cbed2a2b663d79cdfa89595f833ccb164e0927ac27837f1e6731f01c

Observation f250d59f-7c12-45d5-ac2b-c74a27ad96b2 · outbound

This paper cites arXiv preprint arXiv:2602.23699 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2602.23699 , year=

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.535753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.535753Z digest=sha256:5e5b303ee13478ca8b4daed860e540e52a502d586d4969a0408b17ce164d82a0

Pith citing papers

No inbound Pith citation observations are available.