Pith. sign in

Paper Citation Record · LEDGER

Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2405.05803.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.05803 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:28:37.284569Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3b2d0748-67b0-40b2-8e29-7a8d8f00469a · inbound

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression cites this paper.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.284569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.284569Z digest=sha256:1854138f4cf5547ca07049a168f60886343afabf1e5a26c3c40cc61078df7b9e

Observation a26d361b-458d-46ad-8683-25451e33e88f · inbound

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings cites this paper.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.937545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.937545Z digest=sha256:85923f8e4437d9311673e87a7d0133fa9f95e67f60208e84e8d779ba389488f4

Observation d0092223-0a8c-475a-9347-a225787ceb6a · inbound

Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction cites this paper.

Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T05:21:10.859821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:21:10.859821Z digest=sha256:0830caf08e08c543e1a9a8635d176d6a66b8a36e3c983a7ac557e8fdf7cf9675

Observation e31ec82e-f815-4370-bf47-b7610183a5ff · inbound

MBQ: Modality-Balanced Quantization for Large Vision-Language Models cites this paper.

MBQ: Modality-Balanced Quantization for Large Vision-Language Models Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:52.233673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:52.233673Z digest=sha256:c297d3c6a1bc04e26c35c22250c620865c76514f79dd705f2fb19e3cd765199a

Observation 1257f398-6e4f-412b-9219-cf04de81e5a9 · inbound

ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming cites this paper.

ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:39:18.114263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:39:18.114263Z digest=sha256:42d872ef49aa9ba08820badf07cc5960c46a82697087adbe82b91aefe89717c7

Observation 67ef1458-76b2-41f6-87ca-0c0b71ef2bc8 · inbound

What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph cites this paper.

What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:20:45.731407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:20:45.731407Z digest=sha256:7dcc1e8b689657037894bc96daa1d3ba9aab5504f8ad4d41b76c75a340200b02

Observation 533e2edb-6634-435f-a09b-2c7f9bc28248 · inbound

HiMix: Reducing Computational Complexity in Large Vision-Language Models cites this paper.

HiMix: Reducing Computational Complexity in Large Vision-Language Models Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T19:18:49.288803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:18:49.288803Z digest=sha256:ded2b99a1b175d594652aab3fa45b43f0b8c5d5847e46e94c9b3c8913fc9206c

Observation 8904b5bd-f1e7-4be9-8402-952d135fd0e0 · inbound

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs cites this paper.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.727495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.727495Z digest=sha256:ce5bdbb565af0eb71b72c3d126d9b614d660288dd117a9d7a14d2a8988d5e0cc

Observation b00a0740-01fd-48c8-a130-bbcfc683fd6c · inbound

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models cites this paper.

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:53.659544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:53.659544Z digest=sha256:ca298b99a40658938bb5021478c7bb6ff5afca0e6b0fb90648c250b29ad69900

Observation a8c4bba1-5541-4156-9d32-07021750a33a · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:28.571315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:28.571315Z digest=sha256:3e3a522e1c4ce73aede660479928fcb4b57b28b1da4fb3c19f2b11909c9bffc3

Observation 8b21d098-658d-4f82-8300-bd55089c689f · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.514560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.514560Z digest=sha256:473ecd3dd19dfaa3b5bf977f60e35eb4de0d153b6b3a3661fc1fbf0bb8018a20

Observation fdc7355a-a878-48c0-a559-eb2e5a9c13c0 · inbound

Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers cites this paper.

Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T10:55:18.205488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:55:18.205488Z digest=sha256:9f4651192670ffc6923eaa74b80831c0e66fdcfb0b9f342faa077ea0f655133f

Observation 8d4a8035-c364-4586-ba74-5874fe2d81fa · inbound

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models cites this paper.

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T05:51:15.965046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:51:15.965046Z digest=sha256:1d6ab42df3aa2909253873af43149fff4f488a0b3624dc42803a898a539bddb5

Observation 65e85f8d-6a2d-495b-ae7e-823283bcc887 · inbound

IKOD: Mitigating Visual Attention Degradation in Large Vision-Language Models cites this paper.

IKOD: Mitigating Visual Attention Degradation in Large Vision-Language Models Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T04:29:52.320117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:29:52.320117Z digest=sha256:93946d618e3752e9f8cddc9f9a9cc3ddd272b4db590f72ac9f7e388a503d8b9d

Observation be74d94d-1371-41f4-9704-d21c8c7b1298 · inbound

Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturation cites this paper.

Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturation Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-27T16:41:03.330850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T16:32:09.784716Z digest=sha256:1f0d026c556026b3a0d8be64a26be7199f977f18d6eecf98660386cd38a40a27

Observation d8e5c122-354f-4bf7-851f-23d6203502e0 · inbound

SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models cites this paper.

SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T16:41:28.862652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:41:28.862652Z digest=sha256:084557db436d661d84b4864625d3f71af16516f04f330ec57ac5ebb1cc6cc16f