Pith. sign in

Paper Citation Record · LEDGER

Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2409.10197.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.10197 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T06:04:22.053400Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:15:44.600996Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 393329a3-b647-4b1a-8236-92e1c3c38c42 · inbound

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings cites this paper.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:22.053400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:22.053400Z digest=sha256:6ea1b031f2e9b5410ae463b9389c9fa9b3292c1f1e9729144db6cddb37b72f3b

Observation a0cc11a0-a1b1-4e58-9058-1d1139de049f · inbound

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs cites this paper.

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T00:57:32.553659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:57:32.553659Z digest=sha256:639a160cb33d2891a55a338244fd11a44c8fd45b70ce3b6f759cc6c16e6b2e0c

Observation 55a8b314-8652-4f49-8c5e-a1fc978e722e · inbound

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression cites this paper.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.393098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.393098Z digest=sha256:e1a67f95112677abd2283e3a0c15473f298064e2f169619d681373fde7234055

Observation 1bc03161-2ce6-4819-8add-7b320ac75a8a · inbound

Pruning All-Rounder: Rethinking and Improving Inference Efficiency for Large Vision Language Models cites this paper.

Pruning All-Rounder: Rethinking and Improving Inference Efficiency for Large Vision Language Models Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T19:42:46.637405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:42:46.637405Z digest=sha256:f114b9ef8244e5735280c899c18d62f7f8be0336a25396c23e940827ed75ada1

Observation 97da0c74-c780-4d18-be93-225255075fca · inbound

AdaFV: Rethinking of Visual-Language alignment for VLM acceleration cites this paper.

AdaFV: Rethinking of Visual-Language alignment for VLM acceleration Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:00:12.651532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:00:12.651532Z digest=sha256:08a0c11dcb310f835990ad0df18ed9a1fe0416286c89d416c63523173e3c818d

Observation 04a75efd-fe91-4a21-b94f-83b37927bfcd · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.764089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:90a10631a11f26d26ca8714231ad84c0b2d643379cf80e1d6a0208cf15698220

Observation 35d14d73-0835-4788-9491-3e49152ba743 · inbound

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models cites this paper.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:31.752838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:31.752838Z digest=sha256:b083bb133a8af929e04500448990db4269083e89d66ce249388d32c4cc34a630

Observation 9d18b4bf-9bd0-4340-9b6c-470f3b9efc82 · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.524477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.524477Z digest=sha256:5a8ffa77b60d6e3c7db257a4b83528ff24681e19de0a4335c67583017af5f791

Observation ab771970-2e0f-4ef3-9662-e90719abfbf4 · inbound

SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models cites this paper.

SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T19:47:08.863100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:47:08.863100Z digest=sha256:521713dbab97fe4a85597e266df45c285e4a9aa748fcc1c5f7f6ab86a00b1727

Observation f60cfa21-2421-4bf9-bc2c-8e4b2c0a7cc9 · inbound

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling cites this paper.

ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.638913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T00:56:47.841355Z digest=sha256:81cec150fb230bd6219302bde1ee20f3142bb44d31f9b5a58228ca4bc54368dc

Observation 8cd29f63-cb68-4c29-943a-5da731e9ceee · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 204

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.765270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:e2345c2c37d526933ac8175134066246014f3d076edcae37f3afe5ebad55ba80

Observation c8cfb87f-aaed-467f-9a20-037a7f0d1b85 · inbound

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs cites this paper.

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.602300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:41:04.184461Z digest=sha256:a953537c0f1b00c90df82e1ce60715f4aaf0860d7b9e2549321063a8f48b72f8

Observation b2a326a4-24bb-40bd-9e70-61d1d6b7b4b8 · inbound

SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models cites this paper.

SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T16:41:28.860018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:41:28.860018Z digest=sha256:f8ccb2a7814a870404df6c1c3304029ce1d9bbd7d5047448331cdbf785527e70

Observation 03fcfec7-111b-4c18-8d98-5ee76158c127 · inbound

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware cites this paper.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.837652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.837652Z digest=sha256:d0d78318ef9e86c2feb6fd8374df09741fea1ffeeedd12e8a8f1da885b46aec1