Pith. sign in

Paper Citation Record · LEDGER

An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2403.06764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.06764 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:57:32.397998Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 36fb42a2-82c2-44de-8e31-e2539c974c31 · inbound

PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling cites this paper.

PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:58:29.106770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T09:58:29.057357Z digest=sha256:f79c8b7e2969b1998faa3724a9390b1ff71b49a8b51718fdbe90eed03269f02d

Observation 9129ef6c-85f5-4afb-9e3d-acebd4def756 · inbound

When Attention Sink Emerges in Language Models: An Empirical View cites this paper.

When Attention Sink Emerges in Language Models: An Empirical View An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:41:03.797181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-16T17:41:03.674759Z digest=sha256:5a14b50dc2c590c9fe2101c047642a86428b12dfeb43a5cbe309ccd17dfbb66d

Observation bd0e560b-a626-4864-acc2-3525f206e579 · inbound

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction cites this paper.

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:12:14.760555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T12:12:14.613620Z digest=sha256:dc6e596dea9589b878e80cf94b4e33a346cbf912ab0348137482d1971b95e498

Observation 2f00ea8d-456e-46c8-bfde-d319e9726ef2 · inbound

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs cites this paper.

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:57:32.397998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:57:32.397998Z digest=sha256:5c695b5252b991a0290f05c286819584692a8cb7607e4094aceeef09bfdc623e

Observation 9ece6cac-4691-4b1a-8bbb-3c84f29e0dc9 · inbound

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey cites this paper.

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 136

Resolution
unresolved
no resolver link, observed 2026-08-11T23:54:23.806149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:54:23.806149Z digest=sha256:71fc90a7940e1bfc47de6b71143a29a45bedef1d6fed1db8d9e8931ee5c64325

Observation 7a680d99-9396-4a95-a4fa-dd313db57ba8 · inbound

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs cites this paper.

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:35:24.150592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:35:24.150592Z digest=sha256:74fd56c1d54a361c11cd395a0a518acb60ead928ce66766dce6717c4ef4ca8da

Observation c5ce3a98-1b80-4e69-8f83-431aa4261138 · inbound

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression cites this paper.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.246133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.246133Z digest=sha256:c27f3de0126492dc254888cebe3ba1116717c9f1cdd662d9b1130feac7847425

Observation a1f51f80-4cd2-4477-b5aa-eaf8a52d8ad6 · inbound

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs cites this paper.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.491097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.491097Z digest=sha256:cfd73eb113cc475beb6b9359ed20af2d8c72d5a28036d6bc9c3860f7c10d50f7

Observation c065c839-cc18-4729-a2f2-bd9c412471da · inbound

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation cites this paper.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.677283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.677283Z digest=sha256:a3b90c58f7cdacb0556ef8665d6f91ebc4e1aa6fb45b31898cdff6f242d6fcde

Observation 2096eee0-1f88-45ef-b2c8-ede51a9c376a · inbound

SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer cites this paper.

SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:33:13.065992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:33:13.065992Z digest=sha256:257d4f62738d74b5ada855747a3f553c3c14dfc80ce03ef153f1f1bac46df794

Observation fd6205dd-cbcc-42d1-ba3b-e588e9721631 · inbound

LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering cites this paper.

LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:13:24.425557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:13:24.425557Z digest=sha256:cd96939e34b80e1eba1cf5141c84bd55b9922d227fe441d65b62001dcac795eb

Observation f29064d1-35f5-4560-81fa-fbeecb340d41 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.350120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.350120Z digest=sha256:0ab0fded36693d59cc1d223898b468b8e4df71825458b9185339390df0746b00

Observation b0a3ddb1-62de-458c-86b4-b28932ba393f · inbound

ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming cites this paper.

ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:39:18.051553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:39:18.051553Z digest=sha256:8abeadaa943a457b5b8c79f0f16892f135dc5f3fe08371e0e30556b3a60133e3

Observation 5d19ddcd-8134-4d6a-bcfb-06d07fa22729 · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.472844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.472844Z digest=sha256:04073b3bff7afcdb73728650658fe42a66bc039f585364818aed168ebbe14804

Observation d03d0134-b4ca-4520-9a68-69b6cf1a6e6e · inbound

What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph cites this paper.

What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:20:45.651882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:20:45.651882Z digest=sha256:629d3e43ddf963148bec65386b0384576c4c93107d3e43e41e923890c1d15c58

Observation 23577145-6b84-427d-b468-7100293c016a · inbound

LongViTU: Instruction Tuning for Long-Form Video Understanding cites this paper.

LongViTU: Instruction Tuning for Long-Form Video Understanding An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:23:57.853494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:23:57.853494Z digest=sha256:40853edd4e3526cf6e515638ce1ee10c401b581d25d32ff180ff86ec5fed12b7

Observation 2a5d5a1f-e4c3-41a6-96a5-2e4d0b889939 · inbound

AdaFV: Rethinking of Visual-Language alignment for VLM acceleration cites this paper.

AdaFV: Rethinking of Visual-Language alignment for VLM acceleration An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:00:12.568815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:00:12.568815Z digest=sha256:bee16fa3a44946a83bcc6751a7dc7d20927ffc439dae80a779dd1b49f51b8f41

Observation ec529799-0e10-4035-b845-b9cc040f1b06 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.780757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.780757Z digest=sha256:d34bb734060989c42a95618d4dc3ff042d0d9f486cea9c8b584b552801211e91

Observation 6dc287d1-f569-467c-a16f-9564fbd18c74 · inbound

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models cites this paper.

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:53.610866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:53.610866Z digest=sha256:32dbe0ccbe22d5e77c6c3d52df8bfbc2b4bdebb197d96426590593051312b67c

Observation eef22555-b744-4038-8e00-8a46b63f3888 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.568606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.568606Z digest=sha256:09b68555992e47653742a5e10bc4fd9cdd07b157dc8cf97f96b85e3ab0437d42

Observation c3b71c14-67de-4ccb-8b88-9e3581a2ebf3 · inbound

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs cites this paper.

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.165421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-19T03:18:11.993413Z digest=sha256:7d3c4486e21ee5f3f5cc2ccbd1018de66aee86e73f736eaa37725d98d9d716a3

Observation b6e18987-f9be-456d-8b71-e8b2d7397529 · inbound

SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models cites this paper.

SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T19:47:08.908760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:47:08.908760Z digest=sha256:827030913cb36ef4d69e1e68dadc2f421416779ad70cfa991b8910e4532ae02b

Observation 9f662f41-9e89-4102-bdfd-12827c0e3094 · inbound

VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness cites this paper.

VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:08.157390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T14:50:58.269817Z digest=sha256:5cabe49317efdfd7bac1a791f272cbbdb4ebbc0dd869ed833b554db5928558c4

Observation fa624b16-8128-423d-b260-c53fb5913a90 · inbound

UIPress: Bringing Optical Token Compression to UI-to-Code Generation cites this paper.

UIPress: Bringing Optical Token Compression to UI-to-Code Generation An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:00:59.145920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T17:21:32.024105Z digest=sha256:02ba08fe1415acee01124db5194f42d3e52e41856c122855a6dbc704f1d24fc5

Observation d2dc88b3-4605-468a-be51-2b4b43c1932e · inbound

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding cites this paper.

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:01.435400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:50:37.022338Z digest=sha256:51cf22a3e69803ca0badf5f57e3df2826581644cb0c9d8a350d28477caac3953

Observation 28558941-b8c5-4d6e-bbee-b078bb5e5ee3 · inbound

MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging cites this paper.

MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.050833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T22:42:20.280777Z digest=sha256:06800f1830959154406e4d7e1fe49fb0be8668469eb25cbbf126e84246dda8d4

Observation e041da95-08fc-4b5b-9c5a-86ac38731e6e · inbound

Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturation cites this paper.

Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturation An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T16:41:03.346578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T16:32:09.784716Z digest=sha256:4922704a61a53cd45a015df3b203d0bc56766d961360671a74f735c236b06850

Observation 6e8e6ee6-fd67-4474-a552-bad8bf3905f7 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 193

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:290ccfec864eb5ac702874bad3929074075e4ee19b4bd648b1b53ab988c7ac75

Observation a8b5b3c8-cc53-4915-913d-0bbd6b662b3b · inbound

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring cites this paper.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.573081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:f7c425bab88b1f84b52c338f302b3a198d02d48004e72741724830adf18b058a

Observation 0dbb35f5-9537-4245-861d-e21d3fa3da83 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:43.980109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:986929057ecc796f2177db7336473ab67c3ca5f2e2586aa4ea0b2fc4fa44b1af

Observation 8c008f01-dadb-4eac-8bac-6ab703a44618 · inbound

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation cites this paper.

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T01:48:49.788167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:48:49.788167Z digest=sha256:7c6d1741920e3f4a975f1874afc0a440a4528efdec674507070633fd50cf8db8

Observation 449bde39-58e9-4d03-a245-820be3235440 · inbound

LayoutLite: Token-Level Implicit Layout Analysis for Efficient Document OCR cites this paper.

LayoutLite: Token-Level Implicit Layout Analysis for Efficient Document OCR An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T05:34:45.516112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:34:45.516112Z digest=sha256:3c7a5d1adfd59397a2dcf8615fe03fc97105584eb4384c7031be269ca2410b6f

Observation 90366d26-3421-42e0-8ded-37882f2138cf · inbound

ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression cites this paper.

ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T00:46:04.099639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:46:04.099639Z digest=sha256:67c19fbfc92aac326d9e745eab41a8c2feeaeaa23b371f5a9948106ef73cc759

Observation 818cb237-c809-4edf-a4ac-13d5b653a7be · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.546718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.546718Z digest=sha256:9c0834ad3c8a9838a2b2c8902e43ee18751063db566f044af40e593bf9e96402

Observation b3db6853-5d03-492e-8259-085086b20159 · inbound

SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models cites this paper.

SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:41:28.812887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:41:28.812887Z digest=sha256:5e13168352b6262b540a04af3f5064406e92e7dc48495f8f24925a2ecc43fde5

Observation 54c530f6-4381-45b3-8282-d55da1520d61 · inbound

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning cites this paper.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:13.935715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:13.935715Z digest=sha256:f9d6ef628bba6f0494b87d69bd9aa5b0da138279f9107929e19f4cdcaad85fbe