Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2403.06764.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:57:32.397998Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 36fb42a2-82c2-44de-8e31-e2539c974c31 · inbound
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9129ef6c-85f5-4afb-9e3d-acebd4def756 · inbound
When Attention Sink Emerges in Language Models: An Empirical View An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bd0e560b-a626-4864-acc2-3525f206e579 · inbound
PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2f00ea8d-456e-46c8-bfde-d319e9726ef2 · inbound
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ece6cac-4691-4b1a-8bbb-3c84f29e0dc9 · inbound
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 136
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a680d99-9396-4a95-a4fa-dd313db57ba8 · inbound
A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5ce3a98-1b80-4e69-8f83-431aa4261138 · inbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1f51f80-4cd2-4477-b5aa-eaf8a52d8ad6 · inbound
[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c065c839-cc18-4729-a2f2-bd9c412471da · inbound
Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2096eee0-1f88-45ef-b2c8-ede51a9c376a · inbound
SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd6205dd-cbcc-42d1-ba3b-e588e9721631 · inbound
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f29064d1-35f5-4560-81fa-fbeecb340d41 · inbound
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0a3ddb1-62de-458c-86b4-b28932ba393f · inbound
ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d19ddcd-8134-4d6a-bcfb-06d07fa22729 · inbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d03d0134-b4ca-4520-9a68-69b6cf1a6e6e · inbound
What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23577145-6b84-427d-b468-7100293c016a · inbound
LongViTU: Instruction Tuning for Long-Form Video Understanding An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a5d5a1f-e4c3-41a6-96a5-2e4d0b889939 · inbound
AdaFV: Rethinking of Visual-Language alignment for VLM acceleration An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec529799-0e10-4035-b845-b9cc040f1b06 · inbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dc287d1-f569-467c-a16f-9564fbd18c74 · inbound
GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eef22555-b744-4038-8e00-8a46b63f3888 · inbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3b71c14-67de-4ccb-8b88-9e3581a2ebf3 · inbound
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b6e18987-f9be-456d-8b71-e8b2d7397529 · inbound
SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f662f41-9e89-4102-bdfd-12827c0e3094 · inbound
VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fa624b16-8128-423d-b260-c53fb5913a90 · inbound
UIPress: Bringing Optical Token Compression to UI-to-Code Generation An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d2dc88b3-4605-468a-be51-2b4b43c1932e · inbound
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 28558941-b8c5-4d6e-bbee-b078bb5e5ee3 · inbound
MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e041da95-08fc-4b5b-9c5a-86ac38731e6e · inbound
Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturation An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6e8e6ee6-fd67-4474-a552-bad8bf3905f7 · inbound
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 193
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8b5b3c8-cc53-4915-913d-0bbd6b662b3b · inbound
AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0dbb35f5-9537-4245-861d-e21d3fa3da83 · inbound
What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8c008f01-dadb-4eac-8bac-6ab703a44618 · inbound
Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 449bde39-58e9-4d03-a245-820be3235440 · inbound
LayoutLite: Token-Level Implicit Layout Analysis for Efficient Document OCR An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90366d26-3421-42e0-8ded-37882f2138cf · inbound
ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 818cb237-c809-4edf-a4ac-13d5b653a7be · inbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3db6853-5d03-492e-8259-085086b20159 · inbound
SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54c530f6-4381-45b3-8282-d55da1520d61 · inbound
Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.