Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2407.02392.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:55:03.126780Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T10:15:44.617819Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 5ecbc9d4-5d8e-43c3-a3a6-0753f84636a3 · inbound
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b9df33b0-369c-4531-afcb-3e94984d212a · inbound
FoPru: Focal Pruning for Efficient Large Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0a986b0-1320-418f-be09-e018279d5e4d · inbound
freePruner: A Training-free Approach for Large Multimodal Model Acceleration TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f87dbc8-cfd3-48d1-934e-08ca0e0c21ee · inbound
Importance-Based Token Merging for Efficient Image and Video Generation TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5a7df4a-714f-4097-893a-1e71631a0529 · inbound
Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb1ee894-a87f-497c-acc4-0e3fbf1ef691 · inbound
Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 715ec7c7-15a8-4847-9b3e-4a3cdc18d2b3 · inbound
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64483dff-5527-4c2b-ba48-8dfb8a2a7b10 · inbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0873528-fedb-49f4-88f7-8c9689af5dca · inbound
DocVLM: Make Your VLM an Efficient Reader TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59e7046c-d4fb-4a3a-b6fb-0d8f77c9a4cd · inbound
SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb05bf86-a2c6-460a-8300-fb344eaa035b · inbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d591470-788b-498c-9f25-6218ac1d9be2 · inbound
ST$^3$: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 726c841c-55fc-46e1-8a32-8eb056635bd1 · inbound
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71909ff6-8976-435a-8fde-f3e3fed2aac0 · inbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 996775d2-0132-4d8e-9dc0-582210f9d4f1 · inbound
DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b7a0790-66f6-455b-a4f4-4dedf1e16902 · inbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b57cb93c-d8e1-4482-b0b8-4d92a2d3f68b · inbound
STAR: Stage-Wise Attention-Guided Token Reduction for Efficient Large Vision-Language Models Inference TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ad7117d-d750-496d-bf95-cb4432bb11c7 · inbound
PixelThink: Towards Efficient Chain-of-Pixel Reasoning TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f5a24b9-6c20-4983-ad8f-d9e2bcdde671 · inbound
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81818563-8738-439b-a096-f401ce88c5b9 · inbound
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1413e60-66b2-4474-8cc5-4e90d4930ee2 · inbound
Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8315937-1c4c-4eed-a4d5-020faa7412bf · inbound
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4adaa931-5207-495e-bbd1-7f0ca55cb3db · inbound
HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c62664f3-9f31-4c76-a0de-2a28e1bfcca7 · inbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4294a87b-37ca-4dd7-8684-3793d217ceef · inbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ede358af-656a-4c9c-a204-21158acab1c5 · inbound
Continual Learning for Generative AI: From LLMs to MLLMs and Beyond TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 118
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c934db64-3e50-42a6-b795-2596f8f744e4 · inbound
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb2fa175-544f-4b44-8abd-a3127afaa569 · inbound
FACap: A Large-scale Fashion Dataset for Fine-grained Composed Image Retrieval TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c424190-ddff-486c-a6ef-16aaf8b4e4ef · inbound
Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3243f840-cfe4-4a19-8f2d-3f79e8bfccb5 · inbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3e74dd7-8293-4ec6-9204-f21aaa904a8f · inbound
HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b04a3492-ffc0-4989-92c4-271b99e7bb8a · inbound
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a3c9763-f343-4d6b-95d7-997527b62566 · inbound
Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9edf9c59-0678-425a-9243-243392fd084d · inbound
CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b131883-97a4-4f14-97af-e2abeb8c3d79 · inbound
Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d937da5-6184-48f7-91e4-6eaf389f73b2 · inbound
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 730fc652-2479-4f49-9878-86f56d5b5466 · inbound
UIPress: Bringing Optical Token Compression to UI-to-Code Generation TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation de47aec6-9bef-40d4-9786-bfe0e2b88f70 · inbound
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2500fc84-b94d-418c-a669-5923dcef4800 · inbound
MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8f953d7f-c32b-4170-910c-9dc353963321 · inbound
MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 029a34b8-1856-4299-8a32-a5378955ea5a · inbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a842a616-3f4c-4e81-9abb-82e1759158ba · inbound
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c74d0f9-56b4-442e-91c8-51e2e8a3ec73 · inbound
Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.