Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2501.03895.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:32:28.293861Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 9da856d3-e288-4a0d-90b0-9d9b6ae8b686 · inbound
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86aeb704-3adc-4170-afc5-5eac48021193 · inbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5d1cee3-db0c-4fd8-b453-1aa778b5ed74 · inbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80dec6b0-91d2-4194-a46b-e89c217429fe · inbound
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fa331de-f931-4698-b381-4bd57deec528 · inbound
DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36dc7495-b72b-4676-9c79-17a83e1ba0e0 · inbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e812d10-ea7d-45f4-80f4-3e28783cfb9e · inbound
Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb4cd5d7-3937-4748-ae83-c7fe93378ead · inbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1f0c0cb-268c-402c-aad2-c20d039f92d4 · inbound
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 220f8bba-cf8e-446a-b0d3-556a1f5f6b26 · inbound
NoLoCo: No-all-reduce Low Communication Training Method for Large Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a45288c7-d67a-44b0-81fd-85d7a70a5977 · inbound
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9433110-79f7-4096-895e-e2f652542aaa · inbound
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 806989ac-a5bb-43c2-aff6-6e3192fa1a0a · inbound
Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7271e21-b9d0-4e84-8ef0-a93975e2687e · inbound
Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation be70dd87-0aee-4f07-a90a-33ab85d0a529 · inbound
Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20621b8b-84bf-4333-bd44-84af059d8666 · inbound
Synthetic Homes: A Multimodal Generative AI Pipeline for Residential Building Data Generation under Data Scarcity LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4730ebf1-852b-481d-81f1-9ddd9edb92b7 · inbound
Synthetic Homes: A Multimodal Generative AI Pipeline for Residential Building Data Generation under Data Scarcity LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28e56081-00fd-499b-8177-d31081ce1074 · inbound
EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4557b983-9211-4401-936e-7bf88f5b7569 · inbound
FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba33d555-e37f-4f04-b8c6-f4f9b1180065 · inbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d3bd22d-f1d0-434f-a862-2d9b7bdc212a · inbound
Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ee25deb0-089f-455c-adea-041508b539fc · inbound
Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5f1b2b21-8ab8-422c-9e15-cc1db7e8f5bc · inbound
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 52b28f62-8889-4351-94fb-d11f64bf0b04 · inbound
Geometry-Guided 3D Visual Token Pruning for Video-Language Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 037ff41f-8d0c-4e46-98a5-f3a2270c6af9 · inbound
VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bec6f267-091b-4cb4-a809-46d6344c8821 · inbound
LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation feb54c7e-bd0f-41d2-8b8f-9b63ec5c1470 · inbound
OProver: A Unified Framework for Agentic Formal Theorem Proving LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 171008aa-deaf-41a7-a9c0-d7ae860611bf · inbound
Focus-then-Context: Subject-Centric Progressive Visual Token Reduction for Vision-Language Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 690383a7-c803-4306-8b92-9682bf0bf7e8 · inbound
VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c62617d9-5947-41e0-b401-21df000f508f · inbound
CIVIC: End-to-End Sequence Compactness for Efficient Vision-Language Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a0720b38-b6e7-43a7-abf4-ca9f78a33327 · inbound
The Hidden Evolution of Disguised Visual Context inside the VLM LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c01c7c92-3aae-491b-a2ad-9db3abb7ac87 · inbound
ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 85f39ace-49c0-4eb0-b2ad-36bfa1e9112e · inbound
MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 500c4645-e0ec-476e-979a-18597bb1ea26 · inbound
EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation af44867e-005c-483f-afe5-ce72a7bc4d89 · inbound
PDD-RRG: Posterior Diagnostic Decision for Study-level Radiology Report Generation LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 100f7f4a-5ab0-4897-aeac-f15ce15c3ddc · inbound
Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cde152f-f8de-45dd-8e0e-f00f4a94f582 · inbound
Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.