Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2406.08487.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:28:36.781880Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T19:52:01.847188Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation a9e8b27d-9201-4c1e-92c0-a0ba5b52c41d · inbound
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c85c23ed-b540-42fa-b5ae-eb93a40efafd · inbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eec66a7f-c702-4e86-bd74-b72d55c96ba4 · inbound
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 050d4782-a562-4c0b-ad4b-9dfb6841be5f · inbound
Large Language Model-Brained GUI Agents: A Survey Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 235
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6e61ed9f-6618-42e0-922a-6d32942ed904 · inbound
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd389ee4-b4b8-4a57-9470-38c4779e0df1 · inbound
Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2293816a-e9c4-45b7-9785-2041aec0a79d · inbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee01eec7-4e0c-4fbb-bccc-f7315f330684 · inbound
p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b17128a6-cf98-48ea-bcac-7682aca48906 · inbound
LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3f4f408-d529-4635-8f5f-6201c447c5eb · inbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 119
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df1d01a3-4cf9-421d-9494-3812d1d04ec7 · inbound
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a1add6b-2778-4f07-9329-22dcd1aed3ee · inbound
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b9ef01a3-a29c-4a5e-9094-8278d42682ff · inbound
CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2894911-4be6-4cfa-b667-da2a3ece02c6 · inbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 110
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06e0b997-010a-4be1-bb5a-16b4b227a6d5 · inbound
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe5d7b0f-5675-4012-8c79-da07e7ebffe9 · inbound
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b382da80-0c94-48de-b22d-f2e8fd4ba5eb · inbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c8b5f18-f5f1-4841-8f7e-5927283508ab · inbound
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8a1bd3f-8f1a-4e44-ac29-90882517e885 · inbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55c7d61c-77b3-48d8-a6b7-88c4aa8ead28 · inbound
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fc0e204c-eed9-4d35-9456-4e775fb30aa4 · inbound
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1c3bbed-f522-4f09-a95f-6f74449b07b3 · inbound
RESAnything: Attribute Prompting for Arbitrary Referring Segmentation Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9b2a7ad-a09d-464c-adfa-027d2a0f65bb · inbound
Beyond Hard and Soft: Hybrid Context Compression for Balancing Local and Global Information Retention Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80f2e1f6-1317-4010-b328-cee0148f0e27 · inbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 309da78b-8d1e-4153-836e-aa64ee4ed65e · inbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e3c2543-df55-4b6f-86a9-bb839450d42e · inbound
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 123
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 858439a1-fa26-45ca-af18-97a616067348 · inbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2f9c8825-b548-405f-a355-b60ea1a8aa1c · inbound
Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bfbd8f8-02de-442e-a511-6d69283bce5f · inbound
CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54d06397-c0d5-4151-a839-fb563950427b · inbound
Mitigating Coordinate Prediction Bias from Positional Encoding Failures Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2825c939-e432-4838-a4b1-6456fd772dab · inbound
HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2517de30-451d-4e1f-ad06-5ddc7ee99579 · inbound
ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation aa753e71-a650-4be4-807e-36b50e213599 · inbound
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 199
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a18acc9-376b-4343-b92c-c026baf70eed · inbound
Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.