Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2406.08487.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T12:54:56.387608Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T19:52:01.847188Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation a9e8b27d-9201-4c1e-92c0-a0ba5b52c41d · inbound
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 050d4782-a562-4c0b-ad4b-9dfb6841be5f · inbound
Large Language Model-Brained GUI Agents: A Survey Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 235
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2a1add6b-2778-4f07-9329-22dcd1aed3ee · inbound
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8c8b5f18-f5f1-4841-8f7e-5927283508ab · inbound
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8a1bd3f-8f1a-4e44-ac29-90882517e885 · inbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55c7d61c-77b3-48d8-a6b7-88c4aa8ead28 · inbound
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e9b2a7ad-a09d-464c-adfa-027d2a0f65bb · inbound
Beyond Hard and Soft: Hybrid Context Compression for Balancing Local and Global Information Retention Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80f2e1f6-1317-4010-b328-cee0148f0e27 · inbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 309da78b-8d1e-4153-836e-aa64ee4ed65e · inbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e3c2543-df55-4b6f-86a9-bb839450d42e · inbound
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 123
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 858439a1-fa26-45ca-af18-97a616067348 · inbound
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2f9c8825-b548-405f-a355-b60ea1a8aa1c · inbound
Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bfbd8f8-02de-442e-a511-6d69283bce5f · inbound
CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54d06397-c0d5-4151-a839-fb563950427b · inbound
Mitigating Coordinate Prediction Bias from Positional Encoding Failures Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2825c939-e432-4838-a4b1-6456fd772dab · inbound
HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2517de30-451d-4e1f-ad06-5ddc7ee99579 · inbound
ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aa753e71-a650-4be4-807e-36b50e213599 · inbound
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 199
Source-reported events for the cited work
Unavailable: canonical work link unavailable.