Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 56 inbound Pith citation observations for arXiv:2408.15998.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:10:13.173431Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T15:09:55.297801Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation df506477-f157-4d6d-9192-7b5a093d8d76 · inbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 520bca3a-9828-4f0a-8aa0-be5507544d76 · inbound
Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6cc9117-80e7-49c2-b2af-c2da7fac13b2 · inbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7aa4b2bf-e988-49dc-b003-ebd1a709932b · inbound
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 144
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16fb8192-cc40-498d-9fdf-55ca80f0de8c · inbound
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16054fb2-0a0c-4f1b-8662-43a7cd5815a4 · inbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09010bec-878a-4b04-b1d3-177e846291a6 · inbound
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16dc2d83-7125-423b-8fce-4acc5191ded3 · inbound
NVILA: Efficient Frontier Visual Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dd7ba175-5922-443a-84ad-b4cd2b439172 · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 211
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7fe8678d-227b-4473-9d60-eff082295c9d · inbound
YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 702ae479-9a6d-4899-9919-2aa79d520fd7 · inbound
Olympus: A Universal Task Router for Computer Vision Tasks Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f58044c-919b-41c3-abff-d74cc9aff3dc · inbound
Apollo: An Exploration of Video Understanding in Large Multimodal Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b98487b-ab3a-4d1d-a282-817d9a3ba1f5 · inbound
FastVLM: Efficient Vision Encoding for Vision Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bcd8123-4256-460d-8e79-bd6bbf39b835 · inbound
Do Language Models Understand Time? Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 133
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ab493a7-7d1a-4b78-ac95-00766ae2826f · inbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa9e2839-d2f8-4f32-b1b4-61908a1ba1e7 · inbound
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44e47afa-7329-4462-b8c0-5b62c58b3edc · inbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 151
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cea5852e-cdf8-4cd4-90d7-0adf3091d588 · inbound
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8093b0b-6d46-4454-a930-67eba09d1e5f · inbound
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c1421ce-8e09-4402-ad0f-3d0f91432404 · inbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 443394bb-4e23-447a-b82f-55b45b339c66 · inbound
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 685e67aa-bf45-4915-bfc3-5aa41164105c · inbound
LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0251a32-8844-48e6-8b63-e3bcce42e86f · inbound
Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b173ffc3-2fd3-4803-a360-30bbb851d192 · inbound
The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 763974f5-2b29-4718-9563-692d710ec268 · inbound
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 533938b4-b2bb-4b3a-bd32-4cf798a355d5 · inbound
Diffusion Instruction Tuning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e50171f-9fac-4c3c-a3d1-e69360510242 · inbound
AIDE: Agentically Improve Visual Language Model with Domain Experts Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fdb54f3-82dd-45cc-baa4-6bc6b85d3de0 · inbound
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d77b2a76-1a39-403b-ba68-6524d9aab019 · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1f2af3d1-7069-4150-abb6-d75a83746ec6 · inbound
Visual Compositional Tuning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3bbb0d25-af3a-423d-8185-218fc55a05ac · inbound
Causal-LLaVA: Causal Disentanglement for Mitigating Hallucination in Multimodal Large Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d47bf71d-1318-4c93-a5c0-783ff35a655a · inbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c418386-4dd9-4852-b58d-4d8fd9c8d79b · inbound
Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dc75d72-0f11-49d1-866f-586263a125d2 · inbound
Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5efff50b-5682-40c5-b5d9-0ed382beb549 · inbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cb2a407-73a5-493f-bfe7-de063c93e6c1 · inbound
Hidden in plain sight: VLMs overlook their visual representations Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87a83c05-843f-4c94-8eb7-9707e641f268 · inbound
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 111
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a1ddb53-fd2a-4b2f-977d-72783eacc4ae · inbound
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a7bcfc6-44a7-4e98-ac0f-e6e35e9920ad · inbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 173447f8-35ee-440e-b605-14493b863059 · inbound
FaceLLM: A Multimodal Large Language Model for Face Understanding Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9b76c5c-bc67-46d1-bccd-02a02a98b93d · inbound
Docopilot: Improving Multimodal Models for Document-Level Understanding Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69540dfa-b95c-4284-99f5-f8b6e9bd2e0d · inbound
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bcf31caa-61fe-4cf8-8152-0b803eaec2a8 · inbound
MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acbad68a-4833-4661-9087-b974a4f3c936 · inbound
EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e3a8aeeb-1a3d-4112-bac7-8ad303b07b30 · inbound
CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 37faf7fb-2289-46ff-93b1-5f5bedc8cb05 · inbound
Boosting Visual Instruction Tuning with Self-Supervised Guidance Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3237e382-857d-4ef3-8222-734daab5f285 · inbound
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation adcf702e-8f17-4fd1-8192-70520c6f4469 · inbound
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5e875590-fa29-4b24-bf74-c5fca19e609c · inbound
VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 50c94100-c9d2-423f-bdaf-e5f455f52f48 · inbound
ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 13c63f98-43b8-4123-8087-67392edd9ea7 · inbound
PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 57aaec14-49e6-4d60-b199-643e31120ecc · inbound
Investigating Adversarial Robustness of Multi-modal Large Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 48c98fa9-8ee2-44fc-b467-2fafe97a36fc · inbound
Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b268367e-8616-40e0-99ff-af7dc2c90b08 · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 90d63a20-dd51-42f5-b4e0-0d5489f55b7b · inbound
An Exam for Active Observers Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49abcba5-801d-4ff5-a088-ac89af31c9ab · inbound
Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.