Pith. sign in

Paper Citation Record · LEDGER

Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 56 inbound Pith citation observations for arXiv:2408.15998.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.15998 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 56 of 56 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:10:13.173431Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.297801Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation df506477-f157-4d6d-9192-7b5a093d8d76 · inbound

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding cites this paper.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.797984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.797984Z digest=sha256:758a13a12e146593afb7258524cb6afc0a18b10c72e7c23f6b24e38b6ebebfc1

Observation 520bca3a-9828-4f0a-8aa0-be5507544d76 · inbound

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings cites this paper.

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T06:04:21.986845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:04:21.986845Z digest=sha256:637c79ad32df3e55bad7431c1a42cc64c330638ecf75cf23858c27da4ac7a238

Observation a6cc9117-80e7-49c2-b2af-c2da7fac13b2 · inbound

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion cites this paper.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.390853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.390853Z digest=sha256:c27ea53864edc05f16ff99eced5020438949524b6b9b091a9f8c4e97afccb975

Observation 7aa4b2bf-e988-49dc-b003-ebd1a709932b · inbound

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey cites this paper.

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 144

Resolution
unresolved
no resolver link, observed 2026-08-11T23:54:23.847163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:54:23.847163Z digest=sha256:e19813558f475005084c6254a0425bdc03b7b283f719aa963862fb05eec37a22

Observation 16fb8192-cc40-498d-9fdf-55ca80f0de8c · inbound

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension cites this paper.

Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:01.551747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:01.551747Z digest=sha256:43824d994f5d64b8bb024f9f7ad48d3aafde7abb6396606e3ba93915502c6f19

Observation 16054fb2-0a0c-4f1b-8662-43a7cd5815a4 · inbound

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression cites this paper.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.369438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.369438Z digest=sha256:116e20bd5406d99682397ce0cb3ec2d91afdaaef21681f7a8cd3ac88fcc9c9ac

Observation 09010bec-878a-4b04-b1d3-177e846291a6 · inbound

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion cites this paper.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.592158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.592158Z digest=sha256:b1717d0181ec6fe183833571862fb56c18ffa61ec159bab1d3df19ce592869e9

Observation 16dc2d83-7125-423b-8fce-4acc5191ded3 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.210588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:ff8b8f257163dcf99ad8a0077e70e1178d06eb0341b75d561acda7a7b59fd8e9

Observation dd7ba175-5922-443a-84ad-b4cd2b439172 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 211

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.975453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:5bb1b89438b3543982df50dc48132b5cf6979d799c4aaf171b972f8ed3e6bbff

Observation 7fe8678d-227b-4473-9d60-eff082295c9d · inbound

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls cites this paper.

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T17:17:49.402538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:17:49.402538Z digest=sha256:be261d9f1cab559021bd91d696ed2a6a69a2a77d6266be72462ca55bf5445cdf

Observation 702ae479-9a6d-4899-9919-2aa79d520fd7 · inbound

Olympus: A Universal Task Router for Computer Vision Tasks cites this paper.

Olympus: A Universal Task Router for Computer Vision Tasks Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.434821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.434821Z digest=sha256:a0592fa41293790c5a5d1f72fa4d9d7d329d9158ea3bb18b5c073b5e937f40c9

Observation 3f58044c-919b-41c3-abff-d74cc9aff3dc · inbound

Apollo: An Exploration of Video Understanding in Large Multimodal Models cites this paper.

Apollo: An Exploration of Video Understanding in Large Multimodal Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T16:11:10.608343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:11:10.608343Z digest=sha256:d95ef45ddb084ebfa5982e554e6d635e1840c2693b8cf17b03100c7f71298015

Observation 3b98487b-ab3a-4d1d-a282-817d9a3ba1f5 · inbound

FastVLM: Efficient Vision Encoding for Vision Language Models cites this paper.

FastVLM: Efficient Vision Encoding for Vision Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T13:19:23.350189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:19:23.350189Z digest=sha256:35c40cebb85b8aaa943c9d0f52bf253159c9caa8583b524cc00c4916cbc87889

Observation 0bcd8123-4256-460d-8e79-bd6bbf39b835 · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:17.711940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:17.711940Z digest=sha256:c0952b14676137d45f93af4bae39945347b144bd0c26f45ffa714a6b0898e78f

Observation 6ab493a7-7d1a-4b78-ac95-00766ae2826f · inbound

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer cites this paper.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.893904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.893904Z digest=sha256:dc2142fcdc8f08cd4c7bdb365c584a9e5cf1572cbd8a907e613607a07d8feb2e

Observation fa9e2839-d2f8-4f32-b1b4-61908a1ba1e7 · inbound

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation cites this paper.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.259469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.259469Z digest=sha256:ab2c64b789357ff7f036513df4eb9687ddcf36473b1ca9fd47896befe4205b06

Observation 44e47afa-7329-4462-b8c0-5b62c58b3edc · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 151

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.927854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:77238d9a23c2dfa0b357a0ec61e323399b3d4fb647667b853141a65959f1c2f8

Observation cea5852e-cdf8-4cd4-90d7-0adf3091d588 · inbound

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders cites this paper.

MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:13.606238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:13.606238Z digest=sha256:24a2dbb9eae4d51162e01d85e103e544255e89c4b3eada572597b20da485b822

Observation f8093b0b-6d46-4454-a930-67eba09d1e5f · inbound

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs cites this paper.

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:16.232974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:16.232974Z digest=sha256:44715d890ba24273453d9b5aed5590a6725306e8f7d65b6291198401ec9a6430

Observation 8c1421ce-8e09-4402-ad0f-3d0f91432404 · inbound

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design cites this paper.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.156640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.156640Z digest=sha256:5163b1d5de10b6a667ac01f74f0c0203bb42c541180036b98fbbbaa4cc80a5a3

Observation 443394bb-4e23-447a-b82f-55b45b339c66 · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:33.906379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:33.906379Z digest=sha256:75d5fd6000de7ec2262f5f915b2b24fae5e77f3970c30af01fa28857047bac3b

Observation 685e67aa-bf45-4915-bfc3-5aa41164105c · inbound

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models cites this paper.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.295553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.295553Z digest=sha256:195db5703b138eddd695dde04c7936c3e3173bb72d58aa74ae1b6e6193699504

Observation f0251a32-8844-48e6-8b63-e3bcce42e86f · inbound

Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models cites this paper.

Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:58:53.820823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:58:53.820823Z digest=sha256:fafd9e67beb6bce26d81683b31fbe2812325ae5404e99471d8e02d0fd67b7e64

Observation b173ffc3-2fd3-4803-a360-30bbb851d192 · inbound

The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering cites this paper.

The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:42.200920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:42.200920Z digest=sha256:c1de7627ae4a9217c46d4ba38d401f164e571d160da36124573bba5bbc750bc7

Observation 763974f5-2b29-4718-9563-692d710ec268 · inbound

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation cites this paper.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.333188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.333188Z digest=sha256:828fd8d2b21fdf6ae27caffb5831d167e0b34d146bbe0a66f5cb11f99ed0b419

Observation 533938b4-b2bb-4b3a-bd32-4cf798a355d5 · inbound

Diffusion Instruction Tuning cites this paper.

Diffusion Instruction Tuning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.912656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.912656Z digest=sha256:f60abb1e507a909525e62e8dcf3641be68f5a5c2ef6fa8f2dbf9faaa17f1f2a1

Observation 9e50171f-9fac-4c3c-a3d1-e69360510242 · inbound

AIDE: Agentically Improve Visual Language Model with Domain Experts cites this paper.

AIDE: Agentically Improve Visual Language Model with Domain Experts Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.140407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.140407Z digest=sha256:e67826cfefb8a63ce7c46f1b0d7fbf568a6e17f91dea0ac07ba4a8a55012f8be

Observation 9fdb54f3-82dd-45cc-baa4-6bc6b85d3de0 · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.857966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:c5d2a46ef4947519c31c86ba65aa2e85f7b4b06560df4cde0ab34dd8e8a839be

Observation d77b2a76-1a39-403b-ba68-6524d9aab019 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.080193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:4e2ee5c8fec6f2b9a0b2ca3a2545345d1e5fb8d3668fb718b0bd22a7553a9e3d

Observation 1f2af3d1-7069-4150-abb6-d75a83746ec6 · inbound

Visual Compositional Tuning cites this paper.

Visual Compositional Tuning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:41:53.159814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T17:39:09.890605Z digest=sha256:e7921a4de1d1a5ffb6081b94718f5076bd127278e0576b579def27ff45e04196

Observation 3bbb0d25-af3a-423d-8185-218fc55a05ac · inbound

Causal-LLaVA: Causal Disentanglement for Mitigating Hallucination in Multimodal Large Language Models cites this paper.

Causal-LLaVA: Causal Disentanglement for Mitigating Hallucination in Multimodal Large Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:04.245081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:04.245081Z digest=sha256:8d6f277f84024976275627b78714953197d81f389b321f87430ddb05d7735374

Observation d47bf71d-1318-4c93-a5c0-783ff35a655a · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:23.941064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:23.941064Z digest=sha256:21c19f2f0bb4d330489e12744a18053f5f350229e8d609854be6773ce045e1aa

Observation 0c418386-4dd9-4852-b58d-4d8fd9c8d79b · inbound

Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts cites this paper.

Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:25:08.007033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:25:08.007033Z digest=sha256:fd462798f37946ae0f3063433f3ee118cece68e903560400467a7b91f7f96d49

Observation 6dc75d72-0f11-49d1-866f-586263a125d2 · inbound

Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs cites this paper.

Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:14.998714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:14.998714Z digest=sha256:f49a0f8900b59a4d68ade5273f60b495b9a0e1a1d519734c903132be7e6084a6

Observation 5efff50b-5682-40c5-b5d9-0ed382beb549 · inbound

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models cites this paper.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.486555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.486555Z digest=sha256:f0dffced88bb654c4be8991f438a480268835b44082b94cf94c278c554b1d380

Observation 1cb2a407-73a5-493f-bfe7-de063c93e6c1 · inbound

Hidden in plain sight: VLMs overlook their visual representations cites this paper.

Hidden in plain sight: VLMs overlook their visual representations Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.936899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.936899Z digest=sha256:0bb17442068a60f681938893b6471209bc8632ec50f1906f8c16bf74d315745b

Observation 87a83c05-843f-4c94-8eb7-9707e641f268 · inbound

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs cites this paper.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:49.320477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:49.320477Z digest=sha256:8eeebe37ff4453fa29083af45ca203f0f691ea78758f4ea6845a912f05230012

Observation 5a1ddb53-fd2a-4b2f-977d-72783eacc4ae · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:25.522410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:25.522410Z digest=sha256:9b82073c5dbba6874d6ed2cdd63156603c8a5e040532313d99452b84fb883659

Observation 0a7bcfc6-44a7-4e98-ac0f-e6e35e9920ad · inbound

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs cites this paper.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.062070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.062070Z digest=sha256:59d748ff2fa1d20849f65b40f75292e7ee6b28be977dc6ad4282cd51d4cac91c

Observation 173447f8-35ee-440e-b605-14493b863059 · inbound

FaceLLM: A Multimodal Large Language Model for Face Understanding cites this paper.

FaceLLM: A Multimodal Large Language Model for Face Understanding Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:39:41.625414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:39:41.625414Z digest=sha256:fe9f9c1f99ac3fc7a64bb21fa9e7a705a3acec3cb0a9daea6d95d344f0720f31

Observation a9b76c5c-bc67-46d1-bccd-02a02a98b93d · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:04.082371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:04.082371Z digest=sha256:72b53f9946e5a45634a725ad2defd5cfdd99433e0343f412d46ee00ec6a40a7c

Observation 69540dfa-b95c-4284-99f5-f8b6e9bd2e0d · inbound

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning cites this paper.

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.064010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T03:18:14.655384Z digest=sha256:23b5211df6d604c8a8e5a3bbd32eaac7e01024e13ed9e79ae4bdcd8e47452628

Observation bcf31caa-61fe-4cf8-8152-0b803eaec2a8 · inbound

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning cites this paper.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:20.951429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:20.951429Z digest=sha256:31918c257a45386ebd3e2211a99b96b6d68104e2ca083cfe7a333f6298566baf

Observation acbad68a-4833-4661-9087-b974a4f3c936 · inbound

EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation cites this paper.

EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:15:54.645089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T14:14:45.982490Z digest=sha256:f16b42b1b9ef1ae30182cdb0775fbe314d2c965c795d89320c2becb67ad7ee65

Observation e3a8aeeb-1a3d-4112-bac7-8ad303b07b30 · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:17.129719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:c164497ae5653d4dc38e5aac6ea2c6daea4eb8ad2cac560ea242d54138e2e1a6

Observation 37faf7fb-2289-46ff-93b1-5f5bedc8cb05 · inbound

Boosting Visual Instruction Tuning with Self-Supervised Guidance cites this paper.

Boosting Visual Instruction Tuning with Self-Supervised Guidance Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:50:58.631506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:27:52.208827Z digest=sha256:6b6dc094de79a33002ccf1cb344be2d97d0999b8755789bdf5ca30e589c3c5a6

Observation 3237e382-857d-4ef3-8222-734daab5f285 · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:57:09.535118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T02:52:43.674969Z digest=sha256:4e6c73754015719f73ff5ce8b0d478a93b07045675bc2d07886d3f3e575a3af3

Observation adcf702e-8f17-4fd1-8192-70520c6f4469 · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:29:28.665397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T21:28:37.680681Z digest=sha256:5d9a918ec47916b49555739f9e966d7561ae8582a433ddeeb8924d5345583b90

Observation 5e875590-fa29-4b24-bf74-c5fca19e609c · inbound

VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation cites this paper.

VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.251717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T13:33:12.116833Z digest=sha256:64824be4309c90fe7886ef6b1b33cfc37f1691f7e0416a08a0e3a40d1101b897

Observation 50c94100-c9d2-423f-bdaf-e5f455f52f48 · inbound

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning cites this paper.

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.581895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T13:41:44.049230Z digest=sha256:99dbc08145e44685c88c01d113d433ebb4dd12d1a0ee97768cb93c6481e329f9

Observation 13c63f98-43b8-4123-8087-67392edd9ea7 · inbound

PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization cites this paper.

PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:36:27.234746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T10:47:33.591670Z digest=sha256:23ec12e4289b743b702b0824691d2dfa6b4aa6e268eebaf91f9df8e44c772286

Observation 57aaec14-49e6-4d60-b199-643e31120ecc · inbound

Investigating Adversarial Robustness of Multi-modal Large Language Models cites this paper.

Investigating Adversarial Robustness of Multi-modal Large Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.647252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T11:11:34.152223Z digest=sha256:5c60b1faaf5ac26936c166ff69b1af0f382ab8221bceea12c7934123cade81cd

Observation 48c98fa9-8ee2-44fc-b467-2fafe97a36fc · inbound

Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs cites this paper.

Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.503741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T11:08:36.451147Z digest=sha256:dccf17063f3d25148d4899a3647c55c8f7fc17ae5dc99487f749501b2f19e09c

Observation b268367e-8616-40e0-99ff-af7dc2c90b08 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.300051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:93938a4a0e589eab7b342176a7d2fb4a8ed8f0ff28fd07b9de74e70c442c0a22

Observation 90d63a20-dd51-42f5-b4e0-0d5489f55b7b · inbound

An Exam for Active Observers cites this paper.

An Exam for Active Observers Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T21:12:07.027866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:12:07.027866Z digest=sha256:ece80085e240769c13a533b460f88969b28bb3f890e7b18129d33070f56ff36e

Observation 49abcba5-801d-4ff5-a088-ac89af31c9ab · inbound

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning cites this paper.

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T15:10:13.173431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:10:13.173431Z digest=sha256:b50aad1db4bb5a2fd7af85f097bd885dc3251e5bc5e63173e74cda207396b4c8