Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2408.12637.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:49:10.837634Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T05:39:39.660194Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 0f304a0f-a3ca-4e71-9e9a-db73504ec9c7 · inbound
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark Building and better understanding vision-language models: insights and future directions
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 33e2b272-aef7-43ed-a3d4-b2fd70c979bb · inbound
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Building and better understanding vision-language models: insights and future directions
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 92ab786f-14a3-4d3f-91b0-43fa501d2784 · inbound
Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models Building and better understanding vision-language models: insights and future directions
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e481b500-b1fb-47f6-a568-685b8dc5b950 · inbound
VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models Building and better understanding vision-language models: insights and future directions
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ca4f020-57c9-4508-95f9-499a99c5fe63 · inbound
NVILA: Efficient Frontier Visual Language Models Building and better understanding vision-language models: insights and future directions
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7fe45a35-7a91-445e-a1e1-5e349c1c25dd · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Building and better understanding vision-language models: insights and future directions
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 020c5bb4-6836-4b3f-9270-e1dfc5e9856d · inbound
Apollo: An Exploration of Video Understanding in Large Multimodal Models Building and better understanding vision-language models: insights and future directions
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bad5c29-2f67-490c-be1d-f7bb08f05aa9 · inbound
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Building and better understanding vision-language models: insights and future directions
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4662635b-f70f-440c-bc44-1957f06b8306 · inbound
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Building and better understanding vision-language models: insights and future directions
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 627db8c1-9b90-473b-b6c5-3bff1107c01a · inbound
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation Building and better understanding vision-language models: insights and future directions
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7969d33-d702-405b-9d60-6ab0036c8e96 · inbound
ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark Building and better understanding vision-language models: insights and future directions
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e4f1832-6a6e-4fa8-8245-cedc8457168b · inbound
MSTS: A Multimodal Safety Test Suite for Vision-Language Models Building and better understanding vision-language models: insights and future directions
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39a1499e-823e-4baa-96a7-1963287bdc7f · inbound
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Building and better understanding vision-language models: insights and future directions
Reference 108
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 603ad2fc-c9ac-432a-963e-be588269380b · inbound
The Impact of Persona-based Political Perspectives on Hateful Content Detection Building and better understanding vision-language models: insights and future directions
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f78ef67f-3e12-4473-8176-d60d129373ad · inbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Building and better understanding vision-language models: insights and future directions
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10b5a689-0ad0-401e-a7ac-e47ca724df58 · inbound
SmolVLM: Redefining small and efficient multimodal models Building and better understanding vision-language models: insights and future directions
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cb1f77ae-97fa-40b4-93a4-fe4f374436a2 · inbound
Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency Building and better understanding vision-language models: insights and future directions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfe53a72-f40e-43b5-b681-e17e28e4a155 · inbound
R^3-VQA: "Read the Room" by Video Social Reasoning Building and better understanding vision-language models: insights and future directions
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dd34dc0-0697-4789-a243-673797ff84f3 · inbound
FG-CLIP: Fine-Grained Visual and Textual Alignment Building and better understanding vision-language models: insights and future directions
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2021fc53-9986-495d-b081-c67fdc160c3b · inbound
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Building and better understanding vision-language models: insights and future directions
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de154433-0b74-42e1-91f1-83e5f4c6ba35 · inbound
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Building and better understanding vision-language models: insights and future directions
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8251c337-49a0-4d1b-ba3b-3da9da9008f3 · inbound
Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data Building and better understanding vision-language models: insights and future directions
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 280158b9-84d2-4448-93ab-a08974e26520 · inbound
Vision-Language Models Can't See the Obvious Building and better understanding vision-language models: insights and future directions
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed95fb85-610e-4847-a8c5-56b3496cd2ac · inbound
Docopilot: Improving Multimodal Models for Document-Level Understanding Building and better understanding vision-language models: insights and future directions
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d39e0ca0-c515-4ee9-844a-d43e7d6201cd · inbound
LMM-Det: Make Large Multimodal Models Excel in Object Detection Building and better understanding vision-language models: insights and future directions
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e254639f-2e33-484c-87c7-28b6ff9e545a · inbound
VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Building and better understanding vision-language models: insights and future directions
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3a56c70-296c-42e0-bd61-7193d84befbb · inbound
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information Building and better understanding vision-language models: insights and future directions
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 887a9bfd-dcdf-4a24-99f5-8af7eb53e7af · inbound
Measuring Epistemic Humility in Multimodal Large Language Models Building and better understanding vision-language models: insights and future directions
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73d926d0-48cf-4725-bec6-a78b63643ab4 · inbound
MultiMat: Multimodal Program Synthesis for Procedural Materials using Large Multimodal Models Building and better understanding vision-language models: insights and future directions
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d11ee3f7-00d1-46f5-8818-e8275eecb668 · inbound
Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models Building and better understanding vision-language models: insights and future directions
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1a1b7cc-4327-4d06-8e40-7aaa6efa2abc · inbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Building and better understanding vision-language models: insights and future directions
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eae8d05-7347-43e8-ae0c-00d41ee2405e · inbound
Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning Building and better understanding vision-language models: insights and future directions
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cb705c73-6cf4-46f0-ab9c-b9f4c94fd856 · inbound
Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models Building and better understanding vision-language models: insights and future directions
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8e96df26-dce8-45a1-946b-530100747880 · inbound
DenTab: A Dataset for Table Recognition and Visual QA on Real-World Dental Estimates Building and better understanding vision-language models: insights and future directions
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1100dfe0-713e-42a0-aa91-4df2231482c6 · inbound
PBSBench: A Multi-Level Vision-Language Framework and Benchmark for Hematopathology Whole Slide Image Interpretation Building and better understanding vision-language models: insights and future directions
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5357719c-62c0-442c-9aa0-19b21d897cfc · inbound
ZAYA1-VL-8B Technical Report Building and better understanding vision-language models: insights and future directions
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dec68707-2fe8-4e97-ac41-1169b7d858c8 · inbound
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone Building and better understanding vision-language models: insights and future directions
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 878d8f83-82c1-45cb-a155-b8526dfe88ec · inbound
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone Building and better understanding vision-language models: insights and future directions
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5b374f96-111f-4955-a00a-e20b6dff8c51 · inbound
VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding Building and better understanding vision-language models: insights and future directions
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6c79169a-4d02-4fbe-8b40-9d751204ce42 · inbound
Zamba2-VL Technical Report Building and better understanding vision-language models: insights and future directions
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8398975b-134e-490a-8000-37d2a02181e8 · inbound
Stellar: Scalable Multimodal Document Retrieval for Natural Language Queries Building and better understanding vision-language models: insights and future directions
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 13e6751d-735c-4c94-8cd0-5ab2e8566b25 · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models Building and better understanding vision-language models: insights and future directions
Reference 143
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 38ab0453-c911-4761-86b7-c1cdacef7631 · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models Building and better understanding vision-language models: insights and future directions
Reference 143
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c636bb41-f0a3-415f-b6c6-15f95f16eeb9 · inbound
MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement Building and better understanding vision-language models: insights and future directions
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cfbcc96-ccc5-4fce-b4bb-fc008816a546 · inbound
MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents Building and better understanding vision-language models: insights and future directions
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e363be95-6a85-4242-aa43-7a3606fa7fdd · inbound
Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution Building and better understanding vision-language models: insights and future directions
Reference 114
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22d9cf6a-c6dc-4053-90f6-c999d79303cd · inbound
DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation Building and better understanding vision-language models: insights and future directions
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa545273-8786-4b5c-bc0e-179698f21fa7 · inbound
Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ Building and better understanding vision-language models: insights and future directions
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.