Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2403.00231.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:37.086555Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T20:48:56.198380Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 3a50e86b-4986-4c4a-b5cd-c8852fac54da · inbound
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 60f2322e-4ed9-4fe8-ad74-0d9acc528876 · inbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d7c5483a-670c-4b63-bdef-68f7bf107a24 · inbound
PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 27fa901c-1d2b-47ab-b17f-7e3e7b658d12 · inbound
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 763d7f88-6794-437d-8ca5-263415c11c38 · inbound
BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 601ddb87-e901-45fb-aa69-054dc81ebdb7 · inbound
ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8d921e0-e543-4d53-8b04-811b89fa6e3d · inbound
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b14f80-db4f-486f-b7ba-a97b6c857c51 · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f7fd042b-5362-4d7c-b127-ea1c40380f22 · inbound
Chimera: Improving Generalist Model with Domain-Specific Experts Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c7cbd6e-9714-4531-9133-d38be1457413 · inbound
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f70a4bce-d6c9-4da7-acce-b9facd44ab2f · inbound
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9527b3a2-c3fc-4e2a-8f70-92a477a19b5b · inbound
Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c93a2ead-e820-4d5d-94bf-0dac3f7f4884 · inbound
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d046f113-cf1f-414c-9e9d-f21215c262d6 · inbound
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d238fc5d-c011-4afe-acad-2d0c4eec1d98 · inbound
Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 238d0a45-84fd-4159-8213-2acd3e8df778 · inbound
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 26d54d24-5ff5-4297-9863-0ca74008e9f4 · inbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 134
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebfac55c-0657-459d-915c-6b25fd72339e · inbound
MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24f3f26e-458f-420b-9245-9692e789d51f · inbound
ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8a5e407-34d7-4780-8eaf-88fd4a99f7d5 · inbound
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2181e933-b798-4bb5-957a-695d7eefdb49 · inbound
FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e003bd98-94d9-497e-8c18-a6ab3d9cb463 · inbound
ChartCap: Mitigating Hallucination of Dense Chart Captioning Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 871df512-c395-40d8-bb72-368936f975b7 · inbound
Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f894236-9082-4dc2-b385-e47f7738b3c1 · inbound
Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7da3b7e4-3912-42b7-8476-7b73c77741b9 · inbound
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eface75-9228-4622-b38e-64ada7d9de03 · inbound
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54c5659c-cbc6-4c45-b550-49f93447698b · inbound
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6317614-23f2-40cf-9376-adde0a86f484 · inbound
DeepEyesV2: Toward Agentic Multimodal Model Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b22e4d8c-5e39-4e47-852b-d0e060a42f9c · inbound
Attention Grounded Enhancement for Visual Document Retrieval Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4f52da81-0a9f-4f5a-9d0a-ed1ddf0e6d11 · inbound
Spectral Imbalance Causes Forgetting in Low-Rank Continual Adaptation Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ab9584a-fef9-4d1f-ba7b-4957d8e3abe5 · inbound
GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a57c095b-5aa7-4e33-b881-eb084d0b059e · inbound
SciFigDetect: A Benchmark for AI-Generated Scientific Figure Detection Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2cf12307-a70e-41da-8af0-c36a875a6a96 · inbound
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1142182e-cf02-4766-a44b-f99b9fedb515 · inbound
Perceptual Flow Network for Visually Grounded Reasoning Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a6ff8011-8189-494e-a938-9a5a1407433a · inbound
Octopus: History-Free Gradient Orthogonalization for Continual Learning in Multimodal Large Language Models Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 056fba95-f8ea-4e27-99e0-371d26e3b8d1 · inbound
Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c5eaa9fa-2347-4767-9243-568217237c48 · inbound
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8d972fdb-c888-4348-b7f9-aeb63c740a87 · inbound
Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 28741f6f-23d8-48ae-b58a-05a2601e5872 · inbound
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 58ba945a-b848-45bf-bee5-0fa8910d1de3 · inbound
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 168
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4c2cf58-85e1-4a9c-9d72-cfc0dda00e0d · inbound
PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03777384-0678-4999-9811-655003dc4633 · inbound
PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.