Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2406.06462.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:31:36.856649Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T17:18:43.988790Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 7a856c60-8a1e-498a-be79-9ab924c70cf5 · inbound
CogVLM2: Visual Language Models for Image and Video Understanding VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5499c804-4a2a-4638-a85c-ae477b44d24e · inbound
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ca1c7e7-eb2c-4a20-bfb9-f79dcbb931e5 · inbound
Bench-CoE: a Framework for Collaboration of Experts from Benchmark VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95d71df9-01de-4c31-9e4c-97a4b2c36310 · inbound
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 81ae5ea7-003b-4672-921e-440813bc8b78 · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
Reference 149
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bfbad0f2-be38-4f89-89c4-7478d982ec02 · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
Reference 177
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0a3e1d49-48b7-487e-b879-7080f6291384 · inbound
Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc1b745c-c1d3-4db7-b42a-d63899731d91 · inbound
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6670b17a-b19a-4e48-8a9f-883f1c63a54a · inbound
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
Reference 292
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e25db140-f90f-4773-b74d-5308c79c73fc · inbound
Qwen-Audio-VAE Technical Report VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c18abae-f539-486e-b13b-46c87adcc56a · inbound
ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.