Pith. sign in

Paper Citation Record · LEDGER

VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2406.06462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.06462 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:31:36.856649Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:43.988790Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7a856c60-8a1e-498a-be79-9ab924c70cf5 · inbound

CogVLM2: Visual Language Models for Image and Video Understanding cites this paper.

CogVLM2: Visual Language Models for Image and Video Understanding VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:10:27.827632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:3a1c6585fc2d3bde9ab53d08dc12850f288b3a031fff5ec3f2689c9ff5382ef0

Observation 5499c804-4a2a-4638-a85c-ae477b44d24e · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:36.856649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:36.856649Z digest=sha256:fe8f0110b44404405d3f60273ed51324ac74ea5adc6907b746295f7ffb6c96e5

Observation 4ca1c7e7-eb2c-4a20-bfb9-f79dcbb931e5 · inbound

Bench-CoE: a Framework for Collaboration of Experts from Benchmark cites this paper.

Bench-CoE: a Framework for Collaboration of Experts from Benchmark VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T21:46:44.880206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:46:44.880206Z digest=sha256:b02a1fa18eaf85f91148a24ffaad7e20f16964e6cab068d4327cc30ea621836e

Observation 95d71df9-01de-4c31-9e4c-97a4b2c36310 · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.236200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:fb2bb34ab80227fab70a4b17e1203aebf87b013d302a4c4051b1e20ec19d048d

Observation 81ae5ea7-003b-4672-921e-440813bc8b78 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.188674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:1a204c2f79cff2abd913ca13c623f6f7fcf7ac21b9bfed0abbe67ed2f82519a5

Observation bfbad0f2-be38-4f89-89c4-7478d982ec02 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.959246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:42bbbed666b37e0d4a213ac123c45a9019c3d95439cda8fd4e395daee91d83c6

Observation 0a3e1d49-48b7-487e-b879-7080f6291384 · inbound

Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving cites this paper.

Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T09:12:21.282651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:12:21.282651Z digest=sha256:db72d7fd484c5037fb51036de2afc7447cfff3ce70712fcbf795410b2e30db96

Observation dc1b745c-c1d3-4db7-b42a-d63899731d91 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.990226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:1b8cea6923bef95ae876dac051b19d9150a95f87f7ccf10bc0be7b7505376aaa

Observation 6670b17a-b19a-4e48-8a9f-883f1c63a54a · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 292

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:2657b3f3181af774b0f1fc6d18d9e16831eaf00f1a13abba5957d6a772c8aac6

Observation e25db140-f90f-4773-b74d-5308c79c73fc · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:16f43eebb30baa5490cc7759f3db98945aa21f1d0db25f00e0e73eb6d3d8c26b

Observation 3c18abae-f539-486e-b13b-46c87adcc56a · inbound

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs cites this paper.

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T04:16:08.341807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:16:08.341807Z digest=sha256:c3ed2c6d26498343ca4677a64aa7025ecad56de7e8de26e85dd3fe96944a3faf