Pith. sign in

Paper Citation Record · LEDGER

VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2406.06462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.06462 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:31:36.856649Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:43.988790Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7a856c60-8a1e-498a-be79-9ab924c70cf5 · inbound

CogVLM2: Visual Language Models for Image and Video Understanding cites this paper.

CogVLM2: Visual Language Models for Image and Video Understanding VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:10:27.827632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:71c50b67ea554519756d05f7a501034c59bebf22a06ee9a6ab9965a4af1cda99

Observation 5499c804-4a2a-4638-a85c-ae477b44d24e · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:36.856649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:36.856649Z digest=sha256:7a38c2ce3addacc9dd4be97e70bbb4177d24d5bec4486102901b3c45b312e2cb

Observation 4ca1c7e7-eb2c-4a20-bfb9-f79dcbb931e5 · inbound

Bench-CoE: a Framework for Collaboration of Experts from Benchmark cites this paper.

Bench-CoE: a Framework for Collaboration of Experts from Benchmark VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T21:46:44.880206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:46:44.880206Z digest=sha256:9d34aff01a8477917e457a523e09d9519aec1814e08835302ee981f92df2687e

Observation 95d71df9-01de-4c31-9e4c-97a4b2c36310 · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.236200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:d70afccffbb2bafa861a30f9c9b018227ec96d024738b8d78860e07dc441d521

Observation 81ae5ea7-003b-4672-921e-440813bc8b78 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.188674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:cb9c9246e50555e28705269f2b8dae84ed821d163ac6bbc90943b35992b660e8

Observation bfbad0f2-be38-4f89-89c4-7478d982ec02 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.959246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:58760e56059f3d1bb1756c2efde2461f5ff7f55ebc021c545b9813a8240c3c58

Observation 0a3e1d49-48b7-487e-b879-7080f6291384 · inbound

Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving cites this paper.

Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T09:12:21.282651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:12:21.282651Z digest=sha256:1148632972de73eac6c44e6ca718650926572cd2f1fcca40934eecdfcb923ad4

Observation dc1b745c-c1d3-4db7-b42a-d63899731d91 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.990226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:9719eae15af14c891b8673183c9e6f361858015c20e9f356280e3d3133803e04

Observation 6670b17a-b19a-4e48-8a9f-883f1c63a54a · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 292

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:f42aa3ccbd542120db02441e56f61121698f4a46ffba8316f64a7118bf2b4a0d

Observation e25db140-f90f-4773-b74d-5308c79c73fc · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:8e9ba28533ce958d6eb11d964120f4ee1752513d15bdcfe54e8a7a34abbb4466

Observation 3c18abae-f539-486e-b13b-46c87adcc56a · inbound

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs cites this paper.

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T04:16:08.341807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:16:08.341807Z digest=sha256:c061e51ee63dffd0d921b03fcfdcce5ab92a84607c9808c5a526401d6afaf279