Pith. sign in

Paper Citation Record · LEDGER

VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2502.12084.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.12084 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:56:56.496169Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:29:37.906841Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 46ce71c2-46a0-49a7-90f9-628b88aa3017 · inbound

Backdoor Cleaning without External Guidance in MLLM Fine-tuning cites this paper.

Backdoor Cleaning without External Guidance in MLLM Fine-tuning VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:56.496169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:56.496169Z digest=sha256:f30b056b8a4a89d0a9d296f74654910823645c57ba4a2dd44fc82a3d3ab4d8db

Observation 0e367bb4-2cbc-4d97-ac8b-e3719455b251 · inbound

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning cites this paper.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:22.722405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:22.722405Z digest=sha256:534c6c9ee8d7695af8ca437b6283ad8166b6682f2e1b0d5559aed235a916b429

Observation 6a3fa5ba-7c77-456d-bf5e-5c5c11ad4605 · inbound

EMCompress: Video-LLMs with Endomorphic Multimodal Compression cites this paper.

EMCompress: Video-LLMs with Endomorphic Multimodal Compression VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T20:41:50.602727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T20:40:42.995833Z digest=sha256:2dc3406dc531421373c69ec371df4bc4ec77249d8a0205676e2763e7a36494b6

Observation b8778d59-f835-4075-af7d-4c744c7679ef · inbound

TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-Tuning cites this paper.

TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-Tuning VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:35:28.676608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T07:33:15.865358Z digest=sha256:93ea49eb0f36f32c6d83b5b438d04cd0e10d7a7b19c1857d376ddb4e30018c85

Observation eb4e2b92-8766-45f4-a673-b6b4515e732e · inbound

ReMoT: Reinforcement Learning with Motion Contrast Triplets cites this paper.

ReMoT: Reinforcement Learning with Motion Contrast Triplets VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-02T19:58:28.853825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:58:28.853825Z digest=sha256:d7cab9db4eed37faca276991f4bec23270950c6e544e3c7a37cbf3a4f822656e

Observation 98632e51-9124-4271-ae73-f8de8729ddd7 · inbound

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding cites this paper.

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.416721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:26:01.568507Z digest=sha256:102369cb601ac84104d90079570635c44f4524b89718cc3cbea6bb1936615b36

Observation f9479c44-8e9e-4aee-8982-e69adddaa370 · inbound

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams cites this paper.

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:29:37.908516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T14:24:19.702588Z digest=sha256:143ea524d60792357e5022e44fc717b2da41d6dfa3de700125026635103f84fe

Observation 8b3092e7-3a28-49a8-8316-d5422de3b53d · inbound

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams cites this paper.

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:38.418104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:38.418104Z digest=sha256:f501a5dcc106c337bfc959dc609eaaa0c7fd73c443ddb826a85623edc5e809ab