Pith. sign in

Paper Citation Record · LEDGER

See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2301.05226.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.05226 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:55.800954Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T22:42:13.506064Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f804bb78-8fa3-4b9b-918e-eff6e51e1459 · inbound

Multimodal Chain-of-Thought Reasoning in Language Models cites this paper.

Multimodal Chain-of-Thought Reasoning in Language Models See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.534776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:8d409f1ba5b4ee2b967e18e72fcf3b90edf21faa69d135c8a34a55da17f4bb16

Observation fdbd6c0a-404a-4908-9f9f-0152b88d2277 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.682387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:260009a0c3a596ba23579e13a797cc3373576d28a2b287c911d4b3b669983250

Observation 0cc0e3a7-0202-4161-86e2-959667e1bc7e · inbound

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? cites this paper.

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:42:13.509384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T22:38:35.969273Z digest=sha256:8ec99a62b83b0f05c2c95196b8beb12667ec4945ee17986429ef5eadb37a0287

Observation 431d7a65-b4a8-4db9-8a57-5d0e8b9fee88 · inbound

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs cites this paper.

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:55.800954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:55.800954Z digest=sha256:3a6779d4dc419467a922037ace676d1d09702ab639b5f2bca22ac28c3906e6a5

Observation 6eee8200-8caf-4adc-98e2-211589b90ad5 · inbound

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? cites this paper.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:31.881851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:31.881851Z digest=sha256:53a70af8ad0421043bf7d08dc125783301cce9726664eb3c09122ad55c8f8228

Observation f50ac4c8-1124-4a79-b6c8-e32f1c9947e9 · inbound

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning cites this paper.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.841863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.841863Z digest=sha256:4596fe8a9e23496b59614cbb7f94422e39d321ce275ba67d3addf8585c3462a2

Observation e53b242c-3402-4e09-9564-67467c7bba45 · inbound

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning cites this paper.

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:02:52.625798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T23:02:14.913189Z digest=sha256:6b3001cab27a64227ee0f8441f9dd1ddbe5b56fcc007e3cc2ee560e0ba771bbd

Observation da99e7ad-f238-43cd-86f2-4d729a8d488d · inbound

LaRe: Latent Refocusing for Multimodal Reasoning cites this paper.

LaRe: Latent Refocusing for Multimodal Reasoning See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T00:15:15.141748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:15:15.141748Z digest=sha256:fb0b450414c0a4ef7afe8c74ca28837d74ba4ef0327fd60a91c0fb10d54cf863

Observation 535c2719-7ef5-4792-bad4-758da2518608 · inbound

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models cites this paper.

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T18:25:01.504726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:25:01.504726Z digest=sha256:86b9e2d6ba958b88b1e834e5554d93048f4e391b584122ad174eba38c2de3503