Pith. sign in

Paper Citation Record · LEDGER

Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2403.12966.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.12966 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:57:40.838682Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:59:57.253374Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cd69af3a-c658-4de7-b693-ad706aefd14b · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.710895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:b4026f84c46308702648e8068a0af81f6db9ecaf2a051bf68895975609dff2f3

Observation cd23a595-bf5b-4d0f-8550-15b82f3f5b94 · inbound

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning cites this paper.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.081795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:fc31ffa34a8756800bd197a929a72003b1e9beda3b9e4cb4757c78a264ba8179

Observation a0d61b9f-e1be-4f8d-9c4f-fbb619126a18 · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:30.540242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:30.540242Z digest=sha256:e23beb1a14c6c7b1a32b2301d5f3af80f6cd330e3f5ce2d51054b4cac4bfff64

Observation fc6636ce-071f-46a3-91a6-23c9f979f512 · inbound

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding cites this paper.

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T10:49:58.803374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:49:58.803374Z digest=sha256:5a2bb5a622eeffec23f9a4e8c2b0c24a4d92c8be6511a4bfd890ecb6795edf6e

Observation 684c86e5-0d91-499a-85b5-0a00f9524f72 · inbound

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework cites this paper.

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T20:19:15.127415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:19:15.127415Z digest=sha256:140adffbe14cfa7dc1ec0ab71f0d48e71e5bf49ea4d0fc08f46bf72baefad39d

Observation f45d257b-0d11-45a7-813b-85ed3811a840 · inbound

Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking cites this paper.

Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T16:55:20.099628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:55:20.099628Z digest=sha256:a23cf2366d5d4b11084a9f8fc6d71fabfab9fe0cd91223ae8a3fbe2d569ec0c5

Observation 23e646b7-5992-4713-afa8-5cd784aa1b6c · inbound

Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs cites this paper.

Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:00.943697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T21:35:12.859669Z digest=sha256:766fa9fd0c0c3f8829597ccb3ed887dd770ce11164257857ad7975465e3140d1

Observation a961c15c-59e3-4194-9a80-6740ef13ffcc · inbound

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search cites this paper.

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.293323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T02:58:46.728868Z digest=sha256:8bec5a1db3a885a1439489f7a967b0fa2a6cf8765abbc0bae6b2c6fb3588fb8b

Observation 12dd8a4c-9c84-456d-94f0-cd1823df109c · inbound

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation cites this paper.

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:13.393013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T10:58:01.621488Z digest=sha256:004e58e6e617f6d38821a44d0c8a8c6663bce16fcd080fc69283ea90cc77f27f

Observation bbef1c4a-05ba-41c9-84bd-d1f8555e4ab0 · inbound

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation cites this paper.

Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.704886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:28:21.605646Z digest=sha256:7c90fd843fbb37ff142d34b65191a7878be51c64588ff3f75e4354e272b674a3

Observation 02bf47d2-71cf-4189-acf2-9dc74a589aea · inbound

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning cites this paper.

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:57.800982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T01:23:40.564561Z digest=sha256:2fe4197ca2a2a6e7ff2b8f265affe537aa5ec9d2e386930bf832b8057732e1f3

Observation 6fec5a22-b305-436b-9315-e96dc9f8b8d2 · inbound

An LMM for Precisely Grounding Elements in Documents cites this paper.

An LMM for Precisely Grounding Elements in Documents Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.732696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T01:51:50.043216Z digest=sha256:f15772cbc4e467d15e53eb7123065262171406989f1f5ab37474d4207eea4372

Observation b6d50c90-17fe-415b-8956-ed10cf4791f2 · inbound

ActiveScope: Actively Seeking and Correcting Perception for MLLMs cites this paper.

ActiveScope: Actively Seeking and Correcting Perception for MLLMs Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:59:57.255666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T01:06:40.481169Z digest=sha256:c1270b7cf6434b7c7b0679f587d7d4fc5492a005c10a6f7522dd0fe557706d9a

Observation b4d01059-1aac-4bd7-882c-e1b6a0ef5325 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.117287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:7c3e452c291b6f904fce519e967abd225af095a1b3ab0168a3af6965afccc3dc

Observation 4cf2e4d6-aaca-4331-a20f-c6aba144b36b · inbound

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning cites this paper.

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:26:58.461700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-02T13:24:17.538850Z digest=sha256:2399640cdcc4db36976c8e8e05bf6860a07ab0413b060b6201440b2b8c8f1e2d

Observation 4a13e359-7a5d-4ebf-ba38-dea51b07c98f · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 210

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:7133bca4eebd4138db1ddc5f88d9fb50e19501f3f13cf3599ca6567e7372216b

Observation 125a6848-5382-4caa-adf9-477492677fed · inbound

Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs cites this paper.

Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T06:18:15.882601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:18:15.882601Z digest=sha256:a205a023fbd7c048cc245817d62be3b6a4076523e464f90e1f1e90f194b01ea3

Observation 43170df6-d614-4c33-94f2-bfce8ce52198 · inbound

Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models cites this paper.

Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T11:57:40.838682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:57:40.838682Z digest=sha256:85c15082fff9d24c9dfcf768837235a77962cb36860d02e73211aa70fe85f900