Pith. sign in

Paper Citation Record · LEDGER

Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2501.03230.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03230 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:18:57.344627Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T08:57:47.799001Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2347e678-cf70-4e00-b626-809cef25365c · inbound

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models cites this paper.

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:15.540187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:36:15.540187Z digest=sha256:d47ff9c159c86f1bf91d340a5ab1326dacb712b23f84ea9fb2435985d00bc4de

Observation 9675d94b-f65d-4fbb-a217-82b0059f9647 · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.838095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.838095Z digest=sha256:f2b867bc1ad19662e05b7704bf27cfd2eee2b610dba9521303ee02a2954cf882

Observation 90ecb6a8-ce9f-4510-9908-270fc34ff20f · inbound

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? cites this paper.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.061606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.061606Z digest=sha256:a4d09d603c558e2b5a78273ca4e8fef63c4cfacdeb5d35c9cddfb7fced725890

Observation 480c83b2-d71e-49d9-a101-d9b45129049b · inbound

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? cites this paper.

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:57.344627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:57.344627Z digest=sha256:360c08d60d76946c873143827df637b806dc5d02bb38f075a0756c35bdebe7b3

Observation 2ce8fd4d-5cef-4121-8368-d09de60cdb0a · inbound

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning cites this paper.

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:53.934003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:53.934003Z digest=sha256:111bbd30098277ead3ebf3674f45c153db78e6d4343e5295aef91d9dfc85e1d6

Observation 01f3cec6-94bc-4467-b9a8-00f90c6b63fb · inbound

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos cites this paper.

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:17.320440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:17.320440Z digest=sha256:c28d4ba19c258b358e34ce98d94aeb4af677ccdf7d315399b00d97de99a2dee5

Observation 4359e34a-b13d-4732-aa9a-1f327d90bb74 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 155

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:57.362974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:57.362974Z digest=sha256:b0b3f4220af9a26a0805d1b73b42e04ae0dcc06691fde063d67f8f97b5891da3

Observation f57066a7-b0b7-4587-9c3a-b9919b65d70d · inbound

ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval cites this paper.

ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T16:05:18.612817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:05:18.612817Z digest=sha256:647d4d775d05690cb69b60810a06c5b3472700f45a39bec4562f4819da7ed4e8

Observation 06ccf859-2cfb-4de3-bcfe-fe36e36c6fe1 · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:36.704406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:36.704406Z digest=sha256:f9772019a7b614c7d8c4d261812c8c36afad4c10784851a669e5382486b19219

Observation 0cab05a2-37db-4e22-9ab5-6e1e9e8df3d7 · inbound

LaRe: Latent Refocusing for Multimodal Reasoning cites this paper.

LaRe: Latent Refocusing for Multimodal Reasoning Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T00:15:15.354915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:15:15.354915Z digest=sha256:a17575b5e3f2b1e3259902202a6f954fc0ecca2dbc2bd7b4672690ed01f3151d

Observation ffab02f4-3f31-4207-9b4f-1da7cea574b9 · inbound

SCP: Spatial Causal Prediction in Video cites this paper.

SCP: Spatial Causal Prediction in Video Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:50:11.257532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T16:47:44.523606Z digest=sha256:a2e7ee54f405c261be6fd347c710207d6e4a46cce5cdad6f1381cf64f8bec19a

Observation 0fdad807-aa9f-4112-b2a9-f42f8e0db055 · inbound

EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models cites this paper.

EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T22:25:04.080629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:25:04.080629Z digest=sha256:23c8e15f5d3fe9389c843285c431e49ea64e4e9d48d609320acae9ebafd54ecc

Observation 0d6f5247-243e-4e29-b227-e69e78685e10 · inbound

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging cites this paper.

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:05.690356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:42:43.948462Z digest=sha256:c1f177b084f6a9aaa6982a46121b009405092b61c50275d9aa79f5e6d9187c4a

Observation 6d566b1f-c33e-478e-be43-1413a3faa1dd · inbound

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding cites this paper.

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:05:22.155788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T12:03:09.408019Z digest=sha256:fe8ac5f2a3571e96d6b3dcfaa426b344a2cee1b21fd58de9bbf65bc3c93ff629

Observation e74bd42d-a812-4456-a4f0-cd6f30bd7042 · inbound

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding cites this paper.

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:37:41.741853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T17:34:10.344111Z digest=sha256:40a5b36de18750b15c5823dc5d7451e3a29505db6563319979db442986c5476a

Observation ab8ee26b-d76d-4004-a8de-38e451d00b96 · inbound

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models cites this paper.

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:46:37.539454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T06:41:59.641410Z digest=sha256:9a56e783fc4bc45c117c206156c88ab908d7f0572b1e002831cf315dfee07f63

Observation aecfa82f-bb1c-4e17-b1ea-9835a388b5e2 · inbound

Act2See: Emergent Active Visual Perception for Video Reasoning cites this paper.

Act2See: Emergent Active Visual Perception for Video Reasoning Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:45:22.936174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T19:34:53.683729Z digest=sha256:0b1e6de540bbf5484333db34a9fcb6f7620672f082df0cb956c8e1933adc8204

Observation 09ead882-9a5b-48be-b467-633f37849428 · inbound

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models cites this paper.

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:35.770004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:00:23.681682Z digest=sha256:001a930d66c182c0707ad9c6dce946c09eb2f50d3e10955db4507e18beb0d656

Observation bbb98392-7b8e-4250-bac0-197f37dd3781 · inbound

AVTrack: Audio-Visual Tracking in Human-centric Complex Scenes cites this paper.

AVTrack: Audio-Visual Tracking in Human-centric Complex Scenes Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:20.158878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T15:00:41.632699Z digest=sha256:66d36e4116c11634e580b19590383f931508c060e6bd195efd96c79bce3734f0

Observation db96b895-cc57-46af-b777-73f535fa5fc5 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.742811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:84be35b8cb4d012f7f21f8e334099a86b2bfe52a731c838ad8019125b40e5d89

Observation d3fa565a-c03c-4bc6-82f0-7c6b65c5f10a · inbound

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA cites this paper.

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:30.197957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:55:35.743040Z digest=sha256:5567bdf42bd5d74ded9bea9d7efd945a52a930ce1e426d0fcff906458d1b47e3

Observation cce4ce6e-70f5-48c3-bd5d-1e146ed7bfbf · inbound

Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding cites this paper.

Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:57:47.800558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T10:37:52.186246Z digest=sha256:ae1a1f9550472bd950eef0ff0be65dfd7edd78340f6c4647c3db92bc06ad4dd1