Pith. sign in

Paper Citation Record · LEDGER

Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2511.19418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.19418 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:16:06.189513Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T13:37:06.883595Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 415c3fb6-0612-4776-906d-dbc5242dc3cb · inbound

S$^2$GR: Stepwise Semantic-Guided Reasoning in Latent Space for Generative Recommendation cites this paper.

S$^2$GR: Stepwise Semantic-Guided Reasoning in Latent Space for Generative Recommendation Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T14:23:44.856958Z digest=sha256:3ad0c018ad47ce74298d8f2bbf6b2259f8b0ec91c7566d28542a5f7256b609b3

Observation 024008d7-3f12-4e05-887c-3a826b5cf2c7 · inbound

Vision-aligned Latent Reasoning for Multi-modal Large Language Model cites this paper.

Vision-aligned Latent Reasoning for Multi-modal Large Language Model Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T07:48:23.272002Z digest=sha256:5e66a8b441379804c347136cd90e10e8d46e7891bddb33d3d1c8729f7451c659

Observation 66860c5b-fb59-4d77-a478-adfcb95ec9b7 · inbound

PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning cites this paper.

PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T20:05:30.324309Z digest=sha256:deab1c0fbf65e66445bdb555ffa5ca5aac6f05df3579d070223b1d52568c60b6

Observation fe43b6ab-f6b0-41d2-9969-3d02a9f9dab1 · inbound

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators cites this paper.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:76e21882447c2ac2e340de5a461b240881f31cdf3414ad686b0ac31e18bd2584

Observation 332ed0d5-66b7-4381-932c-598029270170 · inbound

Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models cites this paper.

Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T00:16:34.362000Z digest=sha256:eea2463aa36745dce908d1b4f217383d4a8ed775063ce14bb40ac0d5e962262f

Observation 2230289f-a68f-4432-a02f-1679240ba3a0 · inbound

Agentic AI for Remote Sensing: Technical Challenges and Research Directions cites this paper.

Agentic AI for Remote Sensing: Technical Challenges and Research Directions Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T04:29:22.477531Z digest=sha256:24367e52f6af7791c972ffee18299b6fcacabe9aff40e9a93962ff199364d8a4

Observation 828b4f3c-b286-43f2-8f83-e45b4e1f1f49 · inbound

Agentic AI for Remote Sensing: Technical Challenges and Research Directions cites this paper.

Agentic AI for Remote Sensing: Technical Challenges and Research Directions Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T20:55:38.841743Z digest=sha256:8d6518b022d78e06ca9c7968955e86949be59c0c12de40ba0861bae55e4eaa24

Observation 1e42d526-7246-4f84-8d8b-b40e242b3769 · inbound

Measuring AI Reasoning: A Guide for Researchers cites this paper.

Measuring AI Reasoning: A Guide for Researchers Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 114

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T18:53:18.586923Z digest=sha256:d24c1af6568c4ca283aa58e58c4111964dbe66d27d9b29bd61dc5ca2901e894c

Observation 29a6f4c4-47ec-4893-b174-e1fb73bfccf8 · inbound

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs cites this paper.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:42a1e20a17a65656b9ac8956c0451e32d8c72ad20c94f5c57404f80beaddffda

Observation dde8b3f4-8951-41de-841d-81380f1b3855 · inbound

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning cites this paper.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:e514827d2f384e4ec363de319d76197a775c7190c72a8dd50538c82a58ef63c5

Observation bd6c5859-2c73-4cd6-9052-9b8f2382596b · inbound

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization cites this paper.

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T01:19:04.590978Z digest=sha256:8fb83369048f1ddea1814d0d48a104e5ef7422fb105845f0b85474eca88e0862

Observation 8f687679-4021-44a3-9f30-f15f0c604874 · inbound

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization cites this paper.

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T07:08:52.879854Z digest=sha256:a08f393780b1dba4ae030d8b5294e430a0e5b207beed9aa2307e054fbea2d524

Observation 79502fa9-73f8-462b-a20f-e29db0484ff3 · inbound

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning cites this paper.

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T02:24:12.349405Z digest=sha256:b35f3de470455d81580e634104dfe8bc269666af1b3317c63ce8d324bc819e7f

Observation 72477733-2df0-4a69-8be5-372fadc98951 · inbound

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs cites this paper.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:8772a69ae27bf242e2a60d30d3e0b0d0d413f8fbf4905973ec30d4835c2f6097

Observation aac11c40-288b-40d2-9883-86ff09096d21 · inbound

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models cites this paper.

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T07:21:40.803291Z digest=sha256:a041485b5872aa8bf6c2c32ecca747abf437c7d6c0babcdc5e1616c200345cee

Observation 5bece33e-239d-49a6-a35e-1810cf6ba7d5 · inbound

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models cites this paper.

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-19T16:48:20.132240Z digest=sha256:ba3fe3975e2665fc24fe7abd1bed4774cb738b43f6ec00d5f714ca7453764c25

Observation 4b889f7c-f74e-4ca4-937c-f06a49d481b9 · inbound

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models cites this paper.

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-21T08:02:00.980836Z digest=sha256:365eb53331c350a9c60e798f9c6dd84078d1d3117e9d680e8658a3e18ac6a79b

Observation 2783319f-9c3d-49f6-b13e-65d03832f5f9 · inbound

Leveraging Latent Visual Reasoning in Silence cites this paper.

Leveraging Latent Visual Reasoning in Silence Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:27:13.859216Z digest=sha256:f403c26af366af2e8d5b1dc52a023b808ef1c24f60ffbad949d26496e54f98b4

Observation 569bb120-c6ce-4032-aca6-33bc406ff07b · inbound

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning cites this paper.

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T06:34:57.483234Z digest=sha256:e16a654729e2318a326ce7d566b85ebd9e6c6a69a209a30a56bd5b59ffc197ea

Observation f2f93d21-680e-48fa-bf2f-d22a1295e56c · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:28c55e31991dcf67b604b4a6f6fade828d24e32a1bbceddc6a62558477a04e96

Observation fd9d4cd1-b958-4a87-bc8c-c0fb0bc8b837 · inbound

Stateful Visual Encoders for Vision-Language Models cites this paper.

Stateful Visual Encoders for Vision-Language Models Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T07:26:01.601557Z digest=sha256:2364d449d3dc00b4b1165290f27f5400b234b803c886ed70e695b32a37c8bd27

Observation e8a03cbc-b5f1-4887-bdbf-4dae7a60d7bc · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:1d5ca7d1de9a3e73cb49b46ad8534bab0cc3cb4061bc929b11a994e26b024595

Observation 2dc9067c-81bc-4c19-96c2-6032e97c53e6 · inbound

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction cites this paper.

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T02:46:50.373450Z digest=sha256:7cd80806d8162b32c6197a998ae15dd8fa8430aa66b863a5a87f58556c50ff6a

Observation 7ca19a15-9184-420b-9c50-ce3bff914cb1 · inbound

Continuous Audio Thinking for Large Audio Language Models cites this paper.

Continuous Audio Thinking for Large Audio Language Models Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T21:52:49.839901Z digest=sha256:3ed48147e1a09ed537a59a1e4c30063a5f8bc721219ec72eaa9319f29781ea94

Observation dd3366ce-45d7-4e7a-885f-b9472943a6d0 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:cfb4f8b7ac29288d52fcb0cb16fbae1c8927dd420bfba0f1d0501ffdea5611c6

Observation a470d566-ba39-4dd5-acc0-9c6f6dbbac73 · inbound

Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models cites this paper.

Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T06:35:16.868238Z digest=sha256:37adf91a5faf97ff2c17fddfc2c8767247e0d392bae64209000afb72b1986e67

Observation 55fc931d-9a60-4fab-b378-0040574e6bbb · inbound

APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts cites this paper.

APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T13:33:17.574108Z digest=sha256:8df7924e8f9626d61a8296dc541985eeb8f92d9615077f1a13589474856f03cd

Observation 44272ef2-07fa-4948-8519-60251738e985 · inbound

Evidence-RL: Towards Evidence-intensive Visual Reasoning cites this paper.

Evidence-RL: Towards Evidence-intensive Visual Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-12T00:39:05.844143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:39:05.844143Z digest=sha256:3481a57c88addc7f9672b3e261c7be6c2c48f8e8af381183e898d90dd5906e5d

Observation d1c4922b-b9c9-4a04-b667-3d21bc76b694 · inbound

Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models cites this paper.

Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:06.189513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:16:06.189513Z digest=sha256:8d0b641dbde76349b90b3f150781b78ca307d2d2e30076c06d235ff9175cec23