Pith. sign in

Paper Citation Record · LEDGER

Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2511.19418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.19418 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:16:06.189513Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T13:37:06.883595Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 415c3fb6-0612-4776-906d-dbc5242dc3cb · inbound

S$^2$GR: Stepwise Semantic-Guided Reasoning in Latent Space for Generative Recommendation cites this paper.

S$^2$GR: Stepwise Semantic-Guided Reasoning in Latent Space for Generative Recommendation Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T14:23:44.856958Z digest=sha256:f3c46251f914b59c9ad69c8479479357865f75d848214352e9e753967470da36

Observation 024008d7-3f12-4e05-887c-3a826b5cf2c7 · inbound

Vision-aligned Latent Reasoning for Multi-modal Large Language Model cites this paper.

Vision-aligned Latent Reasoning for Multi-modal Large Language Model Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T07:48:23.272002Z digest=sha256:c2c931d39c4d9e4e6ddc14396ba803e4c6dfe98cdc57059f1d0dcbbec71564fe

Observation 66860c5b-fb59-4d77-a478-adfcb95ec9b7 · inbound

PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning cites this paper.

PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T20:05:30.324309Z digest=sha256:a261288171d0794b9293205ab045f39c43766657675ed3a8aed1622a8be3650e

Observation fe43b6ab-f6b0-41d2-9969-3d02a9f9dab1 · inbound

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators cites this paper.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:f9fc046511326ec6b010d81d2061859e458ee6960b346281dbdb292f5edc54e0

Observation 332ed0d5-66b7-4381-932c-598029270170 · inbound

Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models cites this paper.

Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T00:16:34.362000Z digest=sha256:b93ab2a3abb61cf65ab9b35e7808b1ed0dd1e9b207aedae7bd5bd441b7bf8a27

Observation 2230289f-a68f-4432-a02f-1679240ba3a0 · inbound

Agentic AI for Remote Sensing: Technical Challenges and Research Directions cites this paper.

Agentic AI for Remote Sensing: Technical Challenges and Research Directions Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T04:29:22.477531Z digest=sha256:b4a2436d9c9243051ae54f0e53f6445f8f4a80f45ee7864ebe0a52a8ab167bb5

Observation 828b4f3c-b286-43f2-8f83-e45b4e1f1f49 · inbound

Agentic AI for Remote Sensing: Technical Challenges and Research Directions cites this paper.

Agentic AI for Remote Sensing: Technical Challenges and Research Directions Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T20:55:38.841743Z digest=sha256:77d7b53f5ec06a24cec03fdb85c337c45153d4c1c8ae8fbe5f932132b627b558

Observation 1e42d526-7246-4f84-8d8b-b40e242b3769 · inbound

Measuring AI Reasoning: A Guide for Researchers cites this paper.

Measuring AI Reasoning: A Guide for Researchers Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 114

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T18:53:18.586923Z digest=sha256:9fc96d6f216ca555eee6bafe634b2c13c7f1e1af773b9261aa1fb854fae3c702

Observation 29a6f4c4-47ec-4893-b174-e1fb73bfccf8 · inbound

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs cites this paper.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:f721a71718e800582c4126451afbe27e53a8aab948371989ccc245a4299de233

Observation dde8b3f4-8951-41de-841d-81380f1b3855 · inbound

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning cites this paper.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:5c1febeb3728debc9a9cbe47db384f5ead51ec7362e7f2dc01be015aa61da2d0

Observation bd6c5859-2c73-4cd6-9052-9b8f2382596b · inbound

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization cites this paper.

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T01:19:04.590978Z digest=sha256:9d4d00d5a5d13c0b7373bf7bd21ba7d421c9a74094a4ee5e426c67636065eacd

Observation 8f687679-4021-44a3-9f30-f15f0c604874 · inbound

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization cites this paper.

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T07:08:52.879854Z digest=sha256:8206a157e383d86a95249b87b5e56daada9ce67141285b0871bfaf41c122823d

Observation 79502fa9-73f8-462b-a20f-e29db0484ff3 · inbound

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning cites this paper.

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T02:24:12.349405Z digest=sha256:df3e3b39941415a79cec66b91252f90e01f2e8ea021a4b272795dad6ff98a3e7

Observation 72477733-2df0-4a69-8be5-372fadc98951 · inbound

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs cites this paper.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:9306c9fbcc931975b7982a1cb9ff6fb6fbccfe3586627ec496c48288b14dbbc8

Observation aac11c40-288b-40d2-9883-86ff09096d21 · inbound

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models cites this paper.

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T07:21:40.803291Z digest=sha256:fd8d5091a581e569ba97c9c4cb0c3a3799a93182072f6910ee114d2a2c14c249

Observation 5bece33e-239d-49a6-a35e-1810cf6ba7d5 · inbound

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models cites this paper.

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-19T16:48:20.132240Z digest=sha256:39e4af00ac0a4a55eb73257b2b1cc85ffd01d9ced0f191c18ad743d7279b7f63

Observation 4b889f7c-f74e-4ca4-937c-f06a49d481b9 · inbound

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models cites this paper.

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-21T08:02:00.980836Z digest=sha256:8d5b7f1b34017cb52f79c5358d22c54b9de923e5cfd99fcebcb6d0bcfab82aec

Observation 2783319f-9c3d-49f6-b13e-65d03832f5f9 · inbound

Leveraging Latent Visual Reasoning in Silence cites this paper.

Leveraging Latent Visual Reasoning in Silence Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T10:27:13.859216Z digest=sha256:3ea170905f69cd93137ef0e11fca18f699b9973ab080f3e44975f29659f3d0d2

Observation 569bb120-c6ce-4032-aca6-33bc406ff07b · inbound

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning cites this paper.

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T06:34:57.483234Z digest=sha256:8c33aa9cae4fb70f58a23206c47bf000582019aa242273a20d498a4b7d092663

Observation f2f93d21-680e-48fa-bf2f-d22a1295e56c · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:e7cfe3a01afe986e5b8e65bf7fb3af781592ae4490af4327a094120f6efa31b0

Observation fd9d4cd1-b958-4a87-bc8c-c0fb0bc8b837 · inbound

Stateful Visual Encoders for Vision-Language Models cites this paper.

Stateful Visual Encoders for Vision-Language Models Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T07:26:01.601557Z digest=sha256:4df2f253e36624e972c370ccd9d420886f4f777f9dea5b4ffd72e067b6636118

Observation e8a03cbc-b5f1-4887-bdbf-4dae7a60d7bc · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:66299cd46288774c39cdf78607699fcab6e0d775b19b3037e927d917627b6947

Observation 2dc9067c-81bc-4c19-96c2-6032e97c53e6 · inbound

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction cites this paper.

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T02:46:50.373450Z digest=sha256:f8cee93b9163ada1aa9b1a2ac3d98ca53ae846f6f32b4fd9e2ea8359e569b475

Observation 7ca19a15-9184-420b-9c50-ce3bff914cb1 · inbound

Continuous Audio Thinking for Large Audio Language Models cites this paper.

Continuous Audio Thinking for Large Audio Language Models Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T21:52:49.839901Z digest=sha256:823464cec1679963a0ed0aef85fd8794cfa32516baaa7be5262031dd1f81c004

Observation dd3366ce-45d7-4e7a-885f-b9472943a6d0 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:cfcd25bb98190e3c6c704407b4fbc07f9a836b7a378fd1fb3a52491d54bb5ceb

Observation a470d566-ba39-4dd5-acc0-9c6f6dbbac73 · inbound

Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models cites this paper.

Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T06:35:16.868238Z digest=sha256:22406684672dd5a136c6a2562b1bc5a9b84e1711c267b78ea8adbeed7e77186b

Observation 55fc931d-9a60-4fab-b378-0040574e6bbb · inbound

APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts cites this paper.

APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:01:32.731977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T13:33:17.574108Z digest=sha256:2596ed281e1eef09983a0a29d0b833ec4bdee11163e366f018bffeebe6f2e050

Observation 44272ef2-07fa-4948-8519-60251738e985 · inbound

Evidence-RL: Towards Evidence-intensive Visual Reasoning cites this paper.

Evidence-RL: Towards Evidence-intensive Visual Reasoning Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-12T00:39:05.844143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:39:05.844143Z digest=sha256:3481a57c88addc7f9672b3e261c7be6c2c48f8e8af381183e898d90dd5906e5d

Observation d1c4922b-b9c9-4a04-b667-3d21bc76b694 · inbound

Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models cites this paper.

Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:06.189513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:16:06.189513Z digest=sha256:8d0b641dbde76349b90b3f150781b78ca307d2d2e30076c06d235ff9175cec23