Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T01:12:22.931098Z
Paper Citation Record · LEDGER
As of 1 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2605.07106.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T01:12:22.931098Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T04:55:06.999018Z
A source-named dated measurement, never combined with another source.
Source: cited_works
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f30be200-80c3-4379-a9f3-a4352cdede76 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Chain-of-thought prompting elicits reasoning in large language models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2a2659ef-4497-472f-947e-1bc1c23f2b93 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b718ad6d-f801-4db6-aea0-66706690e3f9 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 19d474a0-ae56-435c-92a8-b4c16eccf2b3 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3adfb8e0-4d71-4e5f-b3a2-9fb6e0822618 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e5d4afc5-7a43-4bcf-9b08-e5f02efe6293 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Llava-plus: Learning to use tools for creating multimodal agents
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 730654a4-c88c-4cf8-9781-d7fc512a27e9 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Jigsaw-r1: A study of rule-based visual reinforcement learning with jigsaw puzzles
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 772fc35f-fb3f-4c72-8320-ed114569a506 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Visual programming: Compositional visual reasoning without training
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 049a28f5-82cd-4c0b-9b02-2f215d44fe27 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Vipergpt: Visual inference via python execution for reasoning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation fd48fd69-35e0-4bb2-804c-a5b2753e6e72 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Visual program distillation: Distilling tools and program- matic reasoning into vision-language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 29593fa8-184e-4cd6-a4a2-33cad8972be6 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 402a1fec-1c14-4069-919b-65d696137a6b · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Latent Visual Reasoning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 5da0f996-d261-4397-a636-c57c874ce105 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Monet: Reasoning in latent visual space beyond images and language
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1f98e667-4a78-42f7-afa9-c586a6b07fe4 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Imagination Helps Visual Reasoning, But Not Yet in Latent Space
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 151ac1cc-7389-4b8c-9675-c8f91395a498 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Multimodal Chain-of-Thought Reasoning in Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3ea7456d-c685-4ce1-bd30-b3387d2be9ef · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 24f70aeb-2147-4bdd-98a5-b9771943db9d · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Training Large Language Models to Reason in a Continuous Latent Space
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e5910887-2f0f-4c79-8e06-dd73ae5c1bbc · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Codi: Com- pressing chain-of-thought into continuous space via self-distillation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 4cbd606c-8f58-4820-8a63-acba136633d7 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Fan, Y ., He, X., Yang, D., Zheng, K., Kuo, C.-C., Zheng, Y ., Narayanaraju, S
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b3257744-8fbc-4791-9ae0-3a7d3ed086b0 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation ade04805-f270-47f2-9564-23bc3b7da21f · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Hudson and Christopher D
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 17493145-aec6-408f-828c-c56646fb4f58 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0ecea5ee-58eb-4429-967f-0d32834f2e75 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Qwen2.5-vl technical report
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 4fd9394f-8fca-4044-86be-7736e1e946ca · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Qwen2.5-VL Technical Report
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a309c390-f6e4-4aca-8955-d4c92ae5057b · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Emogen: Emotional image content generation with text-to-image diffusion models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 53566c6c-b41a-4d84-b6ff-701b49d4a4b9 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b2a722a3-2039-47e8-943d-ca67f1c46170 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation ab48df36-fd47-46a8-af7d-1414478f447a · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Blink: Multimodal large language models can see but not perceive
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation dde8b3f4-8951-41de-841d-81380f1b3855 · outbound
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Chain-of-visual-thought: Teaching vlms to see and think better with continuous visual tokens
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 7a903fe6-ff2c-4282-b772-bc3acbc5e319 · inbound
Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.