Pith. sign in

Paper Citation Record · LEDGER

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning

As of 1 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2605.07106.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.07106 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T01:12:22.931098Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T04:55:06.999018Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact14
  • verified fuzzy13
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f30be200-80c3-4379-a9f3-a4352cdede76 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Chain-of-thought prompting elicits reasoning in large language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.373546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:a32726a4d9267f3cc4db45a36888b59b74e718908b24f2de585bc41c8c6ccd3e

Observation 2a2659ef-4497-472f-947e-1bc1c23f2b93 · outbound

This paper cites Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.383327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:33c69aa9bc18395683bfed0e6ccb370734cccb6a987ca4c82701c6eb2fca94ac

Observation b718ad6d-f801-4db6-aea0-66706690e3f9 · outbound

This paper cites Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.366423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:4676bdb0dff4d4bfba7283e4f5c1cfd9b579522c51933036e6e126ed6b2cb33c

Observation 19d474a0-ae56-435c-92a8-b4c16eccf2b3 · outbound

This paper cites Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:34:23.540456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:44e5bc7d079218ee44f2a4b7950f4e818566eae7655c30f6de197600867a700d

Observation 3adfb8e0-4d71-4e5f-b3a2-9fb6e0822618 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:17:58.977376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:b1ce1e7ff831a260956f2493bda4e34309cb5930fa8f926e6c6b3da6efe88251

Observation e5d4afc5-7a43-4bcf-9b08-e5f02efe6293 · outbound

This paper cites Llava-plus: Learning to use tools for creating multimodal agents.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Llava-plus: Learning to use tools for creating multimodal agents

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.445108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:76a3601325960a10d14882550de04b17ac4bd07f85532a75098831862149774e

Observation 730654a4-c88c-4cf8-9781-d7fc512a27e9 · outbound

This paper cites Jigsaw-r1: A study of rule-based visual reinforcement learning with jigsaw puzzles.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Jigsaw-r1: A study of rule-based visual reinforcement learning with jigsaw puzzles

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.640305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:14c0c9c8b7e33c6b0e38281a0992bfb2c2179bc8ab27476cb58db7b08baa62df

Observation 772fc35f-fb3f-4c72-8320-ed114569a506 · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Visual programming: Compositional visual reasoning without training

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.452265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:11cdba0e3b992a34f7546efb0ef09a296324dfb2430f5b1149c67a2d3f63b7f6

Observation 049a28f5-82cd-4c0b-9b02-2f215d44fe27 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Vipergpt: Visual inference via python execution for reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.468839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:4f52088bcfcc19129afd1a633bde2aeb8d01e14bf5354b87613bde98e9547f48

Observation fd48fd69-35e0-4bb2-804c-a5b2753e6e72 · outbound

This paper cites Visual program distillation: Distilling tools and program- matic reasoning into vision-language models.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Visual program distillation: Distilling tools and program- matic reasoning into vision-language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.437297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:762ffd83330faea9f3d9c70c7daa9a55adf8d963d88029f1c42d0d10abf251e3

Observation 29593fa8-184e-4cd6-a4a2-33cad8972be6 · outbound

This paper cites The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-08T02:03:55.088792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:6b114bd36de67453ab7e2b22f0419ecc6d28e9fbae3da647ab9be46cd631ae17

Observation 402a1fec-1c14-4069-919b-65d696137a6b · outbound

This paper cites Latent Visual Reasoning.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Latent Visual Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:41:30.500607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:7f3b74a29b21283341094b4fa19d9bb6d2a54ae06cec8b42d76cf8553a719e3a

Observation 5da0f996-d261-4397-a636-c57c874ce105 · outbound

This paper cites Monet: Reasoning in latent visual space beyond images and language.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Monet: Reasoning in latent visual space beyond images and language

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.550348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:1e3270589a9764069c0a21900d41a1ea3a3edf7cb07381ef7838d735157a9000

Observation 1f98e667-4a78-42f7-afa9-c586a6b07fe4 · outbound

This paper cites Imagination Helps Visual Reasoning, But Not Yet in Latent Space.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Imagination Helps Visual Reasoning, But Not Yet in Latent Space

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-09T02:06:14.979939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:76a308d1e040852e4d03362f024c53293aa4cc7c4b568f14bdbb8c679d70fa2f

Observation 151ac1cc-7389-4b8c-9675-c8f91395a498 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:9a9f84a1a83cf32643f2e7a542ed09594c5225e956aa72ba40ce7951ec2db99a

Observation 3ea7456d-c685-4ce1-bd30-b3387d2be9ef · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.866178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:f832b54ca1834192229386c387a4ce7012dbdf152c80ca1f832032199932b788

Observation 24f70aeb-2147-4bdd-98a5-b9771943db9d · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Training Large Language Models to Reason in a Continuous Latent Space

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:29:05.896693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:bcd71132193fa40f4b663aede3db7b7d7544f704d1810cce40c61d8bf627a85c

Observation e5910887-2f0f-4c79-8e06-dd73ae5c1bbc · outbound

This paper cites Codi: Com- pressing chain-of-thought into continuous space via self-distillation.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Codi: Com- pressing chain-of-thought into continuous space via self-distillation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.461386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:b68712c1b3e54342211e4f98979cd46e6a055c358f4ce97d7041efd2b310c8cf

Observation 4cbd606c-8f58-4820-8a63-acba136633d7 · outbound

This paper cites Fan, Y ., He, X., Yang, D., Zheng, K., Kuo, C.-C., Zheng, Y ., Narayanaraju, S.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Fan, Y ., He, X., Yang, D., Zheng, K., Kuo, C.-C., Zheng, Y ., Narayanaraju, S

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:35:59.524353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:94d012e9987ad29a73d27456931ad31e22c68ddbfda4959e348e27e01d3d8013

Observation b3257744-8fbc-4791-9ae0-3a7d3ed086b0 · outbound

This paper cites Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.653902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:a73cc78bd31d3c34512f8036189e5d33a34469862e8bc36d9a62c89428f41fb6

Observation ade04805-f270-47f2-9564-23bc3b7da21f · outbound

This paper cites Hudson and Christopher D.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Hudson and Christopher D

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.388478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:da9290d2d90ba76a2f3d229c0c3206e7456336ff124a4dfd1300940f9903495a

Observation 17493145-aec6-408f-828c-c56646fb4f58 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.609673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:cdd81609275ebfa338320a18ea96f3e300f342e902e4d840604ac594d56d635e

Observation 0ecea5ee-58eb-4429-967f-0d32834f2e75 · outbound

This paper cites Qwen2.5-vl technical report.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Qwen2.5-vl technical report

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.399164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:9a6ea80ea6f15737ce042bb55eeabaf9f941dd164b9edfe6cb84a7f20e414457

Observation 4fd9394f-8fca-4044-86be-7736e1e946ca · outbound

This paper cites Qwen2.5-VL Technical Report.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Qwen2.5-VL Technical Report

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:35:59.633185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:bb508edcc0d1842711948f864327e952f81f6873b57704f60bd49302c6eac6fc

Observation a309c390-f6e4-4aca-8955-d4c92ae5057b · outbound

This paper cites Emogen: Emotional image content generation with text-to-image diffusion models.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Emogen: Emotional image content generation with text-to-image diffusion models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:15:51.823830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:18d8b03105df594586a96b41aa6a6e617c7ba13aaf0e3795225e9e8c41d1a878

Observation 53566c6c-b41a-4d84-b6ff-701b49d4a4b9 · outbound

This paper cites Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.409051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:78c0799de7e9ad8faa9672868e5c048db19ed73b34aefc5762226280e0c04e23

Observation b2a722a3-2039-47e8-943d-ca67f1c46170 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.428672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:34c378a254347dd3ebbe643f76203c06a0e790a89b22078fdb55d5ac648f3473

Observation ab48df36-fd47-46a8-af7d-1414478f447a · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Blink: Multimodal large language models can see but not perceive

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T17:27:31.420059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:2c96052f214f0510bbc319f43065f1dbf5e8d83ba62eebc39b69929c89dac9a9

Observation dde8b3f4-8951-41de-841d-81380f1b3855 · outbound

This paper cites Chain-of-visual-thought: Teaching vlms to see and think better with continuous visual tokens.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Chain-of-visual-thought: Teaching vlms to see and think better with continuous visual tokens

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.662703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:9521873396ed4d5ee627f1c68a64072a6f19158c32704c10959c3bc5429a48b8

Pith citing papers

Observation 7a903fe6-ff2c-4282-b772-bc3acbc5e319 · inbound

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding cites this paper.

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T04:55:06.999018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T04:55:06.999018Z digest=sha256:1a8a3a816206a29ee052e03e4b75b1dcb47f32918bd3bf7b7052da923428b4db