Pith. sign in

Paper Citation Record · LEDGER

Semantic-Enriched Latent Visual Reasoning

As of 16 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2605.19342.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19342 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T18:48:04.370230Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact16
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2d0d8c2c-85fc-4243-8bb4-ab2794584cac · outbound

This paper cites Qwen Technical Report.

Semantic-Enriched Latent Visual Reasoning Qwen Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.802055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:5b54fbccb2c9302f851870b5769e653b98af4df618d35578f0f32cb5ba6b58fd

Observation 5b0361da-21aa-41f1-b420-ca32b8214f6b · outbound

This paper cites Qwen3-VL Technical Report.

Semantic-Enriched Latent Visual Reasoning Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.773699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:5bddbb51f361a3f7e0e706159c0118dbac3e10b5e39e95b54645df86970cb91c

Observation c9831f50-bdaa-4a63-9999-ce89286902a1 · outbound

This paper cites Fan, Y ., He, X., Yang, D., Zheng, K., Kuo, C.-C., Zheng, Y ., Narayanaraju, S.

Semantic-Enriched Latent Visual Reasoning Fan, Y ., He, X., Yang, D., Zheng, K., Kuo, C.-C., Zheng, Y ., Narayanaraju, S

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.787211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:7214e79efde86a5e5080baccd9c4a721f4843c304d8f6eaa338c81dfa10737d9

Observation e3fadffb-52c9-4660-ad93-24a21d412f49 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Semantic-Enriched Latent Visual Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.779157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:cca9cf3fc16e9d99938fded2d5e89715b483eb0c1cc9b07ab49300e2e47a5607

Observation a2d50dbf-c337-47f2-8d10-e4b1338a2ba8 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Semantic-Enriched Latent Visual Reasoning Training Large Language Models to Reason in a Continuous Latent Space

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.781451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:2bd9c24db05d6b9dd082965d87d42d46a223f0158f1e611f4a79cf01890a63dd

Observation 23c6f834-5a0d-4bd6-b3a3-3c5dc009c600 · outbound

This paper cites A diagram is worth a dozen images.

Semantic-Enriched Latent Visual Reasoning A diagram is worth a dozen images

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.863213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:b3a0c2fe013050f11426b0bc0eb3bacc4a2fd4fbbfbc1a61d5b8d1c4fe066b6a

Observation 7981f732-ccff-4dd8-a4d2-2522bdb2b282 · outbound

This paper cites Latent Visual Reasoning.

Semantic-Enriched Latent Visual Reasoning Latent Visual Reasoning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.789619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:69958088a5d6bf43d922edf2f3b139e4ddca1af72f8ff785ac4ac9f0cb821a21

Observation d0de086b-62db-40ae-92d4-842f5e7410dc · outbound

This paper cites Self-Rewarding Vision-Language Model via Reasoning Decomposition.

Semantic-Enriched Latent Visual Reasoning Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.783814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:90e220b70edba362a7eda04c4f8f3e0a616e7703f021370178b387ad5ba7e18b

Observation be1a0c49-92a6-4ada-b386-815925ec4914 · outbound

This paper cites VisionReasoner: Unified visual perception and reasoning via reinforcement learning.

Semantic-Enriched Latent Visual Reasoning VisionReasoner: Unified visual perception and reasoning via reinforcement learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.771162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:83097eee3c5386e9594f93457c7fb43e4af931992c5cd04bc3a3567a516bb17b

Observation be14c65a-cfbd-49b9-bfbe-8ab71e326f53 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Semantic-Enriched Latent Visual Reasoning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.792405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:0c2b4b549309aed8efa848eaded634e0d114088830b42b26b35802ebc0d59d8f

Observation 912d9698-11e9-4ddf-a44e-09154c0087c4 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Semantic-Enriched Latent Visual Reasoning Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.784282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:c69115433ab6484097a223d3959087d078533f5ef12cd56c4e04863ad4a6364b

Observation 5917e295-2c5d-499d-bed6-0bab35cd9b31 · outbound

This paper cites CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning.

Semantic-Enriched Latent Visual Reasoning CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.787051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:509bb2a78ed278a48c4783bd2f9c9d5d0272bfe0aab05e97d50bd2d34b727b15

Observation bffc31fc-6cf2-4ccf-a88b-f356890a06bf · outbound

This paper cites Mull-Tokens: Modality-Agnostic Latent Thinking.

Semantic-Enriched Latent Visual Reasoning Mull-Tokens: Modality-Agnostic Latent Thinking

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:55:00.773515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:f1a6a96f375a0041a8ee9142e60fef0778f3864bacb2df17af55686e8064a574

Observation db4ee2f6-c4f7-4e38-97bb-966aba36fff3 · outbound

This paper cites VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge.

Semantic-Enriched Latent Visual Reasoning VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:55:00.779176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:765ba8f8be27e69c4c8051332042ed390dab1d9ad3b1d060a49227ab8c73fe2c

Observation f70a3516-3401-4dc4-b33f-363901e667a1 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Semantic-Enriched Latent Visual Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:55:00.771422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:fb50b2a5015c6b4b66fa2a22fae81898ff5800fb7555fd7f265eac961ab3fa95

Observation 5b40eb2a-5fc4-4f12-be67-d2b34629640a · outbound

This paper cites Monet: Reasoning in latent visual space beyond images and language.

Semantic-Enriched Latent Visual Reasoning Monet: Reasoning in latent visual space beyond images and language

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.795638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:4912dce3c2a901d439449c8df154210b2bfe2e9a2ac38034b7f6402fb572e12a

Observation 545241dc-6884-4843-9c04-a2b7507f72d5 · outbound

This paper cites DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning.

Semantic-Enriched Latent Visual Reasoning DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.797119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:9923ee7ba97929720dd16aa05e25efb86465486a4a7fe39314286da84a6bbd67

Observation eb7fdc95-373b-4b52-b6ed-90d17d125bc0 · outbound

This paper cites Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens.

Semantic-Enriched Latent Visual Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.794774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:ae61e6cccf3f29d0ba7d2f403d8c37eb1ec3e94e06f2590a1ec6b2b26d920a55

Observation 4c60180f-f141-4bd9-add2-7747d4321919 · outbound

This paper cites Latent sketchpad: Sketching visual thoughts to elicit multimodal reasoning in mllms.

Semantic-Enriched Latent Visual Reasoning Latent sketchpad: Sketching visual thoughts to elicit multimodal reasoning in mllms

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.799636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:207f6c3ebd1454a3af5aebc82e56e42e0df6fd03372686d62303ba2d05573c9d

Observation bbeead68-561b-43bb-af50-b9751db0a840 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Semantic-Enriched Latent Visual Reasoning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.798016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:8e471582490342bba49763d8f5073f85637f09440eee0e57c66384819fd713ba

Pith citing papers

No inbound Pith citation observations are available.