Pith. sign in

Paper Citation Record · LEDGER

Semantic-Enriched Latent Visual Reasoning

As of 16 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2605.19342.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19342 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T18:48:04.370230Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact16
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2d0d8c2c-85fc-4243-8bb4-ab2794584cac · outbound

This paper cites Qwen Technical Report.

Semantic-Enriched Latent Visual Reasoning Qwen Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.802055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:5525a69996879c776798fbf6e4cb157bb08ac24abc09fb442d5c6bdf66a40ba9

Observation 5b0361da-21aa-41f1-b420-ca32b8214f6b · outbound

This paper cites Qwen3-VL Technical Report.

Semantic-Enriched Latent Visual Reasoning Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.773699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:c587181cce97c31333fd1e76d24f15de33696bd2399aa70353bcc391dc7b3ad7

Observation c9831f50-bdaa-4a63-9999-ce89286902a1 · outbound

This paper cites Fan, Y ., He, X., Yang, D., Zheng, K., Kuo, C.-C., Zheng, Y ., Narayanaraju, S.

Semantic-Enriched Latent Visual Reasoning Fan, Y ., He, X., Yang, D., Zheng, K., Kuo, C.-C., Zheng, Y ., Narayanaraju, S

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.787211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:503d3818fdce28124815385e512a4b7344466ff557d0e85e20256252ebcfbdd5

Observation e3fadffb-52c9-4660-ad93-24a21d412f49 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Semantic-Enriched Latent Visual Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.779157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:d0db3bc6617a3f22cb89cbb8e36dc504da9142a7475e073e3d7d0c41c78701e4

Observation a2d50dbf-c337-47f2-8d10-e4b1338a2ba8 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Semantic-Enriched Latent Visual Reasoning Training Large Language Models to Reason in a Continuous Latent Space

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.781451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:06f5eb42e657c927cdc245e489891dbccae312eda4c149acb0a7936e407f9b91

Observation 23c6f834-5a0d-4bd6-b3a3-3c5dc009c600 · outbound

This paper cites A diagram is worth a dozen images.

Semantic-Enriched Latent Visual Reasoning A diagram is worth a dozen images

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.863213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:f2a97fb25c116bad086794bf234a7c89a872019ba5c8776e56c09350948d7519

Observation 7981f732-ccff-4dd8-a4d2-2522bdb2b282 · outbound

This paper cites Latent Visual Reasoning.

Semantic-Enriched Latent Visual Reasoning Latent Visual Reasoning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.789619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:866568ea7339acdc3b99d9ba1ebf107c320c3c0e64fa7dbd0ea9d48e07778900

Observation d0de086b-62db-40ae-92d4-842f5e7410dc · outbound

This paper cites Self-Rewarding Vision-Language Model via Reasoning Decomposition.

Semantic-Enriched Latent Visual Reasoning Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.783814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:c32f18d503184b8603c467b188797b966e373e2f11c8c293a06f7a44641d3d7e

Observation be1a0c49-92a6-4ada-b386-815925ec4914 · outbound

This paper cites VisionReasoner: Unified visual perception and reasoning via reinforcement learning.

Semantic-Enriched Latent Visual Reasoning VisionReasoner: Unified visual perception and reasoning via reinforcement learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.771162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:6b98936b9d1f7e5d1ff8b2952045c3e2a769886ffa755a4c7c77bf2e323d031d

Observation be14c65a-cfbd-49b9-bfbe-8ab71e326f53 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Semantic-Enriched Latent Visual Reasoning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.792405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:b35c77842e44e72b5d315305743ba250f5b4c4dd859fb24ccbb4971aa3eb7c7d

Observation 912d9698-11e9-4ddf-a44e-09154c0087c4 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Semantic-Enriched Latent Visual Reasoning Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.784282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:210cbb3c2f4dda9e0960b9febc12ad4f65e4daa856dadc356b417ad3228adea5

Observation 5917e295-2c5d-499d-bed6-0bab35cd9b31 · outbound

This paper cites CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning.

Semantic-Enriched Latent Visual Reasoning CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.787051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:7168dd2324b71978fdaf78a39f2f49c35bbaf0f6190cc9e66717fb7fc6e2c28e

Observation bffc31fc-6cf2-4ccf-a88b-f356890a06bf · outbound

This paper cites Mull-Tokens: Modality-Agnostic Latent Thinking.

Semantic-Enriched Latent Visual Reasoning Mull-Tokens: Modality-Agnostic Latent Thinking

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:55:00.773515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:fd3488b3c500312254dcc1c7366cb8aecba4341afd621c59035afdfddce52d37

Observation db4ee2f6-c4f7-4e38-97bb-966aba36fff3 · outbound

This paper cites VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge.

Semantic-Enriched Latent Visual Reasoning VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:55:00.779176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:924b63583c3c7bf7b1d7709462df88dc18370a2462b1406f2f4c71e1b61efe54

Observation f70a3516-3401-4dc4-b33f-363901e667a1 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Semantic-Enriched Latent Visual Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:55:00.771422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:4f100a55ce2415e68e2039ce0b6914bcc5b515c009f04582aaba3250f06bab28

Observation 5b40eb2a-5fc4-4f12-be67-d2b34629640a · outbound

This paper cites Monet: Reasoning in latent visual space beyond images and language.

Semantic-Enriched Latent Visual Reasoning Monet: Reasoning in latent visual space beyond images and language

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.795638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:1c46fb7282fd1eafad015bd0de0dad6eba70cdd39b79c428ed9c82fc4ccd537b

Observation 545241dc-6884-4843-9c04-a2b7507f72d5 · outbound

This paper cites DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning.

Semantic-Enriched Latent Visual Reasoning DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.797119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:46c6e6b0a0c0f99727bdf2b3be061dba6d81a0ea50345ce108fcdd63b2fee3a6

Observation eb7fdc95-373b-4b52-b6ed-90d17d125bc0 · outbound

This paper cites Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens.

Semantic-Enriched Latent Visual Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.794774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:652e7a05815a17460b947a5f98f7e2965892e984d55f756cca18bfa5f6f9c9a6

Observation 4c60180f-f141-4bd9-add2-7747d4321919 · outbound

This paper cites Latent sketchpad: Sketching visual thoughts to elicit multimodal reasoning in mllms.

Semantic-Enriched Latent Visual Reasoning Latent sketchpad: Sketching visual thoughts to elicit multimodal reasoning in mllms

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.799636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:9e060c142e3fc6d7fce43343023e0b813e9f103cc5bfe1b51f757db4f268b57a

Observation bbeead68-561b-43bb-af50-b9751db0a840 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Semantic-Enriched Latent Visual Reasoning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.798016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:b3a91566ec6ee2e086b1af94e59065b2bf3aa3b1c02ca4a89a58be2afa5d4dd0

Pith citing papers

No inbound Pith citation observations are available.