Pith. sign in

Paper Citation Record · LEDGER

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs

As of 1 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2605.02735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.02735 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T15:47:49.982564Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T16:13:17.573307Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact14
  • verified fuzzy31
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3a2c6d5f-8b2d-4508-80e2-23be1da2faab · outbound

This paper cites Qwen2.5-VL Technical Report.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:36:09.579251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:6900dd7cc87901e614a023a02212806c2804985f69e9024780070b0c3a0c3549

Observation cc5c100b-b3ca-4721-848f-e2f4db540260 · outbound

This paper cites Qwen2.5-vl technical report.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Qwen2.5-vl technical report

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.991467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:d6c917e26b0b2e57caa02aabc861166c6180a2929aad80d2970c87cf6873989a

Observation 66037d08-075a-43d3-afba-1f975d3c4ec1 · outbound

This paper cites UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.631193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:2da396bac655f708a748ebb1eda08fe9ea31163ac00fc77d65dd0da1eac06967

Observation 519fb987-1521-4309-885f-e59d3d4a5cf7 · outbound

This paper cites Sft or rl? an early investigation into training r1-like reasoning large vision-language models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Sft or rl? an early investigation into training r1-like reasoning large vision-language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.978226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:09fb09c86e54fc692eab2476dfd6b7e95a88b41427e4c158b5cbf0bb77fcb87e

Observation ba3b1d82-cfba-413d-aa70-88101c9f74d8 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:41:44.612219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:93899c946d5e35ab375631bd7c630240f02ec926c88fdb7576d0cd42c435a20b

Observation 529ad48e-7849-43da-8eb9-57d01a694020 · outbound

This paper cites Think with 3d: Geometric imagination grounded spatial reasoning from limited views.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Think with 3d: Geometric imagination grounded spatial reasoning from limited views

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.917571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:c022f95eb5d0e45032c376342f1fa8f9bb8e4d44a410d55113a7f89e98b488fc

Observation a6291f43-4fd3-49e4-ab53-0d6edfb426e7 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.984613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:a56b1e8cf5f4e8ffb44a5ac06423a0c726a80ae1913f52ac31e72ebc2f02a88d

Observation ed1cadea-6207-4b5c-8aa3-40ecee404002 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:d56dcc56424ec41b85d6a34c630f693bb2ee6a07e08aaee51111599e8e0ce6a8

Observation 8f4b7016-b63a-4096-99f4-7ea0ca3e514f · outbound

This paper cites Refocus: Visual editing as a chain of thought for structured image understanding.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Refocus: Visual editing as a chain of thought for structured image understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.998588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:9f4b3e3b1e1e581b56c648af7682777b97f8cd49fe2127aa7361e2f867e9f1c8

Observation c7ec3e4d-a5c5-44ff-bfb8-fa94413f0e62 · outbound

This paper cites Omni-MATH: A universal olympiad level mathematic benchmark for large language models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Omni-MATH: A universal olympiad level mathematic benchmark for large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.949019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:c9635458b003feb6727e4b86c2404076246d74779178562d1ea29df6104396db

Observation 8c2ee88f-a43e-4f6a-9657-550f5c015b93 · outbound

This paper cites Interleaved-modal chain-of-thought.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Interleaved-modal chain-of-thought

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.926002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:85c5e2837b740a9966e7d5fa629a1e70bb97613c388ac508a95df6d823045b72

Observation 8733b059-e4b1-4135-a7a5-2a44b73fe9cc · outbound

This paper cites Hallusion- bench: An advanced diagnostic suite for entangled language hallucination & visual illusion in large vision-language models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Hallusion- bench: An advanced diagnostic suite for entangled language hallucination & visual illusion in large vision-language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.929162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:f326c362b5811f2ca3a1d47aa70037367e0aab3480a786f0de4ac1d48ddbe6c9

Observation 2e6e0cbf-20a7-46d9-8a0a-bb77addb9cae · outbound

This paper cites Training large language models to reason in a continuous latent space.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Training large language models to reason in a continuous latent space

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.988336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:82443c3dc238c3c9bdd4f2ae45de9616c4f9769c23eebcc89ef6b52051b9b62a

Observation 7334ff85-fe9c-46cf-b2cb-208cf32169fb · outbound

This paper cites Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.Advances in Neural Information Processing Systems, 37:139348–139379.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.Advances in Neural Information Processing Systems, 37:139348–139379

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.900376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:666631126711d9653d71a67eff2d25bb0e66d04dc4c49fc3c7ffa05784332bc0

Observation f6492854-552e-441c-886a-1a135ae0c3e4 · outbound

This paper cites Latent visual reasoning.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Latent visual reasoning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.981224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:beaa25fd22b305443e3c04817819d15f60197389e4c7c96d62fa91c211160d3a

Observation 89b6da3e-4c93-40b6-9ea1-d17d776ad77d · outbound

This paper cites Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:36:09.585493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:6f34ebc8a75ac5ed11f808489a8f8425182a63906f98eaf1cb242d5fae4f113a

Observation c32a2eec-c6af-4597-af8e-4d48820cc54a · outbound

This paper cites Deliberation in latent space via differentiable cache augmentation.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Deliberation in latent space via differentiable cache augmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.888577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:7bddf003787927d8610e54374d44347ed341e56fffb77b40cd4c7318ea6842a6

Observation 68ca0288-8ca9-4a17-af42-fc0c5bc3dc2d · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.880107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:e2b8fb5471e12ba2fcd6a079ad1b00dc5eee276f222be80a75a9d6f8c8469aaa

Observation b958fd74-c7a4-425d-8d34-807959c82865 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.974192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:8bcda41aa243c73596a050a3c9106952ecc4fc9c30880a7a93ff1ec41064d601

Observation 68d1ffcb-7056-4b6f-a2bb-f2ef99f2c862 · outbound

This paper cites A survey on latent reasoning.arxiv.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs A survey on latent reasoning.arxiv

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.995586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:0c080eb726be1078481b983707a99f7d0b82ed2fec55d1376d296c1788e9e2a2

Observation 4c330125-d170-4c8c-b72b-e0ff100b5b08 · outbound

This paper cites LaRe: Latent Refocusing for Multimodal Reasoning.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs LaRe: Latent Refocusing for Multimodal Reasoning

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-27T02:05:07.580934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:ab5fa507bd901b99a7b8a104ec57d06b3f251f3fd5f940c207dc44b6b4f79346

Observation b4a50e53-5de6-4bc5-885f-d95cabce7c8c · outbound

This paper cites Compositional chain-of- thought prompting for large multimodal models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Compositional chain-of- thought prompting for large multimodal models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.884686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:0f13507f1fa07cc62b1ae137160fdadc346e746d6d33915583add7bc75f06153

Observation 29a6f4c4-47ec-4893-b174-e1fb73bfccf8 · outbound

This paper cites Chain-of-visual-thought: Teaching vlms to see and think better with continuous visual tokens.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Chain-of-visual-thought: Teaching vlms to see and think better with continuous visual tokens

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.615772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:1c57d70bd6ba6bd8d50322c5724512c1e41610f1ccbbae23a1e05a6692c33b38

Observation c279abb1-582c-4ea8-a1fa-71ac6c0e3188 · outbound

This paper cites an unresolved cited work.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-26T01:06:25.921888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:33257fb553106b92d9f8482bd7bb92b34362d8190d8f9c522f634fa206fc223f

Observation 7f7ffa2f-596d-483c-9e4e-3a37b04ef130 · outbound

This paper cites Codi: Com- pressing chain-of-thought into continuous space via self-distillation.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Codi: Com- pressing chain-of-thought into continuous space via self-distillation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.913984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:b747051e6b16240287e0b348341fba7ff7a9b0921e5439fa4462d92f944910ba

Observation f4718f76-c2a1-4d4d-9cbb-71c311ed7cce · outbound

This paper cites Swireasoning: Switch-thinking in latent and explicit for pareto-superior reasoning LLMs.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Swireasoning: Switch-thinking in latent and explicit for pareto-superior reasoning LLMs

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.892022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:bfa7c195123d5832780fc16c25653bd74c06fb59c1899c5771c6e1a1a39a0654

Observation 239432c9-d6be-4f55-b0e5-7af43b159614 · outbound

This paper cites Think silently, think fast: Dynamic latent compression of llm reasoning chains.Advances in Neural Information Processing Systems.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Think silently, think fast: Dynamic latent compression of llm reasoning chains.Advances in Neural Information Processing Systems

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.952508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:98a0249bd5209fe764c51ecbd9f2482c003dd423b592475ff16c1a3a7a04bb43

Observation df45b34b-7cbc-46f8-9c89-5636f31c0502 · outbound

This paper cites Visual position prompt for mllm based visual grounding.IEEE Transactions on Multimedia.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Visual position prompt for mllm based visual grounding.IEEE Transactions on Multimedia

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.942767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:4f8e6f229142b4fe415614a471255dcb601ec6d29fc3b8a1ee9a478b07123e0d

Observation eeb952a7-f6e8-4310-a861-0b4fbc97e246 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.896455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:cd45e0ebda1fcad3f1ed44f64c7843ff0796083dcc2b47d51199143431c542cf

Observation 5283bfd8-3b5b-4718-8c67-41d8414ae3e9 · outbound

This paper cites MLLM can see? dynamic correction decoding for hallucination mitigation.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs MLLM can see? dynamic correction decoding for hallucination mitigation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.933325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:788db5827915ea2c7f89a1f15ef26ed26e9b12b0e2b0d547bd3d64e207acdae1

Observation f569092e-9845-4108-8277-7acd9e655a91 · outbound

This paper cites Monet: Reasoning in latent visual space beyond images and language.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Monet: Reasoning in latent visual space beyond images and language

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.653976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:322c71e64d09d5eb0f48452ea037493154e9d36585fe0c6d3f0df4b25a3420bb

Observation 12eb500d-62b4-4d9f-aa2d-6fbc85357130 · outbound

This paper cites Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.558238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:75fdb455182bb6f3831272a45e296707ec14aa4b6efe4a19280689f0261af672

Observation 4cda328c-e917-4ec3-b5fa-2aeecf482a0d · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.866178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:a9366bec9095c1afaa99bb5e6cead3345a33d98d192105f3c973ff6262bd679e

Observation 85060b44-8a35-4beb-a7e4-89efad4d3055 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Chain-of-thought prompting elicits reasoning in large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.960259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:e83bffbd37afcd167227f68239429d979a9ae92abd1a3e29e9865435c7dfa7e5

Observation eafd2daa-e09f-407c-8b0e-5ec8fd47dbd8 · outbound

This paper cites Deepscientist: Advancing frontier-pushing scientific findings progressively.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Deepscientist: Advancing frontier-pushing scientific findings progressively

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.963118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:ea43da85f3b55f2ad4254b2b69c7b10d35b28ead99e6781e6df4762021610504

Observation f4a84c9d-0c51-4dc0-baac-ab8e51beee06 · outbound

This paper cites Mind’s eye of llms: visualization-of-thought elicits spatial reasoning in large language models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Mind’s eye of llms: visualization-of-thought elicits spatial reasoning in large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.936641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:84e65ff235a300b8b257a70e840ca88a9ff20255cff6ea9b8b1c685365cc4016

Observation d98ebe88-1f9a-4510-ba6b-49e297132419 · outbound

This paper cites Mini-omni-reasoner: Token-level thinking-in-speaking in large speech models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Mini-omni-reasoner: Token-level thinking-in-speaking in large speech models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.906496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:2f09b991402ee4cc59e1f0eb06730280f512f7344c14debee3abfa06f1d2559c

Observation 464b87df-fffc-4add-9810-35f6daf5ca5f · outbound

This paper cites Thinking in uncertainty: Mitigating hallucinations in mlrms with latent entropy-aware decoding.arXiv preprint arXiv:2603.13366.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Thinking in uncertainty: Mitigating hallucinations in mlrms with latent entropy-aware decoding.arXiv preprint arXiv:2603.13366

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.643792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:f664ef786401efec326e0f3ad4e05b628906e1e6fc357101f19ca4f2543721fb

Observation d8d1f104-679d-44e3-a497-29f2b09baa2e · outbound

This paper cites R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.945927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:7a3d002192a8ba1a36cbeb4d38e1481fabe1e10fd42b1b341c35ffa1f17ea91e

Observation 9a002af4-fc4d-42c0-aece-bc49f199e477 · outbound

This paper cites Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.544273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:4da72beee0f5990eefddb7731e271709dcae4ef3f9a78af931271aa604a19562

Observation 6aa3f6b0-6b4b-4750-8c29-d17ec79da595 · outbound

This paper cites Diffusion of thought: Chain-of-thought reasoning in diffusion language models.Advances in Neural Information Processing Systems, 37:105345–105374.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Diffusion of thought: Chain-of-thought reasoning in diffusion language models.Advances in Neural Information Processing Systems, 37:105345–105374

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.968119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:32bb7b57a6c52e903dd3863efd96fd6d0a61b6a488a108c4b02e9c4760716188

Observation a8175318-fcaa-4440-94db-57ad8b7afcd5 · outbound

This paper cites A survey on multimodal large language models.National Science Review, 11(12):nwae403.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs A survey on multimodal large language models.National Science Review, 11(12):nwae403

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.002163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:5b498a90b49b6f3ee7d5ece8876bb52977d7ed805d15bfbf83dcddb38fe3df7a

Observation 15e7bcb6-1472-4bf6-9258-728eb6b02d07 · outbound

This paper cites Mm-cot:a benchmark for probing vi- sual chain-of-thought reasoning in multimodal models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Mm-cot:a benchmark for probing vi- sual chain-of-thought reasoning in multimodal models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.609725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:6392519d465d2b13f276e502012a81329d27b8e0fd7689757f5df7987dff1a96

Observation 41d65bfa-371a-4fec-a898-e8b6bd0e72bb · outbound

This paper cites Multi- modal chain-of-thought reasoning in language models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Multi- modal chain-of-thought reasoning in language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.910229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:d2f7b2a4511edc0f1ef0f8c6cda336787ac9891f51888064971127f598f271c0

Observation 19905017-6415-4da5-86b3-dd82f0ef7ca5 · outbound

This paper cites Promptcot: Synthesizing olympiad- level problems for mathematical reasoning in large language models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Promptcot: Synthesizing olympiad- level problems for mathematical reasoning in large language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:25.955941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:0bf8a792f52a7183ea7679fac03bdbda32334853f30c6e4ff0d65cf22b0bd20c

Observation 82f72304-2dd1-4caf-b824-de36c125a1ef · outbound

This paper cites Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.568637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:dbb976f291b5689aec5e4f163c87f6986b79dea9e8aa4d57d9f4e6b52f401325

Observation 6c795bb6-3c09-41eb-bc64-b841891ac706 · outbound

This paper cites Intern-s1-pro: Scientific multimodal foundation model at trillion scale.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Intern-s1-pro: Scientific multimodal foundation model at trillion scale

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.664141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:dfac298b081d77a42950630af3cdd24d58e005ff0d76f693597756a9d1e9c45e

Pith citing papers

Observation 577b9ca7-52c3-4b31-a8f4-fca40918c88f · inbound

OPLD: On-Policy Latent Distillation for Multimodal Reasoning cites this paper.

OPLD: On-Policy Latent Distillation for Multimodal Reasoning Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T16:13:17.573307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T16:13:17.573307Z digest=sha256:47e87b7dabb750618eb58e3a30c7adb1b7f5ad63930cebb952b5d047c257b46b