Pith. sign in

Paper Citation Record · LEDGER

Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 51 inbound Pith citation observations for arXiv:2506.17218.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17218 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 51 of 51 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:03:11.626570Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 36665cc5-a23c-48f6-bd39-a2d07728ebc6 · inbound

Artificial Phantasia: Emergent Mental Imagery in Large Language Models cites this paper.

Artificial Phantasia: Emergent Mental Imagery in Large Language Models Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:44:22.816571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-21T21:41:39.111769Z digest=sha256:87e8cd92c12d9d4baa610ecc2b0f7bc9b0b4ac2131c9dace150b84feb594b6c7

Observation 6a920d1b-defb-48d5-99e2-44aef54f2b2a · inbound

LaRe: Latent Refocusing for Multimodal Reasoning cites this paper.

LaRe: Latent Refocusing for Multimodal Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T00:15:17.918780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:15:17.918780Z digest=sha256:22d64787d3ab250b75079fe51ba63ae9e429e29f6153a9fe464a42e2016d5a23

Observation 7755be9a-e860-4f09-9ecd-3c582b502929 · inbound

Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space cites this paper.

Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:03:38.956148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T23:02:28.588225Z digest=sha256:1385aac0970d9d3ec08033eee4283633e641f5c66970350571e203091ea7d3da

Observation c9dfc7df-df83-4685-b22c-9e0930945a01 · inbound

Forest Before Trees: Latent Superposition for Efficient Visual Reasoning cites this paper.

Forest Before Trees: Latent Superposition for Efficient Visual Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:01:05.142885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T15:58:03.654150Z digest=sha256:63ba0a12098e15d21c45947b9fe1966fde97f5d08f397d2bd35ac83f69cd7f5f

Observation c462f9b0-d339-43ff-b7b2-c46c9af7496d · inbound

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery cites this paper.

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T05:23:37.113526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:23:37.113526Z digest=sha256:24d841998e2e2605a95c5acf3b779262846aaf2f43ffea3b3f4fea808e333ba7

Observation 2c0febcc-f470-4614-8e5e-a846544225a0 · inbound

Towards Explainable Industrial Anomaly Detection via Knowledge-Guided Latent Reasoning cites this paper.

Towards Explainable Industrial Anomaly Detection via Knowledge-Guided Latent Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:10:32.068867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T03:08:58.617137Z digest=sha256:0dd18161f25818251308f900a4b211e35cb37a7f32faf4aa37fddd84fa79d6f6

Observation e4899f4c-3a74-4efb-9bb9-cff07c02b3f7 · inbound

Thinking with Drafting: Optical Decompression via Logical Reconstruction cites this paper.

Thinking with Drafting: Optical Decompression via Logical Reconstruction Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:30:32.901989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T03:28:32.264246Z digest=sha256:bcdbae12cfd68089d9628321f42fade7b7eee994e7a8f516e0dd6cd482e22c77

Observation da6a259a-1b97-4c95-a63f-63120205b382 · inbound

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook cites this paper.

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 268

Resolution
unresolved
no resolver link, observed 2026-07-13T14:03:01.974171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:03:01.974171Z digest=sha256:7168a3c9ca98010f6677847f87a9b5bc9c8241cf57120fa2197e4ae10a86a0f4

Observation de5ec7b8-5e89-464c-b21a-d2d6639880a3 · inbound

Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models cites this paper.

Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:16.251755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T20:48:52.130130Z digest=sha256:3ae913700298c0e24b6c3255eb80d3de8d908f1476b062ec3f9af9873b4c2362

Observation 570e609f-3dc3-41ae-9f50-edcad274807a · inbound

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators cites this paper.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:58:28.561653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:b8cb71b9c6fde5a984f7ad647b6aa4d9babe1a29ccc4d6121e09106542167429

Observation 612ebab2-b9f0-4a81-80bb-7c8312b62360 · inbound

Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs cites this paper.

Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:21:00.710558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T17:12:33.231488Z digest=sha256:c061925be361fe3cef46be36ed3a8e9e9f01f3788b2f1cc2b53d7f710514bf0e

Observation d03206f0-f22e-4813-ad06-6ceae2bd69f5 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:04.298952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:46:36.010169Z digest=sha256:41d6087ef89c9a759452f999a03c745e809cd3f752fc29f01af005f21abd2114

Observation 458f7de4-7dea-4b6a-852e-a8888eeb12bf · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:24.030033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:38:06.774877Z digest=sha256:c45aee8c9370583f02c491ec7fa98d28a22d8054d70405bb170c96393d33f227

Observation d4c4efca-738a-4fc9-aaf6-3766dfcdea06 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:27:29.419952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:26:59.917150Z digest=sha256:44ed6f1412dab37739d394c69f2b983b428db02af250dda77caa59095395088f

Observation 4c6ad6c9-fdf3-43df-8778-46a1964bd261 · inbound

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs cites this paper.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:26:01.935651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:708b706a7626c160639b1e9d453b94239525bfd772bb100db4fe3db7a4832308

Observation ab543ed3-ed1c-4606-ab82-51888183bee4 · inbound

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization cites this paper.

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:41:04.107147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T01:17:33.668451Z digest=sha256:b00ea6bb7b89c934a9d105f649f0f683c7c5d40ad82fc2cf468ecc76a17e06b0

Observation eaf0e5d4-9274-4335-b35c-4c0a6de31ea1 · inbound

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization cites this paper.

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T04:30:40.700508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-05T04:29:52.480873Z digest=sha256:21afe6e67b1dec823e1608ebd2c1434b06246c8a5227543c98a496458f2a6741

Observation e606c722-48a3-4a3e-a93c-928e86959529 · inbound

Using Machine Mental Imagery for Representing Common Ground in Situated Dialogue cites this paper.

Using Machine Mental Imagery for Representing Common Ground in Situated Dialogue Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:44:44.370094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T23:44:12.093852Z digest=sha256:0b2367690d0efb4b0a0bb0f854237e7b8a3e78f6dc3943938027b06c7a5f8eba

Observation e55f0731-9dd4-4a79-b85f-7c009246b541 · inbound

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning cites this paper.

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:31:30.107289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T05:47:17.494531Z digest=sha256:8e4a9ee58604e2ed64be60bda8d3147932be757babfcce071dfe7d656e6a1086

Observation 23a8ebd4-8ab7-409d-981b-149f75048b9b · inbound

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning cites this paper.

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:17:01.556370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T03:00:26.352130Z digest=sha256:c264b7849ad03c28e7a55f296aed9202c7622ceebd2f825199aec04bd835c68b

Observation 9a002af4-fc4d-42c0-aece-bc49f199e477 · inbound

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs cites this paper.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.544273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:902f92f5fec53f18b5baf766b02a219e50a051507d33a316528f0681fdde1c58

Observation b3257744-8fbc-4791-9ae0-3a7d3ed086b0 · inbound

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning cites this paper.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.653902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:0d96e2674e6f41b95db974bec700f0065a29899a482b67826ec9af347c80fc5d

Observation 74cd7e14-0902-4f29-aff7-f5af4f4c9fdb · inbound

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization cites this paper.

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:28.732533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:19:04.590978Z digest=sha256:d651f2ca0de2fa4764b854622d3977d59a18916160f928dbb2de7f59acff1c2b

Observation d0b8133f-092f-4eeb-af0b-6605c8d66833 · inbound

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization cites this paper.

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:12:28.475592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:08:52.879854Z digest=sha256:9dfaaf44a631ddcc88765d8194cacdbc0e38abada3644ea658d16260bc629b64

Observation 5f99935c-c814-4684-883d-47d9657b2ae0 · inbound

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs cites this paper.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:32:29.393294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:52919ad9d7456f4ec419d156b839cad0c3683faffd525c0957d90305d6db13b4

Observation 1e03b941-5181-467c-84b0-33d1b30f8bb6 · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:29.031174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:14:48.918959Z digest=sha256:545ccb0a902f6e00fbf365cbea632a712d0473af89a14f817beed1b5d4677886

Observation 1990d930-630b-401f-bcc9-866ca4e740da · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:48:00.596427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T21:47:50.595481Z digest=sha256:2235e84d5a2271870e1839d66d0fb94e53c23d48f972abae2bc8082f974f7ea7

Observation 3e38a023-09e2-4746-815b-27852730357b · inbound

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models cites this paper.

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:22:28.475379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-13T07:21:40.803291Z digest=sha256:611a0b00909ea8da87058fb9d2afb9f8c38df9ffac40929e97aeb2b008c34abc

Observation 80d05957-b336-4b20-81ef-ef27be586518 · inbound

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models cites this paper.

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T16:52:40.176676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-19T16:48:20.132240Z digest=sha256:346dfcf9e899dae91740965e395286726f945b178fc13fc44f53e14829b596c1

Observation 4003f055-b9ec-4812-9534-a002f7a4233c · inbound

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models cites this paper.

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T08:04:02.670279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-21T08:02:00.980836Z digest=sha256:4b305ad1ba80ceb9a3d604fb6e39ed37852e6c2556aed822eb6931d3c2f99bd1

Observation cfd3c977-46db-48ce-b02b-29fbdf57809d · inbound

Semantic-Enriched Latent Visual Reasoning cites this paper.

Semantic-Enriched Latent Visual Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:43:05.879132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T06:40:55.537488Z digest=sha256:5d12286950d7365827185bcd2500514e4cbdbb9277ad8fc19781fd95517add53

Observation eb7fdc95-373b-4b52-b6ed-90d17d125bc0 · inbound

Semantic-Enriched Latent Visual Reasoning cites this paper.

Semantic-Enriched Latent Visual Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.794774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T18:48:04.370230Z digest=sha256:17d9ecdaca3c58cb8cc79bea2edc3a42abb470daf94df1cf1142eaa6be58fce9

Observation ead86fe9-3639-44a2-b16f-34c923b51c50 · inbound

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models cites this paper.

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:02.161748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T23:04:21.463842Z digest=sha256:859363775ff68077a3599767ba9cc157861a0688cce8d51e1628c74b6359aef4

Observation bb8b15b3-3e80-4378-9f3e-ad26b684b386 · inbound

ReGuLaR: Relation-Grounded Latent Reasoning for Large Vision-Language Models cites this paper.

ReGuLaR: Relation-Grounded Latent Reasoning for Large Vision-Language Models Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T07:43:13.719590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T07:39:30.470848Z digest=sha256:18079f5c1c966972faae26d68f3fcd604c69c56f07abdfeaf1b595eee6c4a57d

Observation 6ba625fa-0b2a-4d40-a6da-45ce7edfe9b5 · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 78

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T19:52:36.136273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:23b1fe03a034e77c72606e6c501306c136337c5a3978754cf679cafccefc3c46

Observation 52cd2762-eb65-4afb-ad5e-50ad477b734d · inbound

MUSE: A Unified Agentic Harness for MLLMs cites this paper.

MUSE: A Unified Agentic Harness for MLLMs Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:26.900218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T11:14:31.395762Z digest=sha256:d031c616c64e81bc8d6915e12a365ae07b0b9678615bc6880ad556e0a8eb4a21

Observation 4a2d7c9e-0a34-482d-b5d2-cb01938830fa · inbound

Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models cites this paper.

Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:36:29.341511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T09:56:05.047862Z digest=sha256:bbd70c8a50818171a1d579775079386256ced02a314e2bbfe465d205ec1b3e8b

Observation 63ddab03-9136-4c40-b3b6-1604641457a1 · inbound

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation cites this paper.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.449658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:ecb4e7e397aba14d87e037571a83fa9403f4bdab94e205f7785646f016b2eca7

Observation f4828936-c2d8-4ae9-b5fb-d69f471cdd29 · inbound

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction cites this paper.

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T11:56:55.390969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T02:46:50.373450Z digest=sha256:a2734faf73754bb977da9c2a380f50137325982d01c36ef214dbe6c8507daf68

Observation ce7a9a48-e4f8-463f-9f72-d77fae6558b2 · inbound

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming cites this paper.

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T08:49:42.497880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T10:54:49.774490Z digest=sha256:7b328fa945898be105455dbd76de64a234dfbfc005ccb72f621cbfa8447012c7

Observation 4566a277-a066-4f04-b0c2-2e3f3257905d · inbound

Latent Visual States for Efficient Multimodal Reasoning cites this paper.

Latent Visual States for Efficient Multimodal Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:29:57.230432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T00:38:11.619574Z digest=sha256:06ba6e364261639ac93042c093c8edc5303c7c742aeccc13ab52323141659e11

Observation 9a33ec06-f553-4f50-86d2-08e2a6d7f46a · inbound

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning cites this paper.

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:30:06.790782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-25T21:23:10.051805Z digest=sha256:19b262dca6578b19c787bc05223891baa48e73514c168a507f171e231aa74673

Observation a49e6b72-f388-4693-968b-e6cf3c29b957 · inbound

Why Struggle with Continuous Latents? Interpretable Discrete Latent Reasoning via Rendered Compression cites this paper.

Why Struggle with Continuous Latents? Interpretable Discrete Latent Reasoning via Rendered Compression Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:44:18.300595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T06:43:20.893142Z digest=sha256:18ebdd4027cb9905434b3ea317200efd4593031d6ffc1e1e6201c2f32613cde7

Observation f9abef83-1b04-44f2-82ef-22b99f6802f1 · inbound

Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models cites this paper.

Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:44:19.792879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T06:35:16.868238Z digest=sha256:845be6aa25b13fa0fc5482107647dd81006b735725b5a5b2b04708135f6896fd

Observation 28cff490-528d-4d2c-8319-8c703119c9e4 · inbound

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning cites this paper.

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-02T15:07:03.807955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-02T15:06:38.215137Z digest=sha256:1b2569620fa740adf3164581d4a4b15911a6b6600951978dca9f68921aca2832

Observation 2e019a78-9093-467e-877b-216c3b32a019 · inbound

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space cites this paper.

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T06:13:16.894125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:13:16.894125Z digest=sha256:cfab75f775f12fec3468184f737b812052ab661f74466affbf17bde573caead7

Observation 4ee2f962-6b5e-4826-b422-3518eb5e0a6b · inbound

APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts cites this paper.

APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-10T13:37:06.835539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-10T13:33:17.574108Z digest=sha256:719d757a34107d95e5b6e97bd69805134752c23525404e6b5ccb9a797623a049

Observation 17d753cb-9dbb-475e-9f20-82b549c88feb · inbound

Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference cites this paper.

Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-10T03:06:43.680965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-10T02:59:21.353359Z digest=sha256:98a9da1dca6d11eaa029af428e3612f033fec0cc20e01004a4ec6395a367e993

Observation 2dd301e0-1288-4f03-aad9-e60e3a6b54d8 · inbound

OPLD: On-Policy Latent Distillation for Multimodal Reasoning cites this paper.

OPLD: On-Policy Latent Distillation for Multimodal Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T16:13:17.497272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T16:13:17.497272Z digest=sha256:627d0671cb97baa483abecca6b12db6c698d5fbf80bcd6e06308bb627bf8ab81

Observation 27386c4e-2c67-4476-a2b4-570bcd9b295d · inbound

LUT: Latent Utility Training for Visual Reasoning cites this paper.

LUT: Latent Utility Training for Visual Reasoning Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T00:28:16.479658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:28:16.479658Z digest=sha256:3ef3b8295b7ad5e7b22bcd867c090e8b5c4e3a66f228e09f5b56a6976edf18c4

Observation d9f77a4a-2a24-479b-8e50-eb7f1c101a61 · inbound

Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs cites this paper.

Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:11.626570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:11.626570Z digest=sha256:3074d1162b8e9aa1ca96593e516050d9e2fc6812e3b046ed406f1b3df7bda588