Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T07:32:03.466222Z
Paper Citation Record · LEDGER
As of 1 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2605.11856.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T07:32:03.466222Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T07:47:31.725395Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T13:38:19.098461Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6e4c205f-be54-4385-afe7-442d47abb0a9 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d0dedd2a-68cf-4a47-a1a3-37b8197b89c6 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs V*: Guided visual search as a core mechanism in multimodal llms
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 70a06127-e3e8-49ae-9dad-b6a44a274cfc · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation c90818c4-39fd-4989-b189-fcdca557321d · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a46a3294-d575-4971-983b-365af24137ea · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Mme-realworld: Could your multimodal LLM challenge high-resolution real-world scenarios that are difficult for humans? InICLR
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation fa674638-a829-47fe-889e-8a772044b1bd · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Smith, and Ranjay Krishna
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 7b3dd458-1509-4e74-8f23-263ef115effa · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Vtool-r1: Vlms learn to think with images via reinforcement learning on multimodal tool use
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 188e7e8a-7044-4727-ac5e-b9fa6597ddb4 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 35e617ca-f4f6-43c7-a018-ee7baef37a02 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 5f99935c-c814-4684-883d-47d9657b2ae0 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 9aded464-3ad4-4290-9513-98cad20d3116 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Latent Visual Reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 708d1a0c-3374-4dbe-834f-65601273b857 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Monet: Reasoning in latent visual space beyond images and language
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0c1a9b78-3b78-444c-b17e-5a4bb08faa5a · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Sketch-in-latents: Eliciting unified reasoning in mllms.arXiv preprint arXiv:2512.16584
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 78d87c49-8154-448c-8817-e8647146a4a3 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 72477733-2df0-4a69-8be5-372fadc98951 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Chain-of-visual-thought: Teaching vlms to see and think better with continuous visual tokens
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation da58669f-b94e-4167-81c6-8168cbfae508 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Training Large Language Models to Reason in a Continuous Latent Space
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d1bb62b3-0fc0-489f-a284-00343f25ba29 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e35ead35-e169-48d5-bda4-5f564d0dfaed · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e01bd607-1f63-4cb0-a68a-461c49279de7 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Visualizing thought
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation c020d510-63f3-47c0-971c-b970a9622734 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs MIT press
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation ae4437ed-a8ee-40f5-bddb-ff61fc9a7353 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs DeepSeek-OCR: Contexts Optical Compression
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation dd9681af-dae1-4dcd-a15c-e18ebf63b01d · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Deepseek-ocr 2: Visual causal flow
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b8e1891c-c662-413a-96ae-15ede4279f62 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Qwen2.5-VL Technical Report
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 7c45dfc5-734a-47f1-b215-58d34ba55fc4 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Qwen3-VL Technical Report
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 19073b14-44b4-4473-bd2f-80103f7a1c60 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Onelatent: Single-token compression for visual latent reasoning.CoRR, abs/2602.13738
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a4a6595e-2d3e-4a5b-83c2-bacef615f717 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation ce2199f9-5350-4d7b-aaf3-e8bfac8a39c3 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs GPT-4o System Card
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d97c809e-1c30-4871-9a3e-afda60a8cfa2 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Token fusion: Bridging the gap between token pruning and token merging
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2d9ec589-cfb8-4756-bf8c-b3f019891e01 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Layer by Layer: Uncovering Hidden Representations in Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 729e7149-e502-4063-ad96-9c568b3a3011 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Representation alignment for generation: Training diffusion transformers is easier than you think
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2d4cae9c-be33-475b-b9ad-9b15033c3085 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Tamp: Token-adaptive layerwise pruning in multimodal large language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8dc8993a-6458-4839-b7ee-4f59ac2fe259 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Your large vision-language model only needs a few attention heads for visual grounding
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 92a01683-c2ac-4333-97da-26c6d49c4f63 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f249090c-00d3-4db9-8b6d-37e4c327e92f · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Zebra-cot: A dataset for interleaved vision language reasoning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 59f6eab1-3774-490f-80d0-3adf561923e2 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Towards vqa models that can read
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e5426542-f55c-4b5d-a579-277c59fc8689 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0866d340-7582-42bc-8547-3398a0513d59 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Evaluating Object Hallucination in Large Vision-Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 76e35c7d-dd7a-4c35-842b-eb8a79b59493 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d09703d3-6e32-438f-9a3d-a83ae7aba2ef · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0ac2d947-d46e-4000-82c8-b39ce1bf6c7c · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs As outlined in Algorithm 1, the canvas width is constrained by a minimum threshold and the scaled auxiliary image
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation ee2cde56-338c-47f8-a770-801a626c7756 · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs As detailed in Algorithm 2, the canvas has a fixed width but dynamic height
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0aff2764-704c-4721-8767-5dd9fb13983d · outbound
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs As described in Algorithm 3, the canvas dimensions are rigidly constrained (e.g., 1024×1024 )
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 23770e47-94ca-4045-b471-e45bb3e7c596 · inbound
Language-Guided Abstraction for Visual Reasoning UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.