Pith. sign in

Paper Citation Record · LEDGER

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs

As of 1 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2605.11856.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.11856 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T07:32:03.466222Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T07:47:31.725395Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T13:38:19.098461Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact26
  • verified fuzzy16
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6e4c205f-be54-4385-afe7-442d47abb0a9 · outbound

This paper cites Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:34:23.540456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:c7ae6437390e261be96a7563e6dcfcf942028834d36f0124096898942239ad88

Observation d0dedd2a-68cf-4a47-a1a3-37b8197b89c6 · outbound

This paper cites V*: Guided visual search as a core mechanism in multimodal llms.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs V*: Guided visual search as a core mechanism in multimodal llms

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:32:30.748855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:45b3db2b97ef568b31af843c4dea1fcd42cf74a5085b5f0c07010af61394d61f

Observation 70a06127-e3e8-49ae-9dad-b6a44a274cfc · outbound

This paper cites Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:32:30.764546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:bbe69380695fae2d0271031e4eb4414528350245011959bd1b44a6a6e9916f5e

Observation c90818c4-39fd-4989-b189-fcdca557321d · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:32:30.752238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:e46ffdf75d9ad96742c695ad529eaa2485a239bc702bf571a5049a64d53bb89e

Observation a46a3294-d575-4971-983b-365af24137ea · outbound

This paper cites Mme-realworld: Could your multimodal LLM challenge high-resolution real-world scenarios that are difficult for humans? InICLR.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Mme-realworld: Could your multimodal LLM challenge high-resolution real-world scenarios that are difficult for humans? InICLR

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:32:30.728858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:5049122760f4c27f283cd5192ff1dbf352d388038a8e0224714b1d500b14e2c7

Observation fa674638-a829-47fe-889e-8a772044b1bd · outbound

This paper cites Smith, and Ranjay Krishna.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Smith, and Ranjay Krishna

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:32:30.745369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:2897b2299507f874d3fb07cd8e36161ccf279a5d26786557f875977f2b7ad34d

Observation 7b3dd458-1509-4e74-8f23-263ef115effa · outbound

This paper cites Vtool-r1: Vlms learn to think with images via reinforcement learning on multimodal tool use.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Vtool-r1: Vlms learn to think with images via reinforcement learning on multimodal tool use

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:32:29.374623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:a2553c1b81ed50bedfdd5faf9e3abddd93335473199218d473271aef0537e13f

Observation 188e7e8a-7044-4727-ac5e-b9fa6597ddb4 · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:32:29.420726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:b3a0ee98ce4bb0a67eccc0b6ef4610132c616ecdb2f1d73707805c1241ae14f0

Observation 35e617ca-f4f6-43c7-a018-ee7baef37a02 · outbound

This paper cites Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:22:27.090901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:0005f189d11dcad682c740710fae67ae0195d11f6bb6c6fa245909cff659998c

Observation 5f99935c-c814-4684-883d-47d9657b2ae0 · outbound

This paper cites Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:32:29.393294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:e8300cb16bf60f6a1bc6aaa74612b3c5d586c6aa1cd7bb20bdb9a6593820f4ec

Observation 9aded464-3ad4-4290-9513-98cad20d3116 · outbound

This paper cites Latent Visual Reasoning.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Latent Visual Reasoning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:41:30.500607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:3615b636e06d04943256e89e4ed2d650ceb446318321bce2b9c8df7bccf3ac0d

Observation 708d1a0c-3374-4dbe-834f-65601273b857 · outbound

This paper cites Monet: Reasoning in latent visual space beyond images and language.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Monet: Reasoning in latent visual space beyond images and language

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:32:29.387207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:6b0bebd6bd9f1ba2971068cef8b35c8ddad8c702a7081d708000437adfadf98f

Observation 0c1a9b78-3b78-444c-b17e-5a4bb08faa5a · outbound

This paper cites Sketch-in-latents: Eliciting unified reasoning in mllms.arXiv preprint arXiv:2512.16584.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Sketch-in-latents: Eliciting unified reasoning in mllms.arXiv preprint arXiv:2512.16584

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:32:29.309729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:8efa9185c83717db79bb61fa35f2f640350630b127ddf9ec49d7f5c2cde7f7d9

Observation 78d87c49-8154-448c-8817-e8647146a4a3 · outbound

This paper cites Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:32:29.443170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:8881266b9744a6324ffed9fc6c63047c461e4a26eff9864abbba27958f4ad0c6

Observation 72477733-2df0-4a69-8be5-372fadc98951 · outbound

This paper cites Chain-of-visual-thought: Teaching vlms to see and think better with continuous visual tokens.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Chain-of-visual-thought: Teaching vlms to see and think better with continuous visual tokens

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:32:29.321234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:85747675f6bac2ce776c5f458e115e35a991b5f601a0518e52c46d25ed27a369

Observation da58669f-b94e-4167-81c6-8168cbfae508 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Training Large Language Models to Reason in a Continuous Latent Space

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:32:29.315676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:9c14449565587736ce61f7de97799d9964c29138f98862e0cae9cf65cb8abdf5

Observation d1bb62b3-0fc0-489f-a284-00343f25ba29 · outbound

This paper cites Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:32:29.327633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:c4fbb457d98bfd7e3814d39d7a5f5d37133cb9f46e0e113dfd5291fdfb8ef33f

Observation e35ead35-e169-48d5-bda4-5f564d0dfaed · outbound

This paper cites The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-08T02:03:55.088792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:578ae64c908c2f04350ad1d9d7bbc388a39d66c27c9bef8494902db418b38418

Observation e01bd607-1f63-4cb0-a68a-461c49279de7 · outbound

This paper cites Visualizing thought.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Visualizing thought

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:32:30.710078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:3224371524404d1996c8b06026ab329ce4d11f770f89bc87d27d3b4b86983e9f

Observation c020d510-63f3-47c0-971c-b970a9622734 · outbound

This paper cites MIT press.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs MIT press

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:32:30.714164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:fe29032a66b3266cbbdc70365cc93d3f04c8ede62f45d8c29db302fd6f4f87a3

Observation ae4437ed-a8ee-40f5-bddb-ff61fc9a7353 · outbound

This paper cites DeepSeek-OCR: Contexts Optical Compression.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs DeepSeek-OCR: Contexts Optical Compression

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:32:29.448698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:6934f207b0b7507e99c954d51b8d394ea7fa6ae922af81f8abd57d90f509f0b0

Observation dd9681af-dae1-4dcd-a15c-e18ebf63b01d · outbound

This paper cites Deepseek-ocr 2: Visual causal flow.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Deepseek-ocr 2: Visual causal flow

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:32:29.350855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:2160d8551ca0764b8ff31c60da5a8a6f0e1e27d0f59647669f76ebe829c6cece

Observation b8e1891c-c662-413a-96ae-15ede4279f62 · outbound

This paper cites Qwen2.5-VL Technical Report.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Qwen2.5-VL Technical Report

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:32:29.362335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:5ee7f807700503d95cc478890be140704b0a00c97372d6f1603a78d86edddc48

Observation 7c45dfc5-734a-47f1-b215-58d34ba55fc4 · outbound

This paper cites Qwen3-VL Technical Report.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Qwen3-VL Technical Report

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:32:29.398745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:528ff18bf6b56b5f902c9dfc462e8925abf39bf6bcc04184a3c0a325e40c5e10

Observation 19073b14-44b4-4473-bd2f-80103f7a1c60 · outbound

This paper cites Onelatent: Single-token compression for visual latent reasoning.CoRR, abs/2602.13738.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Onelatent: Single-token compression for visual latent reasoning.CoRR, abs/2602.13738

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:32:29.357250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:27483f1692429de72309e268daefe8c2b81ed564387e82d334c9283403bd20a9

Observation a4a6595e-2d3e-4a5b-83c2-bacef615f717 · outbound

This paper cites Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:32:29.303557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:0a76601672b8d116780c849cc864a389a32a6edf4079765d0db714b8112af2dc

Observation ce2199f9-5350-4d7b-aaf3-e8bfac8a39c3 · outbound

This paper cites GPT-4o System Card.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs GPT-4o System Card

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:32:29.380375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:58443c3c9b1229bcf3b6026192a49ea81ed3b6d7691a3c5fa229c6a642ab3acb

Observation d97c809e-1c30-4871-9a3e-afda60a8cfa2 · outbound

This paper cites Token fusion: Bridging the gap between token pruning and token merging.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Token fusion: Bridging the gap between token pruning and token merging

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:32:30.717384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:d901be22f6ef2da7813f202b409046eac28ca1d0ed65453c3cda0245c0e4d34a

Observation 2d9ec589-cfb8-4756-bf8c-b3f019891e01 · outbound

This paper cites Layer by Layer: Uncovering Hidden Representations in Language Models.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Layer by Layer: Uncovering Hidden Representations in Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:30:37.615120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:7916becc782dc5c6d2073a8a7454ea25262fb2207231788b5a2d3444eef2c51b

Observation 729e7149-e502-4063-ad96-9c568b3a3011 · outbound

This paper cites Representation alignment for generation: Training diffusion transformers is easier than you think.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Representation alignment for generation: Training diffusion transformers is easier than you think

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:32:30.724247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:c8a75e9b8a4872a66a529cdc47e94bfe526e1ef9cffcf22a972d4c3be26d5e40

Observation 2d4cae9c-be33-475b-b9ad-9b15033c3085 · outbound

This paper cites Tamp: Token-adaptive layerwise pruning in multimodal large language models.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Tamp: Token-adaptive layerwise pruning in multimodal large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:32:30.720835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:a205d95362943f21f9ab00d1220d22083ae215c1568f32b42e2538a488a2bc6b

Observation 8dc8993a-6458-4839-b7ee-4f59ac2fe259 · outbound

This paper cites Your large vision-language model only needs a few attention heads for visual grounding.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Your large vision-language model only needs a few attention heads for visual grounding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:32:30.736427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:aafb27e23e2830f030f9db7539dc8aa2042dac7152fba59a0acce479fdd72075

Observation 92a01683-c2ac-4333-97da-26c6d49c4f63 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:32:30.732983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:c05cce098321f1e4484ab48030b49bd65fa77a5a12a8458f86be5a7b5a824708

Observation f249090c-00d3-4db9-8b6d-37e4c327e92f · outbound

This paper cites Zebra-cot: A dataset for interleaved vision language reasoning.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Zebra-cot: A dataset for interleaved vision language reasoning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:32:29.333597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:18c75ecd684a2d998363ff4880ea7cf3f70eac75cfcc797a41ef4f1d6e593d70

Observation 59f6eab1-3774-490f-80d0-3adf561923e2 · outbound

This paper cites Towards vqa models that can read.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Towards vqa models that can read

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:32:30.769502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:2ac5a2fdaab78712dcdad317c34260185f47eaeb0221fd76d819ff71ffe0f38e

Observation e5426542-f55c-4b5d-a579-277c59fc8689 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:32:29.410502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:9db65e30cee6c2ae43599126f23e003d8c0ecbb3df0e20b2e192d2a1480c6000

Observation 0866d340-7582-42bc-8547-3398a0513d59 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Evaluating Object Hallucination in Large Vision-Language Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:32:29.404046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:7e42df3d6d54dddef6a435f18a5053058813252f33a94c8e21f68d5c2ca22a4f

Observation 76e35c7d-dd7a-4c35-842b-eb8a79b59493 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:55:41.398838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:95f9c35f029052f5ef92c2a630a1f9e527b26f191f94a3f57e613982c0882100

Observation d09703d3-6e32-438f-9a3d-a83ae7aba2ef · outbound

This paper cites Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:32:29.437886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:faabc65c5eca9177f8faf5d4985d226722dc99178ca8bfd61a2b7549f2acd8b9

Observation 0ac2d947-d46e-4000-82c8-b39ce1bf6c7c · outbound

This paper cites As outlined in Algorithm 1, the canvas width is constrained by a minimum threshold and the scaled auxiliary image.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs As outlined in Algorithm 1, the canvas width is constrained by a minimum threshold and the scaled auxiliary image

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:32:30.760221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:48de4c8c1ea671250146e58fe37b0b157c54e6e39d72f2e7e642566154fd6082

Observation ee2cde56-338c-47f8-a770-801a626c7756 · outbound

This paper cites As detailed in Algorithm 2, the canvas has a fixed width but dynamic height.

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs As detailed in Algorithm 2, the canvas has a fixed width but dynamic height

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:32:30.740309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:4cac9bcd357ef980381106567f16c4ef73961793cc256ff3cf0c6764a65f8470

Observation 0aff2764-704c-4721-8767-5dd9fb13983d · outbound

This paper cites As described in Algorithm 3, the canvas dimensions are rigidly constrained (e.g., 1024×1024 ).

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs As described in Algorithm 3, the canvas dimensions are rigidly constrained (e.g., 1024×1024 )

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:32:30.756424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T07:32:03.466222Z digest=sha256:468c8109d1fc63d31f52749fb37465edbff0501346581f78b53d66e1bc835dcd

Pith citing papers

Observation 23770e47-94ca-4045-b471-e45bb3e7c596 · inbound

Language-Guided Abstraction for Visual Reasoning cites this paper.

Language-Guided Abstraction for Visual Reasoning UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-03T13:38:19.099762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-06-27T07:47:31.725395Z digest=sha256:281c77fc3a34d2a1e5deac2b146f890397a3e587b0482f0c5f52f45046ad56ed