Pith. sign in

Paper Citation Record · LEDGER

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

As of 22 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 20 inbound Pith citation observations for arXiv:2505.16192.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16192 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:10:14.202049Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:12.926143Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 0950a071-901a-45c6-a06a-8ba43b563435 · outbound

This paper cites https://deepmind.google/technologies/gemini/.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought https://deepmind.google/technologies/gemini/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:20.738216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:07.005856Z digest=sha256:b73a9a952f28fba5837a76c192e54ec8eab64b4343dab3781bcf9d4052dd93e1

Observation de2a7e6e-7fe8-4598-8bdc-e4b044e5c007 · outbound

This paper cites https://openai.com/index/introducing-o3-and-o4- mini/.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought https://openai.com/index/introducing-o3-and-o4- mini/

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:20.707078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:07.072665Z digest=sha256:64d177e66c4766081230f890b9b807ebda2d4f783fe3dc293bc81f57b393d52c

Observation a989b189-95c9-4c95-9d01-a4190fa85214 · outbound

This paper cites https://qwenlm.github.io/blog/qvq-72b-preview/.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought https://qwenlm.github.io/blog/qvq-72b-preview/

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:20.668700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:07.162899Z digest=sha256:a13117e704822cf76cd242529b17fc5cf3e68d4497c3ae4eb08b878be210ee00

Observation a5c7d6b5-0e62-42ef-a7cf-7ac074e493a5 · outbound

This paper cites https://huggingface.co/datasets/TheEighthDay/SeekWorld.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought https://huggingface.co/datasets/TheEighthDay/SeekWorld

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:20.618500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:07.247025Z digest=sha256:fc2247c3d5fc106948cf42e830cfbe3db1ae5324caa7b38ae7651819abdeab04

Observation 29727806-d59a-4088-9f2e-31af322b8723 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Flamingo: a visual language model for few-shot learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:20.586240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:07.334328Z digest=sha256:174be72130fde389b55e87cdba0d4a69a694df3b70f974c40bd1ab3796902460

Observation f82d7041-8c4e-4e48-8aa9-d6e71c94b091 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Gemini: A Family of Highly Capable Multimodal Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:07.421698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:07.421698Z digest=sha256:b1e09394b0c4e127f7f9944ec0046b80d0e83455bde67dcaa965f9eae52a8e27

Observation 43fa7503-a1e3-44dc-b529-0cb4945c3b7f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:07.490150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:07.490150Z digest=sha256:b0e8dfe271196089a669164ab218902057fc41d2315f31652f3e48fc118ae029

Observation 84c52e37-6fd1-4bec-b5ce-fdede12d09c4 · outbound

This paper cites Qwen2.5-VL Technical Report.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Qwen2.5-VL Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:07.588509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:07.588509Z digest=sha256:6103357b8afbe5646faf5079271734832b3b81cc3bcd68268427c6d7d8b09b70

Observation c98685ab-94a2-4e9f-8982-ff40d49a9a6a · outbound

This paper cites M 3cot: A novel benchmark for multi-domain multi-step multi-modal chain-of-thought, 2024.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought M 3cot: A novel benchmark for multi-domain multi-step multi-modal chain-of-thought, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:20.556290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:07.654396Z digest=sha256:3abd48789d88c2a44a4c975aed98680fb00277df65c1e16c746dd46f1762fd16

Observation 4afef9e2-3862-46db-8b3f-ac67746f1a82 · outbound

This paper cites an unresolved cited work.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:10:20.512907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:07.748702Z digest=sha256:be0609cb541e83cebba6850ee4a3c421ca9caf9a50d46badbdf2cd381208f4eb

Observation 512219ac-acd7-4f08-82bd-2822370f9ec0 · outbound

This paper cites Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:07.831896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:07.831896Z digest=sha256:dda62be6d6c1fa9f88d84b1cd365d987b33e5f46f2501095988b234f66b652c9

Observation 450f395a-44f9-4b55-bca0-8b2b8bf3bbc2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:07.932135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:07.932135Z digest=sha256:bd714e3a55c472e2c8cf59662b9beac7ba25d2b9c978d53f834212685155a63d

Observation 75eaa120-3b25-458f-ac2c-201dc42adbf5 · outbound

This paper cites Virgo: A preliminary exploration on reproducing o1-like mllm, 2025.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Virgo: A preliminary exploration on reproducing o1-like mllm, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:20.467100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:08.030691Z digest=sha256:0a6a95e7a33e861058b35ddebb0a214ebf76faf6a40f62325e803f30df955af7

Observation ba5c080c-43ad-43e0-b9ba-ceb6db7b3118 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:08.130017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:08.130017Z digest=sha256:93c7adccf11a8e06c4d0a0f8a748bfa17a6364ac7fe8d097ee73700386c7fb80

Observation 6131cccc-c161-40d4-85b9-7360e6fa1a28 · outbound

This paper cites Hallusion- bench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models, 2024.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Hallusion- bench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:20.430749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:08.243185Z digest=sha256:ae6c9964b5f4945b29c7cc236d712c2fc97a93c86f4d7ba93bf00deac006c80c

Observation 2d87f4ec-d619-4cb2-b5c5-a6634f812c64 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:08.359401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:08.359401Z digest=sha256:13e76bc768175f635380d09189c03346b90c34fe9947a1dcca90771a8e5e834f

Observation 423304cc-522e-4e76-985c-534d7140d033 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:08.471968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:08.471968Z digest=sha256:f49abbcfb1d7197bd4d7737f1e8d15b0838c46338f5315c23f8f157d9c00fed0

Observation 4a111b37-51f5-41c5-8693-1d7eec2198d8 · outbound

This paper cites MathPrompter: Mathematical Reasoning using Large Language Models.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought MathPrompter: Mathematical Reasoning using Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:08.588351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:08.588351Z digest=sha256:150495e18bb6eb0f7f1b745e0db611d971be339f0c40e57a5b9b8944f0a4c9fe

Observation 7ce2b8b3-de00-4422-ab69-f0ead01be8b9 · outbound

This paper cites Tab-cot: Zero-shot tabular chain of thought.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Tab-cot: Zero-shot tabular chain of thought

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:20.379552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:08.669117Z digest=sha256:0e0b60e258e71368acb7fb22e3f001151873f5559275454e1e8844e634dcc9e9

Observation 36ec47c7-fa3a-46cd-8b34-f9cefd8e4966 · outbound

This paper cites Imagine while reasoning in space: Multimodal visualization-of-thought, 2025.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Imagine while reasoning in space: Multimodal visualization-of-thought, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:08.770867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:08.770867Z digest=sha256:f2992b626466f4974171ad26d1b14839f2fad571767b5c64043e174f511bc3b6

Observation 3be5777e-2352-4c7a-a174-753e4a5195bf · outbound

This paper cites Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models, 2024.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:08.918857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:08.918857Z digest=sha256:176ac7351d33933ffa33e6a6b5de4da47a20b57d5f66da855b9d21a9b653d643

Observation b6394dd6-7423-497d-8f53-349fcbf14425 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:20.217454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:08.982717Z digest=sha256:8a575b2d1cdceac915e29dbf86b30f5a2682af48d2941255d3be69500dee50c7

Observation 031441a8-32e5-4bf6-98e9-0c8f5558a271 · outbound

This paper cites Let’s verify step by step, 2023.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Let’s verify step by step, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:09.079435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:09.079435Z digest=sha256:66e97d3387381b5465b116d175c9e31aec11c16d63d23eb176d710a762240f8f

Observation 46af13df-a9ab-4ea2-bd91-0c9967fed09b · outbound

This paper cites DeepSeek-V3 Technical Report.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought DeepSeek-V3 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:09.169201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:09.169201Z digest=sha256:8cfbc34173f7c62b8192c8862e7eab1c336509d485c68c3b5cf0e676532a286d

Observation b34ef889-8ed0-4c08-bfcc-4db52ab25a86 · outbound

This paper cites Visual spatial reasoning, 2023.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Visual spatial reasoning, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:09.284782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:09.284782Z digest=sha256:5de8c12f01af7543ed3d492479e8a09c6aaaab2865c2fbfe09bbe2546a5c8494

Observation 8c310283-a77a-4a0b-8a95-3c006dbed489 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Improved Baselines with Visual Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:09.379608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:09.379608Z digest=sha256:dea1ac2dfdede1b5330322cd3908420da0944fdd69ade69ea49b0d86fb956dce

Observation 3c429e97-4e44-4af1-99f7-8c07904cb77b · outbound

This paper cites Seg-zero: Reasoning-chain guided segmentation via cognitive reinforcement, 2025.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Seg-zero: Reasoning-chain guided segmentation via cognitive reinforcement, 2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:19.943083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:09.468404Z digest=sha256:b5ef96f14e1a28061fc55e9e5f68d13254bac727e5aa15e546488f4d03001b2d

Observation d16f2635-af50-4e65-b935-ab1a39e61149 · outbound

This paper cites Visual-rft: Visual reinforcement fine-tuning, 2025.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Visual-rft: Visual reinforcement fine-tuning, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:09.601229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:09.601229Z digest=sha256:9485323c1cafa601881e5a0e2e1b5ba94e75b52ec38317b7e66a2bf13b0638ca

Observation 98631811-9405-49ef-bfe7-6677c9eaf95c · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:09.689775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:09.689775Z digest=sha256:41b8a18df86887b47832f382bebeeceabafe767d257267e0017c4cc68810c6d7

Observation d8b818d9-0a87-457c-b8d7-64dc11dad799 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:09.801233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:09.801233Z digest=sha256:08fed34f6bbb15847c13059d970a412fc815ede36c2a7e897eff8639962d2bff

Observation d254833d-c3cf-412e-89b0-e1a3cb72c310 · outbound

This paper cites V Jawahar.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought V Jawahar

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:09.954006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:09.954006Z digest=sha256:ee6527d37fe06560d39309efad4648f7e72c4d4e22b28dd9bb664dd4137a2299

Observation 596b13ad-e256-46ac-addb-77152f353952 · outbound

This paper cites an unresolved cited work.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:10.053966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:10.053966Z digest=sha256:1b47c3c25f802626d4823bcf752242ab076ed07a7606fb3e3ea910082042eaae

Observation 0562d9e9-61e4-4f1f-9d53-cc899a91a916 · outbound

This paper cites Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:10.150431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:10.150431Z digest=sha256:dc4529f88726f8717733b745a15a0cc5a7f2e6862d20ca9c57b16a3f141b9f46

Observation 3caf8933-d096-4cbe-9b05-67d07173d7ce · outbound

This paper cites Skeleton-of-Thought: Prompting LLMs for Efficient Parallel Generation.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Skeleton-of-Thought: Prompting LLMs for Efficient Parallel Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:10.218028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:10.218028Z digest=sha256:f43450f118a20c871e6efd1baa60ea65f2c878b6dd318a8d254fd919dbfc584f

Observation 6cf2ee6d-e3ed-41ce-b7ce-18f13abab1ed · outbound

This paper cites an unresolved cited work.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:10:19.624614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:10.339377Z digest=sha256:32844698fcaac21aad408dbd3bb97ecc1b73dc8d2b81ce022671ee0b2cc280ec

Observation 439815e4-6c98-4276-8837-22a22b0ddd60 · outbound

This paper cites an unresolved cited work.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:10:19.302376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:10.436748Z digest=sha256:748be5c2414ccaa8ee47df07391e7162684011eb2e0d0063c5bfe54116a256f5

Observation e42b4b21-f613-43c1-8dae-7d5ef8b0ebf5 · outbound

This paper cites Gpt-4v(ision) system card.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Gpt-4v(ision) system card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:10.559586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:10.559586Z digest=sha256:384d1629c13b04393b47905171fe8a650c21ded0040bf177cf4c9d23e9019419

Observation e8d91474-9220-4939-a7d9-feb4bd2b6d51 · outbound

This paper cites an unresolved cited work.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:10.699141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:10.699141Z digest=sha256:1a897e588d70c55a4166ce9a958b469ce8a6ffe52797ea0df8ac10cffa802412

Observation 6cdfce82-46d1-41ca-97ae-3d97c0ffd5e0 · outbound

This paper cites Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl, 2025.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:10.865476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:10.865476Z digest=sha256:1372695a2cbd153bb7a434f0a18fa6e575cc4991a7a74f776fc6fd032375f8a7

Observation 088d5aa3-919b-412f-bf99-72f47975f052 · outbound

This paper cites Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:10.998275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:10.998275Z digest=sha256:203f37ceecf28428ed70663b88656c86c32d3b333b9567fb86f5ec46d4dbd6af

Observation e73181cc-24de-4bc7-a1ef-0206e9c48b12 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:11.170130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:11.170130Z digest=sha256:c31e19581cbe62cf8ced571e4d822a7543d0f0fca4dfe0431f4409520165659f

Observation 824d62e9-3f98-438e-9cc0-bbd5e9de4f34 · outbound

This paper cites Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:11.266165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:11.266165Z digest=sha256:2ad5db0b6b5df743eed7761cd320651c324437ae80ce92724b60522d0e53be2b

Observation c91f2f17-9755-406c-b8f4-543711379eda · outbound

This paper cites Proximal Policy Optimization Algorithms.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Proximal Policy Optimization Algorithms

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:11.390596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:11.390596Z digest=sha256:856d2f7435b5b749fb512f392320ceb80cb30c0c5c1f90d1b3ee92908de34151

Observation 9394e39f-b78c-4e6a-99c2-16d01cbda5d4 · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning, 2024.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:18.961220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:11.571811Z digest=sha256:63a997e86ce7555c1883d687aeb797a54928b95f43f7d20a1e7cd3316088eec3

Observation 2dda3982-92c7-4584-a460-adedfb2bc661 · outbound

This paper cites an unresolved cited work.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:11.757221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:11.757221Z digest=sha256:7ddc1b7e4484e32413177792bb27b1d6a3bcf5aad5e74781aa26c18cf811eead

Observation 4d7b0369-51d0-4b37-92e5-34cc502a2485 · outbound

This paper cites Vlm-r1: A stable and generalizable r1-style large vision-language model, 2025.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Vlm-r1: A stable and generalizable r1-style large vision-language model, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:18.828984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:11.926764Z digest=sha256:71d80f3cd2a826fe415f1023b1e9cb170e9bca88e46be8ee08f7400a609c8c8e

Observation 8b1b5965-1ea6-45b3-ad43-b39e1c67ed30 · outbound

This paper cites Towards vqa models that can read.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Towards vqa models that can read

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:12.110439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:12.110439Z digest=sha256:210cf62c3d4981ef2ff99a3dd3b264f0f23f4d3ced63fd59a81692536c5c2125

Observation 770be980-d70c-4be1-8953-22287adf3c29 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:12.222882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:12.222882Z digest=sha256:bcb450d8bb78a9af875ad2d8aab653b86a4c0656ddb7bf9c89d7ff9e5a71ce22

Observation 0867eecc-0353-42d4-b8f0-7fe6c80e3aed · outbound

This paper cites A Survey of Reasoning with Foundation Models.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought A Survey of Reasoning with Foundation Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:12.352073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:12.352073Z digest=sha256:89eeba264359d88258237ffa84dacf811fcd19e8904ffc573ccfa81fb10d7cf6

Observation 7f1c7586-bf8e-4c45-ba24-25554f6cb3b5 · outbound

This paper cites Mm-verify: Enhancing multimodal reasoning with chain-of-thought verification, 2025.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Mm-verify: Enhancing multimodal reasoning with chain-of-thought verification, 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:12.520560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:12.520560Z digest=sha256:b288bfe87967d170ba72ba589f23a74f9955dd983d3afc96b6a3539d4ce0ce6f

Observation 22ad4768-6ba8-472b-94d3-56213fab3cf9 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:12.592934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:12.592934Z digest=sha256:9c86a36d08b167010476a6c645ea42d1732ab63170f5df3f2300c423958e2435

Observation fdb46ffc-3c0d-443e-9558-9af7b145e9b2 · outbound

This paper cites Llamav-o1: Rethinking step-by-step visual reasoning in llms, 2025.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Llamav-o1: Rethinking step-by-step visual reasoning in llms, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:12.601254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:12.601254Z digest=sha256:f61d5ca7b9c12024e60958dfbfd5ceed859794331186666440ac8db9e1daeb25

Observation 15e7fecf-8aac-42b7-b5da-dc2bc16fffd6 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset, 2024.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Measuring multimodal mathematical reasoning with math-vision dataset, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:18.537503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:12.704054Z digest=sha256:1978981b6df7f6d8ce2e2d7988a84693fda1b6b21ce35ae68ee82987eea7b78a

Observation bb8395e6-d08f-4a8e-a7ff-49fd77e1e389 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:12.845722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:12.845722Z digest=sha256:29c66e1b06d1e9d6bf45ed0c8e57d33543897b6ca60ac285686790bd01cdabdc

Observation 45b885bb-646b-43c7-a571-e45c7916ff37 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models, 2023.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Chain-of-thought prompting elicits reasoning in large language models, 2023

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:12.965174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:12.965174Z digest=sha256:4fe6364737fc0e272efbf7e8cb1ad6badb69906d0684092588913ccb3cd9ea2c

Observation df3c4660-9ecc-48a2-ab0a-29eea0d7cab9 · outbound

This paper cites Boosting multimodal reasoning with mcts-automated structured thinking.arXiv preprint arXiv:2502.02339, 2025.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Boosting multimodal reasoning with mcts-automated structured thinking.arXiv preprint arXiv:2502.02339, 2025

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:13.129782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:13.129782Z digest=sha256:939da84dfee0219e637d4f187ca90fbc5f858ab55a87b74e31b8dd0fc9fd0895

Observation e2fbea10-09db-4734-9a55-a9cbe4d558f2 · outbound

This paper cites Llava-cot: Let vision language models reason step-by-step, 2025.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Llava-cot: Let vision language models reason step-by-step, 2025

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:13.254039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:13.254039Z digest=sha256:7f8b2d814ca5e33d41d11fca70767814834795d74a69325ef7c5ae4b7a8d2fa0

Observation 4420ba65-9895-43a3-9c82-cbc4f9dfc36d · outbound

This paper cites R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization, 2025.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:13.356488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:13.356488Z digest=sha256:96ec39f77e7d1607609887a57e1674dbc71c7da47dcc20f7bc1157de72c11c30

Observation 7c4a7202-d320-4948-bafb-d89e722b0d78 · outbound

This paper cites Mulberry: Empowering mllm with o1-like reasoning and reflection via collective monte carlo tree search, 2024.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Mulberry: Empowering mllm with o1-like reasoning and reflection via collective monte carlo tree search, 2024

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:13.471559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:13.471559Z digest=sha256:c0142e94d1c0049fb29405a9e66435ce6c54a2fbd2053ff167c037a3be0c58d3

Observation fa6a6c8a-356e-40b2-b24f-64276641fa8d · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:13.609155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:13.609155Z digest=sha256:912a2fdfe92b8c6f01b23fecc3b91811989d3df04f0f5cd8e2a65d36d38ebd97

Observation bc02ea5e-e333-4105-9fa9-899492236544 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:13.731642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:13.731642Z digest=sha256:84ec93216c95ff79b6d43d42f397dc4df5d78982a96327b4281fe5d3a63041fc

Observation d15f8ad1-dd4b-4635-be59-43d437fd5024 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:10:17.267328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:13.906483Z digest=sha256:231e4db383e5d685118cd92199c5563df5263b56aa098fd3032d7e7a95d65057

Observation 2907a747-20df-4ae8-b728-efaaae2a8939 · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.Advances in Neural Information Processing Systems, 35:15476–15488, 2022.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Star: Bootstrapping reasoning with reasoning.Advances in Neural Information Processing Systems, 35:15476–15488, 2022

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:14.029007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:14.029007Z digest=sha256:fe3a831f6f61306ad304ad6ebb86ff19d6423164f3000e366f158ea19166b23c

Observation 44fa5f25-11c6-4fad-a0ca-ef63fd42e4b0 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Multimodal Chain-of-Thought Reasoning in Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:14.103607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:14.103607Z digest=sha256:597c75fc8b58c4e80264dba5026861a1bd6baac9f506a2e20084093ea2315193

Observation 35e4a926-a3aa-49b2-b595-c63477df9e60 · outbound

This paper cites Reflection of Thought: Inversely Eliciting Numerical Reasoning in Language Models via Solving Linear Systems.

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought Reflection of Thought: Inversely Eliciting Numerical Reasoning in Language Models via Solving Linear Systems

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:10:14.477257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:10:14.202049Z digest=sha256:2f78fa99883e217b650314fbfe071a32a6dfbfa8b98e795f4b588427239758e6

Pith citing papers

Observation 42c6c1bb-1db9-4ab9-9877-dd4a870d0fb4 · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 192

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.926143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.926143Z digest=sha256:14c1375dba21e660f8f93284e5ca5be53e44e6fcfe9fd15e3c1d6dce6fd969fb

Observation 3eebcfac-f75a-4bb9-ba5d-69145ee7db7c · inbound

DeepEyesV2: Toward Agentic Multimodal Model cites this paper.

DeepEyesV2: Toward Agentic Multimodal Model VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:32:29.579338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T05:32:29.266583Z digest=sha256:d86015ff165c0d18c782b90b124c93e7bb4b2d6bc04642f70914150fcddb3806

Observation 90e92e39-f412-4ebf-90c2-fb05c19a2c0e · inbound

PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models cites this paper.

PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:00:27.234207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T17:57:57.263574Z digest=sha256:f60b206f9e76e41e2541346cf881fd149fabf7dc52492c2134bc4636c0cf4185

Observation 6b6f1835-d10b-4e24-ae05-668478a595b6 · inbound

OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving cites this paper.

OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:38:37.829832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T22:34:00.895252Z digest=sha256:a844b4cf64a66dca247d653ef9d738efb0805698e7e2aedbb36d472f714bc52e

Observation 7bd2ed84-da36-4a67-97fb-fe33e2faaae2 · inbound

Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning cites this paper.

Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T12:31:13.737069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:31:13.737069Z digest=sha256:653e2a8378f39ec5d31c0c869ce53ad085d04ee05b73c5bac7265155d5731186

Observation a120669a-8048-4583-8d48-6b584cac864b · inbound

Imagination Helps Visual Reasoning, But Not Yet in Latent Space cites this paper.

Imagination Helps Visual Reasoning, But Not Yet in Latent Space VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T20:40:30.764822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:40:30.764822Z digest=sha256:9dea3f7cc469e27729f19ccee15ed4cf5ac83247ba571f9bf033e0acf5bea417

Observation d20484e1-aae1-4b27-946d-95ca8299eeeb · inbound

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding cites this paper.

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:08:12.762712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T20:07:23.153064Z digest=sha256:9eaea3ef8aae7cc0f8847e1bd15bac1f3c1a1aa8dfa9568b4145726a41396f96

Observation 697bb1cf-2eea-47d3-9e26-3cff66cbf941 · inbound

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection cites this paper.

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:36:17.630361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T04:46:16.497585Z digest=sha256:c3a47408699899dc35aee3cb399c7c7724ed8e8f29ac3bcb27a24eb5e988ddb1

Observation 8243bba0-66b1-4c28-95a8-0c40596e4190 · inbound

Perceptual Flow Network for Visually Grounded Reasoning cites this paper.

Perceptual Flow Network for Visually Grounded Reasoning VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.910256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:1456998e677a5028ae13d19410f09f17702f72a72d2fa11f49dde1b4526fa2c6

Observation 262f28c4-bf5f-4a80-a590-b54756f9abbb · inbound

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models cites this paper.

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:15.625813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T01:20:54.441367Z digest=sha256:d2a626dad71f34ef6b49a791d82d7a70853f31f5ea89b0bd85a60f8371e6c1ec

Observation 60ec0b92-9072-4c86-9b6b-d5661b39d939 · inbound

When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise cites this paper.

When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:21:09.116990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T17:40:14.089081Z digest=sha256:bb31fb85f19abd3ead7788e14b340e3154503d4949a81b8f679e691ae00628a6

Observation 3f53fdc3-3fd8-4364-a2ed-1beb43ac3a8d · inbound

When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise cites this paper.

When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:51:29.288270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T03:52:44.531016Z digest=sha256:147961c5d7d74fa06b59979949fdc03f7fafe0de0f334af3ba705c4c4fd15fa3

Observation ed2eb08a-8cf8-4e43-a016-e37db3fc75c4 · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:29.145840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T07:14:48.918959Z digest=sha256:2d4972a0c0176b280036ab538065377e1018f24d350acc7d48be452fab46fe3e

Observation 09482764-0ac1-4214-b55c-2cbba6b5b528 · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:48:00.615749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T21:47:50.595481Z digest=sha256:d39367317e26fae75c42a1eb69a4ffed17a8b31360989bfb32da2a569fffce32

Observation 9f1803a0-0cc1-4b0c-bd37-dce56e429f34 · inbound

Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning cites this paper.

Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:29:52.008722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T05:08:09.884231Z digest=sha256:eac2a54bda57121566130331f9cedfabd883395d9fec5b6473ae14595831695a

Observation 00ebaf5f-2631-4fd4-a108-c12080a2a6ef · inbound

How LLMs See Creativity: Zero-Shot Scoring of Visual Creativity with Interpretable Reasoning cites this paper.

How LLMs See Creativity: Zero-Shot Scoring of Visual Creativity with Interpretable Reasoning VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:20.015685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T06:50:08.686114Z digest=sha256:ec04b54866becfeeaf8321b981fa4aa947b4a710700348b4c64da25a141113a9

Observation 6595ff6c-e86a-42fe-aaa4-b29ab7409254 · inbound

How LLMs See Creativity: Zero-Shot Scoring of Visual Creativity with Interpretable Reasoning cites this paper.

How LLMs See Creativity: Zero-Shot Scoring of Visual Creativity with Interpretable Reasoning VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:15:28.984824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-01T07:11:29.002616Z digest=sha256:8d56290c7ce44a03b77c6f8abe87f3587a8258d548e8f8b04b6ca736b964e076

Observation b360c749-1614-45a1-8676-f81404cd65e4 · inbound

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing cites this paper.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:25.996486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:25.996486Z digest=sha256:776e8dda20779ab3d923b5f3f5eb7948d601b2153bab853a5ffbdff347e24770

Observation 0608c6b9-0449-4965-9822-d85057b798a4 · inbound

OPLD: On-Policy Latent Distillation for Multimodal Reasoning cites this paper.

OPLD: On-Policy Latent Distillation for Multimodal Reasoning VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T16:13:17.400413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T16:13:17.400413Z digest=sha256:d53fa0a13a2b6394bfcd8f2ecc79e556e6ba99df82c7ca9fd2c369e75a770e87

Observation 81ef07af-f41b-4322-ad2f-254890b8419c · inbound

InSight-doc: Agentic Visual Perception for Long-Document Understanding cites this paper.

InSight-doc: Agentic Visual Perception for Long-Document Understanding VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 146

Resolution
unresolved
no resolver link, observed 2026-08-12T20:43:53.137356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:43:53.137356Z digest=sha256:455f05e8e565eb1f3ac6f596e31538dca2d05b2215ffc9d4ac1a4745a1ebcb9b