Pith. sign in

Paper Citation Record · LEDGER

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2505.16192.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16192 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:31:13.737069Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3eebcfac-f75a-4bb9-ba5d-69145ee7db7c · inbound

DeepEyesV2: Toward Agentic Multimodal Model cites this paper.

DeepEyesV2: Toward Agentic Multimodal Model VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:32:29.579338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:32:29.266583Z digest=sha256:a1e12931775af414ca815f3b19247af9788bfc8cb1a65151d174193274828d7e

Observation 90e92e39-f412-4ebf-90c2-fb05c19a2c0e · inbound

PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models cites this paper.

PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:00:27.234207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T17:57:57.263574Z digest=sha256:ac3008c32b5360d23f49c9c107df5ab801130040142a7ec9a180a6148a51c135

Observation 6b6f1835-d10b-4e24-ae05-668478a595b6 · inbound

OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving cites this paper.

OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:38:37.829832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T22:34:00.895252Z digest=sha256:a2758534570a712123b27e1433c1a0518b8e3e838a6cb1c153f4237c375b144c

Observation 7bd2ed84-da36-4a67-97fb-fe33e2faaae2 · inbound

Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning cites this paper.

Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T12:31:13.737069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:31:13.737069Z digest=sha256:1db106726b3655c0c7b4af8b404c5cc40b180c84ec1b4222a3d3d292c14ed055

Observation a120669a-8048-4583-8d48-6b584cac864b · inbound

Imagination Helps Visual Reasoning, But Not Yet in Latent Space cites this paper.

Imagination Helps Visual Reasoning, But Not Yet in Latent Space VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T20:40:30.764822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:40:30.764822Z digest=sha256:ef66b529505ccc6f3977cdcf8460c570fe99d6f5b4bba3204f3de00be63108a3

Observation d20484e1-aae1-4b27-946d-95ca8299eeeb · inbound

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding cites this paper.

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:08:12.762712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T20:07:23.153064Z digest=sha256:b34612d7dfedcb0941384cd5304c05d4b8e03101116b949a47574dcc8f957fec

Observation 697bb1cf-2eea-47d3-9e26-3cff66cbf941 · inbound

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection cites this paper.

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:36:17.630361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T04:46:16.497585Z digest=sha256:8614bd5e53f7925095c82dcd1b08f15db41ae9a3e2d1db335eee22b1ec97d46a

Observation 8243bba0-66b1-4c28-95a8-0c40596e4190 · inbound

Perceptual Flow Network for Visually Grounded Reasoning cites this paper.

Perceptual Flow Network for Visually Grounded Reasoning VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.910256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:2953d9d7ae301286d51f9df78917cb277f13300a986b9072c8b688b5c4fd43af

Observation 262f28c4-bf5f-4a80-a590-b54756f9abbb · inbound

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models cites this paper.

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:15.625813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T01:20:54.441367Z digest=sha256:9fb59036d6250433f13f978e4ff2f826dda4c73572504ea2427b4a2c45d03aee

Observation 60ec0b92-9072-4c86-9b6b-d5661b39d939 · inbound

When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise cites this paper.

When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:21:09.116990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T17:40:14.089081Z digest=sha256:87d5d2130b30d9ed267a6185a7631d0e6c83acae400b6b52cf9447127c103e78

Observation 3f53fdc3-3fd8-4364-a2ed-1beb43ac3a8d · inbound

When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise cites this paper.

When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:51:29.288270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:52:44.531016Z digest=sha256:71be02c838c67b9cbe80d09fb1c38f56c9efaccba68ad8788999f4e0e859622d

Observation ed2eb08a-8cf8-4e43-a016-e37db3fc75c4 · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:29.145840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:14:48.918959Z digest=sha256:53a707b46d5712a3f3c1b625237db49baab8081bfd96dd8245c5e9bdcf602bfb

Observation 09482764-0ac1-4214-b55c-2cbba6b5b528 · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:48:00.615749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:47:50.595481Z digest=sha256:3739a1eb92e9ecbadf98bd5bd102df7ee91e83bc4ef050fa9da7360dff67bf1d

Observation 9f1803a0-0cc1-4b0c-bd37-dce56e429f34 · inbound

Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning cites this paper.

Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:29:52.008722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T05:08:09.884231Z digest=sha256:313880227b003e105dd5ba729caa6d5d7806467ebc41f732a2f6e853171c0352

Observation 00ebaf5f-2631-4fd4-a108-c12080a2a6ef · inbound

How LLMs See Creativity: Zero-Shot Scoring of Visual Creativity with Interpretable Reasoning cites this paper.

How LLMs See Creativity: Zero-Shot Scoring of Visual Creativity with Interpretable Reasoning VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:20.015685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T06:50:08.686114Z digest=sha256:9b00bafa888bac91f57964075516bdf1a42ae5d23966eb19cc390caeb9a89a01

Observation 6595ff6c-e86a-42fe-aaa4-b29ab7409254 · inbound

How LLMs See Creativity: Zero-Shot Scoring of Visual Creativity with Interpretable Reasoning cites this paper.

How LLMs See Creativity: Zero-Shot Scoring of Visual Creativity with Interpretable Reasoning VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:15:28.984824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T07:11:29.002616Z digest=sha256:f9293503fcb81383b439c9d56f24bf74ccbd434b791833fced5b9db630803233

Observation b360c749-1614-45a1-8676-f81404cd65e4 · inbound

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing cites this paper.

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T00:59:25.996486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:59:25.996486Z digest=sha256:b38e37408d5e5482f940bd7cbe750ea3cf273c4d33f3b421d2062f613f81a0ed

Observation 0608c6b9-0449-4965-9822-d85057b798a4 · inbound

OPLD: On-Policy Latent Distillation for Multimodal Reasoning cites this paper.

OPLD: On-Policy Latent Distillation for Multimodal Reasoning VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T16:13:17.400413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T16:13:17.400413Z digest=sha256:a418651872c35a161115a5985b78351e7b57969f0c9094f041a9e389dba27250