Pith. sign in

Paper Citation Record · LEDGER

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators

As of 2 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 3 inbound Pith citation observations for arXiv:2604.03307.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.03307 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T23:57:47.657243Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T06:45:27.857034Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T13:56:19.173208Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact26
  • verified fuzzy13
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56382967-f85d-4383-bfc7-4f1fdcae5cc3 · outbound

This paper cites Qwen Technical Report.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Qwen Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.490776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:b0b1b29d2ad61b7a1418975fba1f2ec0f8d0cb8c6fbe392d28d5304cf6d83157

Observation bd5e5be2-8470-4c5c-822d-3b59970544e4 · outbound

This paper cites Qwen3-VL Technical Report.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.621212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:7702d302d908197044ae098013f2672cae3aa5076cd96e31e984cdee61555322

Observation 9da4abac-eacf-4abb-8323-84cf38282d8c · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.519794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:3274f4c56509e055af2748d3358a518dc4a41272d7999aa0ddbf443f158a9864

Observation 2c1c0388-0b5c-4c06-b305-3f1797da3633 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic 11 visual-linguistic tasks.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Internvl: Scaling up vision foundation models and aligning for generic 11 visual-linguistic tasks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.519461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:a573ed5d51b30013afa19be9f6b4b2e0c3d21c203510c6b2c48dcc37bbac96f0

Observation e77e0b81-a7ff-485c-a4ad-ea6dc231f864 · outbound

This paper cites Compressed Chain of Thought: Efficient Reasoning Through Dense Representations.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Compressed Chain of Thought: Efficient Reasoning Through Dense Representations

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:47:40.470022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:32fd3a6c83bc4c7d6f2ea07fd92c874a5e423a65180ff78420b55daf938fc0de

Observation 605c65a9-2d20-43ba-beb3-2ee3961beb76 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Emerging Properties in Unified Multimodal Pretraining

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.578934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:7b5a8cd39d30ea79394c9c3a3ec19205bcf3f10d4f1e9e8fceb9df50db4177e6

Observation 6a5d5814-3ba6-443a-8b76-c8c769a3f8f9 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Blink: Multimodal large language models can see but not perceive

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.507759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:3576df444e08069bf57d2d938a7a1da97cc0cdb2b267d7ee22682917a9f06c93

Observation f54922e4-55e5-4b2a-ace6-55795e4ea7d0 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Training Large Language Models to Reason in a Continuous Latent Space

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.634771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:70dc1feb8c9ed6620715b52cebd9ca697457efc146926cab7cf17ec7c0b3908e

Observation 5aef65bd-63e6-4570-85c2-1007937c2f20 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.536112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:3aa4a4e29337fca92328ab46c5048727c4fb7f0f75c7d94c3f048bdd1ed26780

Observation 97d476b3-748c-4412-8e97-f641cc46561a · outbound

This paper cites GPT-4o System Card.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators GPT-4o System Card

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.430718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:40499a3183275149ad4d0ec3eb24f758116d4c8acfd23de8d2abe8cebf7f3ab1

Observation 1ef4b339-5566-48bd-bcdd-e477fa0ed9ea · outbound

This paper cites Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:58:28.440054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:a054bba215097d14883579610506a3cd3631f543b641982670658e94448c8ef2

Observation 1fe3b71d-7ac3-4323-867d-fa0b304081b6 · outbound

This paper cites Latent Visual Reasoning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Latent Visual Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:41:30.500607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:6210e01f7278c91497b1349adf772077b1ac5127f98e4250a97b715df16dc3cc

Observation 59493c6e-3948-4dd4-99fc-9ad9a9619ef9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators LLaVA-OneVision: Easy Visual Task Transfer

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.627283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:b2f75a3aa09ba82036fe894298e95a65411850e50c34821945a71f9b7571bc9a

Observation c7782cb0-5b4e-40a7-b348-15a634e7e941 · outbound

This paper cites Visual-rft: Visual reinforcement fine-tuning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Visual-rft: Visual reinforcement fine-tuning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.576437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:baf28a7f33c4584895e23a5174c77e49a804bc6e1b8a3a81161ad7e0b673192d

Observation 8e5f7f26-dd15-4346-9226-aca42d9c3d33 · outbound

This paper cites Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:58:28.614087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:81574d440d2c789d92b0579945d7bde2f6de241acd2701afce025b860b30919c

Observation 4ae5fe2d-3889-4424-8438-20a3013dad52 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:15:46.589538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:dc4228a6dfa8bce0b8fb59cacf4f2ebaf01712ddcc736b88f2353ef7d5759e56

Observation fe43b6ab-f6b0-41d2-9969-3d02a9f9dab1 · outbound

This paper cites Chain-of-visual-thought: Teaching vlms to see and think better with continuous visual tokens.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Chain-of-visual-thought: Teaching vlms to see and think better with continuous visual tokens

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:58:28.514256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:f3d03cb279860cd3e79e0309c45118d5f3fb5a24bcb42c8badc5b34830454df9

Observation 498e3f9d-8680-4b54-a970-898071c42b21 · outbound

This paper cites an unresolved cited work.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-13T23:58:29.556548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:4f68f6368b0f6d92d7fc2102a2ee9f5a8d30c3ff6a98398311586e774c53f1b5

Observation fecd0122-ba95-4896-9037-0a37a5c66983 · outbound

This paper cites Codi: Compressing chain-of-thought into continuous space via self-distillation.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Codi: Compressing chain-of-thought into continuous space via self-distillation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.561411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:1e5832baf755a0185926bd500d04e6f707923eac96848fa3992bbff0f4c0be01

Observation cd2c0691-27e4-4e05-9179-be2baea0d56a · outbound

This paper cites OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:13:13.788897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:d91176fdb95e17c946745719d97685c2336311087e34af761885a5e1432fa8ba

Observation 005aa765-86f9-47e7-9a9d-25355ef293c7 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv e-prints, pages arXiv–2503.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv e-prints, pages arXiv–2503

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.567637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:f9c5661266b18bd04d9a6aa39a16c3a9f53386836f8946db48c536883bc7c9d7

Observation 4d3a1248-6fc4-4d6b-b1a0-49d4f7afe4dd · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.515196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:8df99b2869f8fc83aa2e22b13e46dbbd6cff08c207ad52fabafe540a2318c885

Observation b4775fcf-8cfb-4ec1-b317-b1c75584d709 · outbound

This paper cites Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:22:27.090901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:ca471a0d47336d3408e2a730f7c3e17c668e44828a6b4230d3ce234d4acfc69c

Observation c06d79e1-8957-4c62-a998-f2bd30d29bfd · outbound

This paper cites Monet: Reasoning in latent visual space beyond images and language.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Monet: Reasoning in latent visual space beyond images and language

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:58:28.543353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:3466e32acc00c319d3a7f27a0b6652d24f181344bf359208e2cc09fc07539cd5

Observation 6285fb97-d596-4582-9ab9-a50195a27335 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.600076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:3cc308e2230028a4820124f9fb388439500246f7bb21d47969df9360e99f127c

Observation 1a511570-42be-4fb4-9955-2269a7bc8953 · outbound

This paper cites Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.544943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:d501cf2ccada353ecc8975ed4b82d5067da82d5ea41e4b115c40a31a6acf7281

Observation 64d0fd5c-6c9a-43a2-8266-583407ae2783 · outbound

This paper cites Perception-Aware Policy Optimization for Multimodal Reasoning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.528420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:f0ea04de7f0a19741105261624cfb8fa6c58b294866a60a6de72d3d5bcfdbba7

Observation b2c3769f-d601-4f3d-8b57-f678d12c6beb · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.549587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:31cf34b5f4f76c41176dd1814b5fb50b90fda3a4a4fafca50cf29efe1504a28b

Observation aa29c6f0-164e-43a3-af49-7f78e3ef4753 · outbound

This paper cites Vtool-r1: Vlms learn to think with images via reinforcement learning on multimodal tool use.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Vtool-r1: Vlms learn to think with images via reinforcement learning on multimodal tool use

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:58:28.593993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:371454215096ebc21119e631f7143acc5abb40f8708fe22960a456eb0b320108

Observation 4a31d6bd-3e61-4c0b-abaf-38c030a5807a · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators V?: Guided visual search as a core mechanism in multimodal llms

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.533713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:c015fb892ea37a3c0d16d4b89bba2f9746ccddb95a7d8be2ea32b87794969850

Observation 54428619-bef9-4b63-b9af-7efeee0f72b9 · outbound

This paper cites Llava-cot: Let vision language models reason step-by-step.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Llava-cot: Let vision language models reason step-by-step

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.538739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:6f77499173882d68b28de672aaa58aaf6a9fcde60ee1c4329100b9ccc29be3f4

Observation 2640e6fe-3f95-44d8-9628-0d999270d5c9 · outbound

This paper cites Mc-bench: A benchmark for multi-context visual grounding in the era of mllms.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Mc-bench: A benchmark for multi-context visual grounding in the era of mllms

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.523946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:439a59a2596afc67bd1964c389e004e28d26398d37a0abd605d5c70f97004b54

Observation ce1f991b-3bb5-4683-84af-b9431f06f8b1 · outbound

This paper cites R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.528009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:353187f92a492bcca5cdbd5dac941b48c239d84740ad0edb39354bcbd30e1dfe

Observation 570e609f-3dc3-41ae-9f50-edcad274807a · outbound

This paper cites Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:58:28.561653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:4de467ad566e6f9752c02388943cb043e024a848796f66cd987f6ea7b06eb572

Observation 83b88eaf-9ed7-4802-aeb5-e0b6c180f8db · outbound

This paper cites Perception-R1: Pioneering Perception Policy with Reinforcement Learning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Perception-R1: Pioneering Perception Policy with Reinforcement Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:58:28.586668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:bc2925e26cce52463b2aacd462d2d934625d3c388a6c1c45b5ef194839dc124a

Observation 35f11d4c-8ca6-486e-b4ba-a6a9c0554951 · outbound

This paper cites Chain-of-focus: Adaptive visual search and zooming for multimodal reasoning via rl.arXiv e-prints, pages arXiv–2505.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Chain-of-focus: Adaptive visual search and zooming for multimodal reasoning via rl.arXiv e-prints, pages arXiv–2505

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T23:58:29.572747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:7bc5d2f664e3909712fc5f7a0cae4238b5fe925eba587a28f1a248180fa444e3

Observation 2f98649b-1879-493f-8174-555efd4cfe0a · outbound

This paper cites Thyme: Think Beyond Images.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Thyme: Think Beyond Images

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:33:29.495992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:b622f2a466740ce7c3bbb7494dff98314c4397f2d7eae6953af24187ef2bd046

Observation aeba2d28-612a-4e41-8a8d-ec6c810ed36b · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.958879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:8169276a436da75f5bb0064f37f0795fc5eea38a8358f12b9ee5a0153380e83a

Observation 4695c34f-1a6e-4728-be36-9b5ebaee0ced · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.607162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:7615da3e97a5d76e78f0430382f8a1a7a5546c600f8ce9f8b8643c84c58659db

Observation e0907d01-1ba7-4334-a8b7-d8447b7d2acf · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.497797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:ac80af093d205ff08d7c9660e71c455ee148f57b68f4df5920f00a384655c690

Pith citing papers

Observation d9cbee03-fca8-467c-b2e4-33c98b9a092a · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators

Reference 86

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T20:22:37.724913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:b5846574d8d1c4b50a87eb7383d66b6d3bd61b903f29794dbe75c64a2aeaabfa

Observation c2ddf2ac-a0d9-4960-9452-93d250756ea7 · inbound

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning cites this paper.

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators

Reference 98

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T13:56:19.174675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-07-09T13:51:49.149342Z digest=sha256:b8f19dc743a21353f0f0da87d121000489c384f94dbb0f92d3ece12f03e230ad

Observation 0276d545-cc89-4624-9e14-bdd0a3a33304 · inbound

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning cites this paper.

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators

Reference 98

Resolution
unresolved
no resolver link, observed 2026-07-13T06:45:27.857034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T06:45:27.857034Z digest=sha256:a43fc1dbdc13ff00c0a5c76a45163ed750c73066b3357f6224cfe1b6c5fcde0e