Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T23:57:47.657243Z
Paper Citation Record · LEDGER
As of 2 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 3 inbound Pith citation observations for arXiv:2604.03307.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T23:57:47.657243Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-13T06:45:27.857034Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T13:56:19.173208Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 56382967-f85d-4383-bfc7-4f1fdcae5cc3 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Qwen Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bd5e5be2-8470-4c5c-822d-3b59970544e4 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 9da4abac-eacf-4abb-8323-84cf38282d8c · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2c1c0388-0b5c-4c06-b305-3f1797da3633 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Internvl: Scaling up vision foundation models and aligning for generic 11 visual-linguistic tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e77e0b81-a7ff-485c-a4ad-ea6dc231f864 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Compressed Chain of Thought: Efficient Reasoning Through Dense Representations
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 605c65a9-2d20-43ba-beb3-2ee3961beb76 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Emerging Properties in Unified Multimodal Pretraining
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 6a5d5814-3ba6-443a-8b76-c8c769a3f8f9 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Blink: Multimodal large language models can see but not perceive
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f54922e4-55e5-4b2a-ace6-55795e4ea7d0 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Training Large Language Models to Reason in a Continuous Latent Space
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5aef65bd-63e6-4570-85c2-1007937c2f20 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 97d476b3-748c-4412-8e97-f641cc46561a · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators GPT-4o System Card
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1ef4b339-5566-48bd-bcdd-e477fa0ed9ea · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1fe3b71d-7ac3-4323-867d-fa0b304081b6 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Latent Visual Reasoning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 59493c6e-3948-4dd4-99fc-9ad9a9619ef9 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators LLaVA-OneVision: Easy Visual Task Transfer
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation c7782cb0-5b4e-40a7-b348-15a634e7e941 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Visual-rft: Visual reinforcement fine-tuning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8e5f7f26-dd15-4346-9226-aca42d9c3d33 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 4ae5fe2d-3889-4424-8438-20a3013dad52 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation fe43b6ab-f6b0-41d2-9969-3d02a9f9dab1 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Chain-of-visual-thought: Teaching vlms to see and think better with continuous visual tokens
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 498e3f9d-8680-4b54-a970-898071c42b21 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation fecd0122-ba95-4896-9037-0a37a5c66983 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Codi: Compressing chain-of-thought into continuous space via self-distillation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation cd2c0691-27e4-4e05-9179-be2baea0d56a · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 005aa765-86f9-47e7-9a9d-25355ef293c7 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv e-prints, pages arXiv–2503
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 4d3a1248-6fc4-4d6b-b1a0-49d4f7afe4dd · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b4775fcf-8cfb-4ec1-b317-b1c75584d709 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation c06d79e1-8957-4c62-a998-f2bd30d29bfd · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Monet: Reasoning in latent visual space beyond images and language
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 6285fb97-d596-4582-9ab9-a50195a27335 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1a511570-42be-4fb4-9955-2269a7bc8953 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 64d0fd5c-6c9a-43a2-8266-583407ae2783 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Perception-Aware Policy Optimization for Multimodal Reasoning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b2c3769f-d601-4f3d-8b57-f678d12c6beb · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation aa29c6f0-164e-43a3-af49-7f78e3ef4753 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Vtool-r1: Vlms learn to think with images via reinforcement learning on multimodal tool use
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 4a31d6bd-3e61-4c0b-abaf-38c030a5807a · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators V?: Guided visual search as a core mechanism in multimodal llms
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 54428619-bef9-4b63-b9af-7efeee0f72b9 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Llava-cot: Let vision language models reason step-by-step
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2640e6fe-3f95-44d8-9628-0d999270d5c9 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Mc-bench: A benchmark for multi-context visual grounding in the era of mllms
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation ce1f991b-3bb5-4683-84af-b9431f06f8b1 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 570e609f-3dc3-41ae-9f50-edcad274807a · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 83b88eaf-9ed7-4802-aeb5-e0b6c180f8db · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Perception-R1: Pioneering Perception Policy with Reinforcement Learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 35f11d4c-8ca6-486e-b4ba-a6a9c0554951 · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Chain-of-focus: Adaptive visual search and zooming for multimodal reasoning via rl.arXiv e-prints, pages arXiv–2505
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2f98649b-1879-493f-8174-555efd4cfe0a · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Thyme: Think Beyond Images
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation aeba2d28-612a-4e41-8a8d-ec6c810ed36b · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 4695c34f-1a6e-4728-be36-9b5ebaee0ced · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e0907d01-1ba7-4334-a8b7-d8447b7d2acf · outbound
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d9cbee03-fca8-467c-b2e4-33c98b9a092a · inbound
DeepLatent: Think with Images via Parallel Latent Visual Reasoning V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation c2ddf2ac-a0d9-4960-9452-93d250756ea7 · inbound
BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 0276d545-cc89-4624-9e14-bdd0a3a33304 · inbound
BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.