Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T13:29:47.529564Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 34 inbound Pith citation observations for arXiv:2505.15879.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T13:29:47.529564Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T05:31:41.728219Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T03:17:51.507547Z
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 53021e62-d29b-4050-b407-e80fd280bd56 · outbound
GRIT: Teaching MLLMs to Think with Images Introducing openai o1-preview
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 63bce2bf-3454-4dab-9d1c-9c5edb01d2c5 · outbound
GRIT: Teaching MLLMs to Think with Images DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6074d927-0032-435b-b456-b504044db977 · outbound
GRIT: Teaching MLLMs to Think with Images Qwen2.5 Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bbf9e3c2-9f00-4b18-8176-8d9ef96eb906 · outbound
GRIT: Teaching MLLMs to Think with Images Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c391cfeb-9da5-46ce-a8c5-6f1e1022a39e · outbound
GRIT: Teaching MLLMs to Think with Images The Llama 3 Herd of Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 62cf520a-1a8a-4664-88ca-11b6e63de363 · outbound
GRIT: Teaching MLLMs to Think with Images Chain-of-thought prompting elicits reasoning in large language models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ec736a2d-336a-4a22-9cae-61fe2c786e0b · outbound
GRIT: Teaching MLLMs to Think with Images Reasoning models don’t always say what they think
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b2a39bc5-e1b9-4c1c-b816-3a72772522de · outbound
GRIT: Teaching MLLMs to Think with Images Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d578c0e4-75b8-4e2c-90be-49bfec98591d · outbound
GRIT: Teaching MLLMs to Think with Images Training Large Language Models to Reason in a Continuous Latent Space
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3dfb7bd1-d7df-41ee-85bf-a5370b257c62 · outbound
GRIT: Teaching MLLMs to Think with Images VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6856c958-98f0-4f8c-94a5-19c6668a4ff3 · outbound
GRIT: Teaching MLLMs to Think with Images Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fca9c436-30b5-427d-9a4c-14fb285614e1 · outbound
GRIT: Teaching MLLMs to Think with Images R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0994b7bf-7d19-4bf0-8e6a-0ce65e4607dc · outbound
GRIT: Teaching MLLMs to Think with Images InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f8ee944a-7da6-49af-8e04-5d328533c36a · outbound
GRIT: Teaching MLLMs to Think with Images Visual spatial reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3fc15d87-0cb2-4b55-bbe2-70ea3dd3f003 · outbound
GRIT: Teaching MLLMs to Think with Images Tallyqa: Answering complex counting questions
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation abd39a97-7d2f-4d16-98cf-d2c9f9c75a3d · outbound
GRIT: Teaching MLLMs to Think with Images R1-v: Reinforcing super generaliza- tion ability in vision-language models with less than $3.https://github.com/Deep-Agent/ R1-V
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5669a613-aa98-4d70-9955-d87363bd152e · outbound
GRIT: Teaching MLLMs to Think with Images Introducing openai o3 and o4-mini
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3baa21e4-711b-4dad-800a-1daeb7ca1339 · outbound
GRIT: Teaching MLLMs to Think with Images Self-consistency improves chain of thought reasoning in language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1637b3ad-0836-4ee4-a4e2-097fff7616fc · outbound
GRIT: Teaching MLLMs to Think with Images Multimodal Chain-of-Thought Reasoning in Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1f66ac9c-5bec-4b6f-a6a6-130f019ef756 · outbound
GRIT: Teaching MLLMs to Think with Images Visual chain-of-thought prompting for knowledge-based visual reasoning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6ba154f8-3fae-4248-b5f6-8935ec9c9219 · outbound
GRIT: Teaching MLLMs to Think with Images Compositional chain-of- thought prompting for large multimodal models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 23d04d59-7f2a-437a-8a1e-806344182d1c · outbound
GRIT: Teaching MLLMs to Think with Images Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 25a769e3-2dd9-43f7-86f3-a0c6b7f535b2 · outbound
GRIT: Teaching MLLMs to Think with Images Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 04ded7dc-4405-40b8-a22d-83dbb6eb014a · outbound
GRIT: Teaching MLLMs to Think with Images DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 57b627d1-387d-4273-a040-2f30f4199571 · outbound
GRIT: Teaching MLLMs to Think with Images Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5e7b5baf-aae9-4a4d-a81f-9470be0b2def · outbound
GRIT: Teaching MLLMs to Think with Images Mme: A comprehensive evaluation benchmark for multimodal large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 497d4ed4-be8a-4635-ae1a-43e62397c992 · outbound
GRIT: Teaching MLLMs to Think with Images MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 45712bc3-ddea-46cb-8d18-992d0bc06afc · outbound
GRIT: Teaching MLLMs to Think with Images How to Evaluate the Generalization of Detection? A Benchmark for Comprehensive Open-Vocabulary Detection
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d61dcef2-1214-43f4-b902-98e7ea7f22bf · outbound
GRIT: Teaching MLLMs to Think with Images Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e8c31818-4823-4ef0-8fe6-ee62462f9c88 · outbound
GRIT: Teaching MLLMs to Think with Images Language Models are Few-Shot Learners
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f1440bbd-83fc-4df3-9e49-8a0c4355dd24 · outbound
GRIT: Teaching MLLMs to Think with Images Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bf26a524-4fa1-4caa-ae18-5990dab4cdd4 · outbound
GRIT: Teaching MLLMs to Think with Images Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4ef8c58e-78f1-4d42-ad0b-608a17815f64 · outbound
GRIT: Teaching MLLMs to Think with Images OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6388e0e0-b733-4138-9ed3-b930cdf92b31 · outbound
GRIT: Teaching MLLMs to Think with Images SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 05b20e05-7c95-4719-84c9-a843862c2d35 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey GRIT: Teaching MLLMs to Think with Images
Reference 246
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8f2e4bf3-41ee-4dfb-9f20-333b58d1e1ed · inbound
A Survey of Reinforcement Learning for Large Reasoning Models GRIT: Teaching MLLMs to Think with Images
Reference 133
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2165189f-3d1f-4442-9437-75c4bef659ef · inbound
DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning GRIT: Teaching MLLMs to Think with Images
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9decbf5c-c00d-4215-9bd4-540abe876579 · inbound
DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning GRIT: Teaching MLLMs to Think with Images
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9c9dc810-3210-4559-aafd-a0dafc2aa787 · inbound
Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning GRIT: Teaching MLLMs to Think with Images
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8b981fc6-76b5-4a4c-a00b-8a8519489161 · inbound
Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images GRIT: Teaching MLLMs to Think with Images
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8243c96f-1865-430b-97fa-0f15d800c5b2 · inbound
DeepEyesV2: Toward Agentic Multimodal Model GRIT: Teaching MLLMs to Think with Images
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2c11580c-89a5-4c9b-a428-0d40fa656a99 · inbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents GRIT: Teaching MLLMs to Think with Images
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cad19910-7fe8-4123-a681-b87d362055f9 · inbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents GRIT: Teaching MLLMs to Think with Images
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7a5c95e-ca17-4949-bfcf-e30dd4b45186 · inbound
Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space GRIT: Teaching MLLMs to Think with Images
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 09682d13-051c-4bff-9f05-6eeffeabaf4a · inbound
MentisOculi: Revealing the Limits of Reasoning with Mental Imagery GRIT: Teaching MLLMs to Think with Images
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3983524-ce36-4706-9ef6-481081bf2eab · inbound
CharTool: Tool-Integrated Visual Reasoning for Chart Understanding GRIT: Teaching MLLMs to Think with Images
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2153c1ed-3c36-4215-9068-cd12049a8599 · inbound
Discovering Failure Modes in Vision-Language Models using RL GRIT: Teaching MLLMs to Think with Images
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9d39e6dc-9789-4e6a-b8d1-72d728f27bfd · inbound
Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning GRIT: Teaching MLLMs to Think with Images
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d6277388-5c56-4714-886c-2f73cd00da37 · inbound
Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images GRIT: Teaching MLLMs to Think with Images
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9871af4a-fca1-4444-bf8f-d4faaf0f86f8 · inbound
Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images GRIT: Teaching MLLMs to Think with Images
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2acb4af-17dc-4eab-ab7c-239ce84866ee · inbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs GRIT: Teaching MLLMs to Think with Images
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eec7037c-3b7d-4415-9e17-781290b83c91 · inbound
Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs GRIT: Teaching MLLMs to Think with Images
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 282fc9d3-ed0e-4203-9244-9bfbfc1e3a45 · inbound
CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models GRIT: Teaching MLLMs to Think with Images
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b110cb73-25d9-4b4d-899a-6d4b6e567fbb · inbound
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming GRIT: Teaching MLLMs to Think with Images
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 01417436-c34e-4ec4-8566-57ecca9d9828 · inbound
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming GRIT: Teaching MLLMs to Think with Images
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2f17d4ff-7439-4adc-b822-908ea9db3231 · inbound
Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning GRIT: Teaching MLLMs to Think with Images
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6e4cb97b-624b-43f8-ac6b-4f2b5583ead8 · inbound
Self-Prophetic Decoding to Unlock Visual Search in LVLMs GRIT: Teaching MLLMs to Think with Images
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d771a75e-8bff-4560-9352-0dabae6d2137 · inbound
Trust Region On-Policy Distillation GRIT: Teaching MLLMs to Think with Images
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c854a493-ce9f-4840-8c0a-4fd43b61c188 · inbound
UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA GRIT: Teaching MLLMs to Think with Images
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a11129eb-b7b5-43b5-8421-b2d0e1606be4 · inbound
An LMM for Precisely Grounding Elements in Documents GRIT: Teaching MLLMs to Think with Images
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dabee48d-3830-4dc2-8591-6ae2e5417b62 · inbound
V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning GRIT: Teaching MLLMs to Think with Images
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2ebcddd0-050f-4521-81ad-2fda7b48cd47 · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models GRIT: Teaching MLLMs to Think with Images
Reference 177
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ed133d53-aeb3-45e2-9415-f99e1dfba12a · inbound
Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models GRIT: Teaching MLLMs to Think with Images
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fb371fa9-c51b-432c-b8b9-6e9e36278253 · inbound
OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping GRIT: Teaching MLLMs to Think with Images
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 96380a47-4cd1-4d20-b38c-26b01ec491ec · inbound
Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning GRIT: Teaching MLLMs to Think with Images
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 017bab76-2cb8-4f60-864c-553c7b794daf · inbound
Transferability Between Understanding and Generation in Unified Multimodal Models GRIT: Teaching MLLMs to Think with Images
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f8d0cdb-a77f-47aa-a1d8-7b4799e5e5af · inbound
Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models GRIT: Teaching MLLMs to Think with Images
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8c79a5ae-e61f-4544-b33e-d96d7766b491 · inbound
Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models GRIT: Teaching MLLMs to Think with Images
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.