Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T01:13:57.368874Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 100 inbound Pith citation observations for arXiv:2504.07615.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T01:13:57.368874Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:03:31.357707Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T03:17:51.904109Z
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 23851756-2cf5-4645-ac97-4be32c9f3c81 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model https://github.com/hiyouga/ EasyR1
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d876a72b-1b40-4db6-8a81-ded882ac0cb4 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0d80caaa-4d7e-48e9-9b76-71be48d9def9 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d722ba59-8902-446f-afaf-86f2f6ff55fe · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Flamingo: a visual language model for few-shot learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1f561667-281c-4ced-a6f5-e96af9111c57 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Concrete Problems in AI Safety
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation befc5d16-58e5-4381-baa3-9fde0d0f54ea · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Qwen2.5-VL Technical Report
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 54ded20d-c61f-4f0d-afe0-acdd9efd7b58 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5cb3c842-f840-4ff6-a274-05f44a0ad88f · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6ce284a4-bdfa-49c4-abff-8f1609d72718 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model R1-v: Reinforcing super generalization ability in vision- language models with less than $3
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f97b9fff-840d-4c74-a7a3-fb38cd2134f6 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b6244dcf-3734-4ed9-a83f-90d621649de1 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 59d81237-41e9-41ce-ba6b-7ef537ba1516 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 60aedddb-afe1-4913-b011-0ae596652a8f · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Instructblip: Towards general- purpose vision-language models with instruction tuning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 87cfbd50-de8b-49ec-b945-027c72db7606 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e2d588e2-fb9b-4a59-a291-399763e5552d · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 29bc51e1-8673-4c9e-99e7-0ccb6c6ff48d · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Open r1: A fully open reproduction of deepseek-r1
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d65c834a-2372-447d-9e05-dbe0880a3c18 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4986add1-f871-42ba-9788-9a1eab24a6ec · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Lora: Low-rank adaptation of large language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ec121407-24bd-44b4-8d9c-34a06036a044 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 92a507b0-ac44-4c86-98cf-444840ef807a · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model OpenAI o1 System Card
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df3eeafd-ec40-49ea-90a6-77f32fe11478 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Chatrex: Tam- ing multimodal llm for joint perception and understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96620c89-e9fd-4328-9187-83e56a18f2ee · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Grounding language models to images for multimodal in- puts and outputs
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 40f9d804-f7f2-4d01-bb7e-06e0785fcac8 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Buy 4 reinforce samples, get a baseline for free! 2019
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d1c924ca-ac83-4baa-97d1-2ef15c5cbd5f · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Lisa: Reasoning segmentation via large language model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 39ab7beb-e441-48b4-b580-beb318e51e0e · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model LLaVA-OneVision: Easy Visual Task Transfer
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4af7bb55-f1f2-412e-ab54-6c910469b73d · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eb622797-c1b5-4bd4-9032-0072634c0e27 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Microsoft coco: Common objects in context
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3ffb07b7-fc52-4abd-88c1-6c08f724bfe1 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e116865a-fc3d-4f9d-9c6e-00bbf0a1694c · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Improved baselines with visual instruction tuning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 519c1a2b-a43b-473b-96cd-1602b7642589 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Llava-next: Im- proved reasoning, ocr, and world knowledge
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a4cb96c4-7587-4556-8540-4b484a9f306f · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Visual instruction tuning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 714d51d7-5cdc-4a95-99ee-c9a4ee2fc06a · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 32ddb30e-97e4-434d-b92a-a0e9df2442ba · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 719603ed-ff0c-4359-a776-32914a220135 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Understanding R1-Zero-Like Training: A Critical Perspective
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0768d0b9-8533-41be-ad6d-9c493fd87e07 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 312176dd-1196-4fa3-8c1e-fcf728b6ac8c · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 576a9a96-81bf-465d-a663-ecf7c733f8d4 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Generation and comprehension of unambiguous object descriptions
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ff7a526d-feeb-4460-bd3f-958d0d85ffd4 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 92950f75-d039-474b-990a-1d00d26232f1 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Training language models to follow instructions with human feedback
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3e905d17-97ac-4822-9855-bdfcdb3f7df4 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Feedback Loops With Language Models Drive In-Context Reward Hacking
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e015528a-d976-4523-bd6f-cd551e166221 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Spontaneous Reward Hacking in Iterative Self-Refinement
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0a5012b6-ab47-4ea5-ba0e-2123d4e15ffd · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4260cbdf-4111-46a3-924b-0795a670a64f · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Learning transferable visual models from natural language supervi- sion
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6a4c6fa0-345c-4a6c-8450-e3c305b51d1a · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c7dd6100-c3bf-4b53-820a-23ed941ebf2a · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Proximal Policy Optimization Algorithms
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0bcfd00e-02d1-49a0-b713-31c1c48b6d53 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0a0c0629-f13f-49fd-8b7a-bdcf99bef568 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 97464a48-6f5b-4f1d-865f-85cc5f5f6761 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Mea- suring multimodal mathematical reasoning with math-vision dataset
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4dc7ce2f-2b29-4453-a42b-bdc1b9ca76e6 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Large Language Models are not Fair Evaluators
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8852da8b-6d3d-4006-b01d-eab1f18e271e · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ac9d3ad7-dcab-4178-9e0e-76e51601f9aa · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Language Models Learn to Mislead Humans via RLHF
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0bb6fb3e-5c44-4609-9991-6cae86b810ba · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Described object detection: Liberating ob- ject detection with flexible expressions
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 32c79bb6-a543-416a-a4de-de98b4996187 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b0d54564-a0a2-4f3b-a323-1f5002bf9c74 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model How to Evaluate the Generalization of Detection? A Benchmark for Comprehensive Open-Vocabulary Detection
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c12a8a19-e14d-4903-b53d-cd91c3f4f356 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Modeling context in referring expres- sions
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4a697026-1be9-46ed-b0f5-473b729784fc · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 983bf605-c8de-4172-855f-4c222e59ceda · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Sigmoid loss for language image pre-training
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4cf301d3-284d-4ff2-826c-b2041dda12f0 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bd5c7bf4-fd3b-4804-9941-6e554ecaee2b · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Omdet: Language-aware object detection with large-scale vision-language multi-dataset pre-training
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f10f13e6-8ffa-40a6-87cc-4a8590c18a86 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Omdet: Large-scale vision-language multi-dataset pre-training with multimodal detection network
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 75430d7b-4914-4aba-98af-b33c3d837f81 · outbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 50d153ac-52ee-4466-896c-13f87e16a723 · inbound
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a3daa44e-cd7e-4449-8cf8-95b9598be724 · inbound
ToolRL: Reward is All Tool Learning Needs VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2d815fe0-132c-418e-85ce-a5460afe14c3 · inbound
LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3dfb7bd1-d7df-41ee-85bf-a5370b257c62 · inbound
GRIT: Teaching MLLMs to Think with Images VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bb345abe-abe4-468e-a46a-62ccefe84103 · inbound
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d2e7c93f-81ec-4202-b55c-70010487c3c6 · inbound
Grounded Reinforcement Learning for Visual Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 39ad2883-297c-44a7-82b1-c22f2c117995 · inbound
VGR: Visual Grounded Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 120e63a5-c52e-4a26-af2f-5fc7486869f2 · inbound
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b33e7b31-7f1a-48da-a873-cd18218f6cbf · inbound
Perception-Aware Policy Optimization for Multimodal Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e2202ed7-56a4-4e24-b511-601fe3c97091 · inbound
Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c062ded-f0fa-4b2a-9f15-fb05a5f08c6d · inbound
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc0e6dd3-3e5a-460d-9af1-a39a376e33f0 · inbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e14cdb62-4b47-44b7-b911-927b3d75e9d3 · inbound
An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2fde7dc-d960-434c-bf0f-741f24278e8a · inbound
Privacy-Preserving Approximate Nearest Neighbor Search on High-Dimensional Data VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e87cc55c-343f-44f9-afe5-824b64073806 · inbound
UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8413d793-d6fe-4cb9-9d0b-d81a2125a553 · inbound
KnowDR-REC: A Benchmark for Referring Expression Comprehension with Real-World Knowledge VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e87eb0b-6014-49ed-9520-7068b771d9c2 · inbound
BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57ac40bf-5d47-480f-9deb-01ad51d2eccb · inbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64f4bad2-df1b-4476-9751-a4f2260e6bc2 · inbound
Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a608ce7-6121-43d2-9c7f-03ef9b84989d · inbound
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2f5d191-9c2c-4fbd-98a7-87d562b885a1 · inbound
Reinforced Visual Perception with Tools VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f7faa12-999d-4112-81b0-bde24626ac86 · inbound
Omnidirectional Spatial Modeling from Correlated Panoramas VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23566f1d-f878-43b8-a99f-e53dd0a8d5e6 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 226
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 084b67f3-fa5d-4c0f-b1bd-84c94ccd18af · inbound
From Long to Short: LLMs Excel at Trimming Own Reasoning Chains VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8ad9147-4c07-4378-9304-b9ce605681db · inbound
Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3735ede-680f-4e86-ba9c-427534286499 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 148
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e93839-5817-4919-88f0-c24341bfb23e · inbound
Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 40732e49-d2d2-40ce-8a94-1ee1e59b3e1c · inbound
Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7810c525-b166-4904-87dd-21826e7fabb3 · inbound
Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 50481eac-6e75-479f-bb53-dabbba35c617 · inbound
VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b282ff1-f0d4-4fb3-982e-3019f96ae079 · inbound
RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68b632cc-79aa-4d74-86a2-db543070fd90 · inbound
SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5446b5af-5cd3-4782-82a3-40fd3211f52b · inbound
NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5b06dcdd-4fdf-4144-841c-ee1bdde8cf1c · inbound
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 922c2d2e-5f95-4679-b54e-25d430a55103 · inbound
Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2c3109a-1f58-43d4-928f-67fa88d2b826 · inbound
REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fcbcb43b-483a-4ad3-b779-e84dab31f7b5 · inbound
Boosting Reasoning in Large Multimodal Models via Activation Replay VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df2d44da-ba2d-454b-b5ea-d86e7af558d9 · inbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d768473f-179d-44a3-8a5f-c31fc1858d95 · inbound
OneThinker: All-in-one Reasoning Model for Image and Video VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6d828343-faca-4796-998e-a3172586f1a0 · inbound
TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e1296e17-1939-44a6-8446-7d675f11bdad · inbound
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e5eac96-d3a7-4977-b03f-6de7d9092c6a · inbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e602e11-1181-4cc5-b439-85e1db6cec2f · inbound
Grounding Everything in Tokens for Multimodal Large Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a367bd6e-2075-498d-a6fc-45066e86f294 · inbound
Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3e7f2e0-f35c-443c-b62a-4bfc073645f3 · inbound
OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 914eb0af-1b85-4fde-956b-89ab985962d8 · inbound
AdaTooler-V: Adaptive Tool-Use for Images and Videos VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4d20d988-de98-49b9-85ca-d52942405647 · inbound
Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc2e5b03-15dd-4679-80c1-38a641ff3aec · inbound
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification? VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c8f080b0-a003-4722-887e-ae769e568b88 · inbound
Structure Over Scale: Learning Visual Reasoning from Pedagogical Video VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eed7d692-0306-4517-b330-34f47c25dd15 · inbound
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ed327e9d-cab1-40b6-89ba-62a38fb440b9 · inbound
Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9ede4606-d07b-4c39-bc6f-5ba72652507a · inbound
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ba2cd276-f181-4a38-8a66-1b6db929aa31 · inbound
OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c4ac23c-cb4d-45c5-a42a-e5687e6cd6d7 · inbound
RL-RIG: A Generative Spatial Reasoner via Intrinsic Reflection VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bab4e334-fae2-4c1e-9a71-35ae3a90e4cb · inbound
HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d24340bd-d641-4b58-9af2-91abae61c2c1 · inbound
Topo-R1: Detecting Topological Anomalies via Vision-Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d88e741c-b566-4753-a169-5bd6d602c2c9 · inbound
From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87762264-c57f-424c-9ee5-131cc2f8cd53 · inbound
Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2ccc4dd-40fd-4eb5-93e6-34465de90ef2 · inbound
Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d98bb4a2-87a3-49a4-8ddc-3a8e3e69a78d · inbound
Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8e58e6d3-6608-4a82-912e-4955c72ecbe9 · inbound
Discovering Failure Modes in Vision-Language Models using RL VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 45ff02e8-4bc7-44eb-830e-861a206ba9b6 · inbound
CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 644fa6b0-d39a-40b9-8181-83ed4e26ef8d · inbound
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 558253d9-6193-4e14-8cd1-4bde4fbb394c · inbound
Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b54b29fe-c74b-488b-846b-361f34758cb1 · inbound
OOWM: Structuring Embodied Reasoning and Planning via Object-Oriented Programmatic World Modeling VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3b667d81-2581-4c67-9edd-12839033cef5 · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 193
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5a1fb6a-5f47-4201-af1b-736d1bcc01b9 · inbound
Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7da63503-002e-4858-ac03-28b130b5c731 · inbound
One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 147
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8a67ee26-dc25-4400-9898-6c88daf9116e · inbound
Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e0ce69b-2b9b-47b9-9ed4-16a00ac1b5e8 · inbound
CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 556148bc-c795-41eb-b6a4-212a585ea028 · inbound
Can Multimodal Large Language Models Truly Understand Small Objects? VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a7c99fd3-1d62-4345-bd59-091916a71229 · inbound
See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d54c7844-3153-4215-be7a-09283c4ae2d5 · inbound
Improving Vision-language Models with Perception-centric Process Reward Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a443c9d2-ee01-4222-879c-4cf925ed77ce · inbound
Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 460e0320-59c6-49d9-9733-236fa7616836 · inbound
Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a81389c6-a2b6-40a6-b5d9-93a8301ae4f9 · inbound
Perceptual Flow Network for Visually Grounded Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 28c62c2e-3fe9-432a-b5be-a4011a0207d0 · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 658ab417-45af-4dca-a5bb-0218e6afd046 · inbound
GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 67237bf4-4e75-4358-a145-afa3a8758ff2 · inbound
GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ad232e79-e2fc-469c-8fa9-801c3359fa6f · inbound
MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation da8d7f24-f5c8-4597-a479-d30a643b25c2 · inbound
RemoteZero: Geospatial Reasoning with Zero Human Annotations VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 59973c7d-2181-407f-b537-ede3c94155a3 · inbound
Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d0cdbd11-541b-46a3-8084-81a3914bf203 · inbound
Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fcf3f8e9-7323-4a66-b103-bb7f5d2fbf20 · inbound
AffectGPT-RL: Revealing Roles of Reinforcement Learning in Open-Vocabulary Emotion Recognition VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 69960409-cda5-4e13-939c-a65f79365450 · inbound
LiteGUI: Distilling Compact GUI Agents with Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f8acd791-5df8-4dd7-bbd6-c5a2498e6c76 · inbound
SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7c24febd-35b3-4a49-9703-83a62fa63925 · inbound
SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b56d0797-9ee0-43fc-8cae-06e47e8fcb9a · inbound
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cfe4d501-f8f1-44c0-b010-d17b2a2fdd09 · inbound
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cc0f7559-f10b-461a-b8f7-da4e4af1109f · inbound
From Web to Pixels: Bringing Agentic Search into Visual Perception VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8612d913-230e-44a5-8cd4-73577d5900c8 · inbound
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2780cb40-bb28-4db9-a1fa-814241f6ab2b · inbound
Dual-Pathway Circuits of Object Hallucination in Vision-Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aa8cacb8-9afb-4e41-bea9-bf05ffde4c41 · inbound
CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cc8e34e9-c739-44a8-b4e4-9ec6fcefc36f · inbound
CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f748a5f6-7558-4d9c-a256-802773430a75 · inbound
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3e34dc19-2157-4cb6-a67f-45f662c6b6cf · inbound
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0755c88d-71fc-4f4d-8870-a80a1415419a · inbound
From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6d2089e9-74ca-4f76-801f-e0c3b3565e8e · inbound
Segmentation, Detection and Explanation: A Unified Framework for CT Appearance Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 390cf3a5-ba9a-4ac8-bb4e-a3208d65369c · inbound
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e0f8ecf4-68d3-45a7-baa9-88106c6b4874 · inbound
ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned Representation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.