Pith. sign in

Paper Citation Record · LEDGER

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

As of 6 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 100 inbound Pith citation observations for arXiv:2504.07615.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.07615 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T01:13:57.368874Z

measured 161 of 161 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 161 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:03:31.357707Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T03:17:51.904109Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact34
  • verified fuzzy26
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 23851756-2cf5-4645-ac97-4be32c9f3c81 · outbound

This paper cites https://github.com/hiyouga/ EasyR1.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model https://github.com/hiyouga/ EasyR1

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.568179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:89752481db526bf2d12950b8a905485b683d96c75ddf69c163aa26a7249723db

Observation d876a72b-1b40-4db6-8a81-ded882ac0cb4 · outbound

This paper cites GPT-4 Technical Report.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model GPT-4 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.403136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:faa44b61ebfc2bcf98fd07677d356e38f145d0e3cf4b2f0ba9ad6110a3839992

Observation 0d80caaa-4d7e-48e9-9b76-71be48d9def9 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:42:22.769824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:6fce138739555d6138a1dd588e06262125a565ed14a341dcd8517f0b64bd01f0

Observation d722ba59-8902-446f-afaf-86f2f6ff55fe · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.537872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:d63f886b581912404f02c8efbdc807e5cae8e1ab9a1e4395cad8956ebf82b861

Observation 1f561667-281c-4ced-a6f5-e96af9111c57 · outbound

This paper cites Concrete Problems in AI Safety.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Concrete Problems in AI Safety

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.515655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:db3308fe556381054be7b61d5c3adccdf037f7580d72cd842a96b2a78c1c7f34

Observation befc5d16-58e5-4381-baa3-9fde0d0f54ea · outbound

This paper cites Qwen2.5-VL Technical Report.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Qwen2.5-VL Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.518985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:372df9422add82fe4fad9326cd447aa0e4b77fdf33c9f3b2f6c452d13969d834

Observation 54ded20d-c61f-4f0d-afe0-acdd9efd7b58 · outbound

This paper cites Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.522812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:6672c6f420a08cb8693ac0a69891c4d2d67784c08ad0bcb19b20643880164d2b

Observation 5cb3c842-f840-4ff6-a274-05f44a0ad88f · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:13.182567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:4c544773252615a0cd360ccbe476feacf1555dc4ca06f17774474736c238127a

Observation 6ce284a4-bdfa-49c4-abff-8f1609d72718 · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision- language models with less than $3.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model R1-v: Reinforcing super generalization ability in vision- language models with less than $3

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.551540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:9bb5763af663daa9a7b0620b46575177927ce6778dcb3d7fab1a2840bac6d651

Observation f97b9fff-840d-4c74-a7a3-fb38cd2134f6 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.530850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:4e23f4e69de8d52fcf54828c739b427ee41d3cc38ce19bec8ac58e286c047d6d

Observation b6244dcf-3734-4ed9-a83f-90d621649de1 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.534553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:57937b5ba74c2ee8066421a9fdeca0d4e58fa7f62c6494e78c028c52f19ced6e

Observation 59d81237-41e9-41ce-ba6b-7ef537ba1516 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.412568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:2a985abdd3949916f61d542e8b5b5e1e5f0731c45d99def02f99b433d7fad220

Observation 60aedddb-afe1-4913-b011-0ae596652a8f · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Instructblip: Towards general- purpose vision-language models with instruction tuning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.562944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:6a83c4024e934e9091ce9222c3e27bddf03e4e142ea3dc0e98eaa721c5f0d22b

Observation 87cfbd50-de8b-49ec-b945-027c72db7606 · outbound

This paper cites Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:13:57.417184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:8667fe90bc74e109f6a4b1e87a221f82b62191a30abea06c3e12b174c59f882b

Observation e2d588e2-fb9b-4a59-a291-399763e5552d · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:2623c47af8ce6ebd9db4ba89cb25845e3b3aa96ef6cfad0a32a5977a59beb962

Observation 29bc51e1-8673-4c9e-99e7-0ccb6c6ff48d · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Open r1: A fully open reproduction of deepseek-r1

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.570408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:fb5a59c683f2d4c49e57d9dee19461587beef3c7ea530d32bd6104617265de2f

Observation d65c834a-2372-447d-9e05-dbe0880a3c18 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.420952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:0bb27f920447271b6e00541344c23595797c812eb4e53c6c943135972dc64034

Observation 4986add1-f871-42ba-9788-9a1eab24a6ec · outbound

This paper cites Lora: Low-rank adaptation of large language models.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Lora: Low-rank adaptation of large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.575484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:2cc8f2552122e84533bd5bfa082fcbf3ed34e672af7212a235329fdd0afd5277

Observation ec121407-24bd-44b4-8d9c-34a06036a044 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.424812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:61816313a0fa48065a0d2ba2846f10e4a78132a70c6d4d43a132ea3b0485c50f

Observation 92a507b0-ac44-4c86-98cf-444840ef807a · outbound

This paper cites OpenAI o1 System Card.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model OpenAI o1 System Card

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.428243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:a9eea725f8deddfba1333cfc45b88582d5c6febdbd3336ce8301de310a559e80

Observation df3eeafd-ec40-49ea-90a6-77f32fe11478 · outbound

This paper cites Chatrex: Tam- ing multimodal llm for joint perception and understanding.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Chatrex: Tam- ing multimodal llm for joint perception and understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.581888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:b2b8df63374575500df445755d8c14fe7bcdc801598701a4aa695270281a2f69

Observation 96620c89-e9fd-4328-9187-83e56a18f2ee · outbound

This paper cites Grounding language models to images for multimodal in- puts and outputs.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Grounding language models to images for multimodal in- puts and outputs

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.584232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:4b77472c0b3b45f776c5ef952499589018a1897da67c9fc1387696a69cc225bb

Observation 40f9d804-f7f2-4d01-bb7e-06e0785fcac8 · outbound

This paper cites Buy 4 reinforce samples, get a baseline for free! 2019.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Buy 4 reinforce samples, get a baseline for free! 2019

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.586321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:3091a6bcb86f776d15064c16332c8cac5cda99093454e1c587a4f246dc5687c2

Observation d1c924ca-ac83-4baa-97d1-2ef15c5cbd5f · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Lisa: Reasoning segmentation via large language model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.588915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:84ebffcb4983e3f2e80b5c3c890ae1f9a3c1c369ea0ce74a06b08273ac0f064f

Observation 39ab7beb-e441-48b4-b580-beb318e51e0e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.431959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:036f51cca361c82b1176a6b5e8694e4383f3f87daeeb27bc1f55a65af102fd7e

Observation 4af7bb55-f1f2-412e-ab54-6c910469b73d · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.593975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:a253785356b5831eb29ea7d078a11f5b22fc2977d993bcdd1405ef85f188a9a6

Observation eb622797-c1b5-4bd4-9032-0072634c0e27 · outbound

This paper cites Microsoft coco: Common objects in context.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Microsoft coco: Common objects in context

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.596465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:92395df2dc04c282b8c6e22d215a6c46ecc8d61336158b1ae0f25f4722c7b25c

Observation 3ffb07b7-fc52-4abd-88c1-6c08f724bfe1 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:18:01.701777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:7afac11393338557a639aa7fa6fe143c918de5ba2d9b69934c6f606b52e1baaa

Observation e116865a-fc3d-4f9d-9c6e-00bbf0a1694c · outbound

This paper cites Improved baselines with visual instruction tuning.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Improved baselines with visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.601079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:d2f07c96ea40932542ef44f88d4839ffba4a9f0d7f040e0ba0e69e6efc18dba1

Observation 519c1a2b-a43b-473b-96cd-1602b7642589 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Llava-next: Im- proved reasoning, ocr, and world knowledge

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.540727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:f113bf24e980647a60867f379e63df8bd29ad692373eac3daa63d2b516e67b3e

Observation a4cb96c4-7587-4556-8540-4b484a9f306f · outbound

This paper cites Visual instruction tuning.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Visual instruction tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.543304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:ab82653bd708e3bfb401422a8d8cc64e1c9e0d87e864a5b90b8ab6c2a07454e1

Observation 714d51d7-5cdc-4a95-99ee-c9a4ee2fc06a · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.546459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:ff4917e5172b16a3c15972526e6b2103428ff1fdd39fa19b7f4a9bd9967c6046

Observation 32ddb30e-97e4-434d-b92a-a0e9df2442ba · outbound

This paper cites LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.439123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:cfadcebd64f7b1e2db5db2eeda5349043178ed706d256ac3fb91fd523f3fc15a

Observation 719603ed-ff0c-4359-a776-32914a220135 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Understanding R1-Zero-Like Training: A Critical Perspective

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.442465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:83620e7bae42a155435a9ccdce4178a60e06eb8e9bbf405bf9de75622ccde1df

Observation 0768d0b9-8533-41be-ad6d-9c493fd87e07 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:16:16.677111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:acebffd1f322cace30eb90d4c808f0b7c92798e4004b9d74cf5d1d7926c63891

Observation 312176dd-1196-4fa3-8c1e-fcf728b6ac8c · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.451871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:f90bc7c29dac0c1cc19aab8d92434b4a46f3552d4782606705e320c73d2e34df

Observation 576a9a96-81bf-465d-a663-ecf7c733f8d4 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Generation and comprehension of unambiguous object descriptions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.565635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:191bff1732598c3e5be2b03599b58e93852897588eba81e43209cffc87a9761b

Observation ff7a526d-feeb-4460-bd3f-958d0d85ffd4 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.456642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:f55f765fa77500a72b62c1d3c85473f9d0a929149023f74a8fde0865a8dbf5e3

Observation 92950f75-d039-474b-990a-1d00d26232f1 · outbound

This paper cites Training language models to follow instructions with human feedback.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Training language models to follow instructions with human feedback

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.577736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:b0fecf64378ca62e657bec26528c8e63d81f4833f5ebebad1a03e41f206442a0

Observation 3e905d17-97ac-4822-9855-bdfcdb3f7df4 · outbound

This paper cites Feedback Loops With Language Models Drive In-Context Reward Hacking.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.461019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:5c37c632992cc61d7116cce8a4ded7790f18ccceb0a6ee23c556802dd124598d

Observation e015528a-d976-4523-bd6f-cd551e166221 · outbound

This paper cites Spontaneous Reward Hacking in Iterative Self-Refinement.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Spontaneous Reward Hacking in Iterative Self-Refinement

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.464983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:fe3bbfb90a6729d21c3ed96e263aa312f9ebe95f77b75b7e158a4459fa2fba14

Observation 0a5012b6-ab47-4ea5-ba0e-2123d4e15ffd · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:15:46.589538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:56b7e2bc78cf762efb60140bcb98226552d5122e9c3f0b8b2ac32fcbba6a215e

Observation 4260cbdf-4111-46a3-924b-0795a670a64f · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Learning transferable visual models from natural language supervi- sion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.549047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:0d91427d3e35acfa3d8dc888cd5cd3488772876d1d98cf339a8b1a13c3a0076d

Observation 6a4c6fa0-345c-4a6c-8450-e3c305b51d1a · outbound

This paper cites Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.473789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:9c0e9de4baab4b2458619ab2b86d992c1e70ad625babc59ac8fbccdf45c11d31

Observation c7dd6100-c3bf-4b53-820a-23ed941ebf2a · outbound

This paper cites Proximal Policy Optimization Algorithms.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Proximal Policy Optimization Algorithms

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.477597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:15c6c16a803f865c79c469c1e6305287273477418a209ce52125b588c68a2fd6

Observation 0bcfd00e-02d1-49a0-b713-31c1c48b6d53 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.481688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:1c0cb575d27d344bc4bd8bc5a9db4bb8527dea117e485119b5cbd2309143f916

Observation 0a0c0629-f13f-49fd-8b7a-bdcf99bef568 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:6593d412dbe0baaeaff1f77552517a4f1aa60440766c9272ec1343507e3eb708

Observation 97464a48-6f5b-4f1d-865f-85cc5f5f6761 · outbound

This paper cites Mea- suring multimodal mathematical reasoning with math-vision dataset.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Mea- suring multimodal mathematical reasoning with math-vision dataset

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.579846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:82daf97ffd2980c420f294a5c8838829b87926225cfa97052dd2884bd80a3046

Observation 4dc7ce2f-2b29-4453-a42b-bdc1b9ca76e6 · outbound

This paper cites Large Language Models are not Fair Evaluators.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Large Language Models are not Fair Evaluators

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:10:42.624770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:44c260403e79abf8b0d824e40589191723f08b191153abc9546055e49460d432

Observation 8852da8b-6d3d-4006-b01d-eab1f18e271e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.493907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:f9c4aa706c8b7163925a01bc4fc486a1ef6f151638e6e969b18ea2c5beb12d9a

Observation ac9d3ad7-dcab-4178-9e0e-76e51601f9aa · outbound

This paper cites Language Models Learn to Mislead Humans via RLHF.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Language Models Learn to Mislead Humans via RLHF

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.497300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:f08992350054de32bc27f17ac49e8de9c242c28ca05208f983369667436e0ff4

Observation 0bb6fb3e-5c44-4609-9991-6cae86b810ba · outbound

This paper cites Described object detection: Liberating ob- ject detection with flexible expressions.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Described object detection: Liberating ob- ject detection with flexible expressions

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.556933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:aef6ef1cf8fbb70b4de3f0f0b9090e797ccfd78907bf93160a3479a9f5c7d16a

Observation 32c79bb6-a543-416a-a4de-de98b4996187 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:19:20.780377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:cd1372eaaeb5bbd2f70bf0f175b8d06afec3c4d628935a2477d8f1e39231b178

Observation b0d54564-a0a2-4f3b-a323-1f5002bf9c74 · outbound

This paper cites How to Evaluate the Generalization of Detection? A Benchmark for Comprehensive Open-Vocabulary Detection.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model How to Evaluate the Generalization of Detection? A Benchmark for Comprehensive Open-Vocabulary Detection

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.504666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:3cc8b4ebfba1352218cd520e8759e075e06f8554a3c6f70c4be1555ff83399ba

Observation c12a8a19-e14d-4903-b53d-cd91c3f4f356 · outbound

This paper cites Modeling context in referring expres- sions.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Modeling context in referring expres- sions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.591143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:8e41b6cac27ad4927afef49c324c4f743b9a8faf120ab22b1288e9b7346c4b64

Observation 4a697026-1be9-46ed-b0f5-473b729784fc · outbound

This paper cites InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.508448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:dcc4f037530ccd981fa78c18f7e1a2464ed7601b7070fd84dd5dcea23c7645ff

Observation 983bf605-c8de-4172-855f-4c222e59ceda · outbound

This paper cites Sigmoid loss for language image pre-training.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Sigmoid loss for language image pre-training

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.554184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:e524455e5e56c7341f28735de4043604f960fdef192b14e2d367ebc0b5498266

Observation 4cf301d3-284d-4ff2-826c-b2041dda12f0 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.559930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:0bdf1479fe6cfaa47f67bb87d71a20ea8c7544d4b8c1124488c1b556e0f86cfd

Observation bd5c7bf4-fd3b-4804-9941-6e554ecaee2b · outbound

This paper cites Omdet: Language-aware object detection with large-scale vision-language multi-dataset pre-training.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Omdet: Language-aware object detection with large-scale vision-language multi-dataset pre-training

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.572920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:0e6758fa38cc40291f1c154af364e66b8b756b62c80249b3c11b290dc386e89b

Observation f10f13e6-8ffa-40a6-87cc-4a8590c18a86 · outbound

This paper cites Omdet: Large-scale vision-language multi-dataset pre-training with multimodal detection network.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Omdet: Large-scale vision-language multi-dataset pre-training with multimodal detection network

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:13:57.598711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:584e070d53314d7939d54111f44e7ee65d01f9d1c7da33caff19ff46c6925951

Observation 75430d7b-4914-4aba-98af-b33c3d837f81 · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:13:47.672124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:6d39c4ebc682c021f70c39186a94b1c07ccfdc57e8afea57a541f483346818b2

Pith citing papers

Observation 50d153ac-52ee-4466-896c-13f87e16a723 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:59:03.247850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:d6704c042be7953cde44bde7cbef568b38d05e76718acc7999a827a9055b4731

Observation a3daa44e-cd7e-4449-8cf8-95b9598be724 · inbound

ToolRL: Reward is All Tool Learning Needs cites this paper.

ToolRL: Reward is All Tool Learning Needs VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T00:26:48.524669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T00:26:48.291431Z digest=sha256:807c87929a31fd21348200deaa0e2f8a3a45f36baba29c1c49cd65031a82faed

Observation 2d815fe0-132c-418e-85ce-a5460afe14c3 · inbound

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models cites this paper.

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:51:37.657422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T13:47:51.436258Z digest=sha256:984e6c7c06ed519abbeb39d7d857772f41eec15c070a526450c25ca37367e03d

Observation 3dfb7bd1-d7df-41ee-85bf-a5370b257c62 · inbound

GRIT: Teaching MLLMs to Think with Images cites this paper.

GRIT: Teaching MLLMs to Think with Images VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:31:36.050924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T13:29:47.529564Z digest=sha256:8460141db48cdcddb3ae7bd5e1346b3d1ffb0794d9ed914b2618af1ffdde78e7

Observation bb345abe-abe4-468e-a46a-62ccefe84103 · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:40:56.038031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:d545812c32eb6b648d20d8d84cbe2c2cef9c4fa74f7b74c59b7fae184f34e3a0

Observation d2e7c93f-81ec-4202-b55c-70010487c3c6 · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:05:52.018468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:2206f09ee1502d0a2a5b95b5fe41e2b00f7b4014f3b45d422ecf9a2509f9cafc

Observation 39ad2883-297c-44a7-82b1-c22f2c117995 · inbound

VGR: Visual Grounded Reasoning cites this paper.

VGR: Visual Grounded Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:12:14.430161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:11:00.295700Z digest=sha256:a00d393ff5df35c30cbc35edd0a14e5ce746e45d4500c82782e4c193d73d8f11

Observation 120e63a5-c52e-4a26-af2f-5fc7486869f2 · inbound

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning cites this paper.

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:52:08.138562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T06:50:02.607136Z digest=sha256:2e760d5e8b86c228b08424b214565cf7f807c73c3e646dc8cb0182b3fc5dd972

Observation b33e7b31-7f1a-48da-a873-cd18218f6cbf · inbound

Perception-Aware Policy Optimization for Multimodal Reasoning cites this paper.

Perception-Aware Policy Optimization for Multimodal Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.869640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:5769f16e0b432eaf75967e7a52b9d6f5d71be676ec8f67418304623b92f9aa2c

Observation e2202ed7-56a4-4e24-b511-601fe3c97091 · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:31.357707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:31.357707Z digest=sha256:d2c1b4772e1d331d9e2d4f797533d023f27a30dd529229ab6f1977db269aff73

Observation 1c062ded-f0fa-4b2a-9f15-fb05a5f08c6d · inbound

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience cites this paper.

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T23:55:50.919858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:55:50.919858Z digest=sha256:878d7f93f0b366693210d8b635d945b72606b2ecef1c1d3745db3b0f92c76080

Observation dc0e6dd3-3e5a-460d-9af1-a39a376e33f0 · inbound

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models cites this paper.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:16.170654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:16.170654Z digest=sha256:d04ee34c9bb98e16ca06541949a905bdb404b377d950690b4db4fd0b03af0347

Observation e14cdb62-4b47-44b7-b911-927b3d75e9d3 · inbound

An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal cites this paper.

An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T23:25:49.136154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:25:49.136154Z digest=sha256:fa88dea9d3fb522fa0dfeb481299a1eb3f70f2940a36b05071bf96bfeeda2eb3

Observation d2fde7dc-d960-434c-bf0f-741f24278e8a · inbound

Privacy-Preserving Approximate Nearest Neighbor Search on High-Dimensional Data cites this paper.

Privacy-Preserving Approximate Nearest Neighbor Search on High-Dimensional Data VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:47.487249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:30:47.487249Z digest=sha256:529985b23e9022c372e853950c1037ecae303f5eec729c3c68e520ef735a008c

Observation e87cc55c-343f-44f9-afe5-824b64073806 · inbound

UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning cites this paper.

UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:21:53.326591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T22:17:41.758059Z digest=sha256:3781fadf752f2ee6f5458d4522c9430004ef0723514171f5eec97fdbff11d5c1

Observation 8413d793-d6fe-4cb9-9d0b-d81a2125a553 · inbound

KnowDR-REC: A Benchmark for Referring Expression Comprehension with Real-World Knowledge cites this paper.

KnowDR-REC: A Benchmark for Referring Expression Comprehension with Real-World Knowledge VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T21:11:52.381315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:11:52.381315Z digest=sha256:60a1d26550449de33ec4fa605c717c50a1d99af416e67e1ccfd88fa8a74d2a0d

Observation 0e87eb0b-6014-49ed-9520-7068b771d9c2 · inbound

BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens cites this paper.

BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T17:04:38.739210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:04:38.739210Z digest=sha256:32029816c6a4128037f50dca49b7274cfc989a71acf023843a2299f96f701b2b

Observation 57ac40bf-5d47-480f-9deb-01ad51d2eccb · inbound

MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use cites this paper.

MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T16:22:51.550260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:22:51.550260Z digest=sha256:62e6028e70faf7fde92acea172d94d7406b0ca2ee792e54df496b7bc73a67ea6

Observation 64f4bad2-df1b-4476-9751-a4f2260e6bc2 · inbound

Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation cites this paper.

Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:34.579065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:34.579065Z digest=sha256:d0a46d295b376b23c4ccaacd8208fe742f335f8395bd42d869ec3b827fc0d2bc

Observation 5a608ce7-6121-43d2-9c7f-03ef9b84989d · inbound

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding cites this paper.

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:16.788157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:10:16.788157Z digest=sha256:232b18c880467d08f7898728d32aaab15561547fab66bac6aa7384f20eb79a55

Observation b2f5d191-9c2c-4fbd-98a7-87d562b885a1 · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:05.031845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:05.031845Z digest=sha256:1bc4da00a9c18b2cf23ce9c6cb4348a844527088ca8d89eff80ce3d25d3102d0

Observation 6f7faa12-999d-4112-81b0-bde24626ac86 · inbound

Omnidirectional Spatial Modeling from Correlated Panoramas cites this paper.

Omnidirectional Spatial Modeling from Correlated Panoramas VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:37.761608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:37.761608Z digest=sha256:a41593dcda796645559b69221044a3097e7c0ef99f2525c6946bab0940b0edac

Observation 23566f1d-f878-43b8-a99f-e53dd0a8d5e6 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 226

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.450263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:0a6b10b00c4ee0622986ee62efa4ff75621dfd4d469c899eeb77b5c442abb39f

Observation 084b67f3-fa5d-4c0f-b1bd-84c94ccd18af · inbound

From Long to Short: LLMs Excel at Trimming Own Reasoning Chains cites this paper.

From Long to Short: LLMs Excel at Trimming Own Reasoning Chains VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T00:04:01.063328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:04:01.063328Z digest=sha256:2683faa852fd35b9193294994ab9932d2c6ef985f58ed54e3a953b424ddce8d4

Observation b8ad9147-4c07-4378-9304-b9ce605681db · inbound

Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration cites this paper.

Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T18:17:17.467108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:17:17.467108Z digest=sha256:449fc3d4c3026aa91074eccd0eded3847d827b534ef8e1301d35f17ff1d43413

Observation e3735ede-680f-4e86-ba9c-427534286499 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 148

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:39.282457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:39.282457Z digest=sha256:c8f39bd1d5f6f162acc4a0282cf18bc39048c4e3e2becdbcd4a0173fa276b17f

Observation c3e93839-5817-4919-88f0-c24341bfb23e · inbound

Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning cites this paper.

Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T14:11:27.514218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T14:09:13.310620Z digest=sha256:a814afdd3645f0ddee768763d0272388fa4027f86091779191c1e387e9d39025

Observation 40732e49-d2d2-40ce-8a94-1ee1e59b3e1c · inbound

Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework cites this paper.

Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-18T12:32:36.398678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T12:31:25.257879Z digest=sha256:fe54c1dbc3cc7935880e612f66dd17b06633f2dd00368c55db6a65389609ae4e

Observation 7810c525-b166-4904-87dd-21826e7fabb3 · inbound

Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards cites this paper.

Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:46:19.520512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:43:31.679386Z digest=sha256:b643556e4d11ee510bb617bf56a14c258c5bec574c0d82408b96407340052d24

Observation 50481eac-6e75-479f-bb53-dabbba35c617 · inbound

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation cites this paper.

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:41.865785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:40:41.865785Z digest=sha256:6061e36ba2315cc9bac24425f2c4993d43ec021dd74e47d0d419a7cd4dc72369

Observation 4b282ff1-f0d4-4fb3-982e-3019f96ae079 · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:40.274255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:40.274255Z digest=sha256:4b9e0705522cc44dd7ccff58f2dfe5efbd53ab7346588ead1de9067a35874b5c

Observation 68b632cc-79aa-4d74-86a2-db543070fd90 · inbound

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning cites this paper.

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-21T20:24:21.295605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T20:24:02.748854Z digest=sha256:64be00796f0df8fc9463bfdd245556f04a47031962486fb251128b6dcef06967

Observation 5446b5af-5cd3-4782-82a3-40fd3211f52b · inbound

NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation cites this paper.

NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:40:53.027955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T04:39:58.296388Z digest=sha256:bfd77d4bcfe3067f05ead0b4c599572fbf70049682b6195e37a02722e177fafa

Observation 5b06dcdd-4fdf-4144-841c-ee1bdde8cf1c · inbound

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation cites this paper.

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T07:57:25.675221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:57:25.675221Z digest=sha256:707d9865d4b688a4a461bbfbeacbc9c974ba18f4a12d01e737d03990cb80fcd2

Observation 922c2d2e-5f95-4679-b54e-25d430a55103 · inbound

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning cites this paper.

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T07:02:43.058582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:02:43.058582Z digest=sha256:806ec6e9d7a27017b1825647fed136b24b734354c5b7bb81a7dc87811b1b6802

Observation a2c3109a-1f58-43d4-928f-67fa88d2b826 · inbound

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding cites this paper.

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:20:22.826442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T22:19:36.366837Z digest=sha256:8f87a91125a183a72f76facdd11aa40e2f13e54d73cf681734c64174d81827d2

Observation fcbcb43b-483a-4ad3-b779-e84dab31f7b5 · inbound

Boosting Reasoning in Large Multimodal Models via Activation Replay cites this paper.

Boosting Reasoning in Large Multimodal Models via Activation Replay VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:09:04.004788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T05:05:48.682057Z digest=sha256:f1978247ac3bc8e2ca16e021d0b991bd1a3c573c0944021b08d1ef0f30b96332

Observation df2d44da-ba2d-454b-b5ea-d86e7af558d9 · inbound

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling cites this paper.

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-22T12:31:32.175943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T12:26:35.347190Z digest=sha256:d4e70a669b9bb10f174968954806f591d635affd7d7470faf7c0d0768e9e88ab

Observation d768473f-179d-44a3-8a5f-c31fc1858d95 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:11:26.532514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:9d6db73892b4f0bf5cc7bfdf1efb24ed24c581cbf19627f419c8a1343f5c4e87

Observation 6d828343-faca-4796-998e-a3172586f1a0 · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.253010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:259b31a9f3773701b97192822fd28f96cd861d1e969a89e780ac4cabc36922a4

Observation e1296e17-1939-44a6-8446-7d675f11bdad · inbound

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension cites this paper.

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T18:15:06.983726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:15:06.983726Z digest=sha256:557eae6cd862a97ce97c0f971f28ffcfd04ad7e6cde33c2b6b00755f2acb890e

Observation 9e5eac96-d3a7-4977-b03f-6de7d9092c6a · inbound

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding cites this paper.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.523998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:21dab0e0c3c90d697fb458e2dc5b2a47137cb7ca947f08e8b2655369c55254f7

Observation 9e602e11-1181-4cc5-b439-85e1db6cec2f · inbound

Grounding Everything in Tokens for Multimodal Large Language Models cites this paper.

Grounding Everything in Tokens for Multimodal Large Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:31:21.862892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T23:31:05.422935Z digest=sha256:5ecc47fa62828bab1777d9fb392a0204eb14908d872bf405467f125148be325b

Observation a367bd6e-2075-498d-a6fc-45066e86f294 · inbound

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics cites this paper.

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T16:27:34.117502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:27:34.117502Z digest=sha256:22ad48f736cae0bfbb2314f266620c03e3899840b6bf7dbb2f294408fa9a70e6

Observation f3e7f2e0-f35c-443c-b62a-4bfc073645f3 · inbound

OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving cites this paper.

OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:38:37.859935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T22:34:00.895252Z digest=sha256:ae5fec7c683c4139afd759970b81574c656400cc64c363c89206eeda679fe3fd

Observation 914eb0af-1b85-4fde-956b-89ab985962d8 · inbound

AdaTooler-V: Adaptive Tool-Use for Images and Videos cites this paper.

AdaTooler-V: Adaptive Tool-Use for Images and Videos VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:28:34.318011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T21:23:33.598026Z digest=sha256:a670a8dda2c74ff3d816884c7b0f8ec706e453e9cc99a69f55b267af6ccf381c

Observation 4d20d988-de98-49b9-85ca-d52942405647 · inbound

Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning cites this paper.

Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T12:31:16.953290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:31:16.953290Z digest=sha256:651c56abea2f68a76337b65529879de20122a668d3d59b3f713411279a444ab9

Observation dc2e5b03-15dd-4679-80c1-38a641ff3aec · inbound

Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification? cites this paper.

Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification? VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:53:00.547163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T14:51:26.368439Z digest=sha256:80a2f32108bb679541859452969882e25296d5b2aa303783339e5314c51752f7

Observation c8f080b0-a003-4722-887e-ae769e568b88 · inbound

Structure Over Scale: Learning Visual Reasoning from Pedagogical Video cites this paper.

Structure Over Scale: Learning Visual Reasoning from Pedagogical Video VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:22:40.191829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T09:22:24.609440Z digest=sha256:f1df71c9d5fd7f30548da2299e514c68e17afe865739a3594995b20cc46ff3d3

Observation eed7d692-0306-4517-b330-34f47c25dd15 · inbound

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning cites this paper.

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:02:42.482049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T10:02:20.477517Z digest=sha256:bab38936cb73ba5da0893fdb8037005584c91aac906e4c5188648a064d8c6d31

Observation ed327e9d-cab1-40b6-89ba-62a38fb440b9 · inbound

Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension cites this paper.

Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T03:30:33.075965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T03:27:53.694506Z digest=sha256:ec9563e90cdbfb41ff904ee168223d5705940f2bbdf836d0d17e9d4bda1380d8

Observation 9ede4606-d07b-4c39-bc6f-5ba72652507a · inbound

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking cites this paper.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:50:17.290213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:65b66daae33d8b67b9107d20c8999565423a300474a512d817dd4b2cc9a55651

Observation ba2cd276-f181-4a38-8a66-1b6db929aa31 · inbound

OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents cites this paper.

OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T22:10:52.040247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:10:52.040247Z digest=sha256:353cd87b0003928e2dc945ddc5f45595c07e607564292133e6954ab28084633b

Observation 0c4ac23c-cb4d-45c5-a42a-e5687e6cd6d7 · inbound

RL-RIG: A Generative Spatial Reasoner via Intrinsic Reflection cites this paper.

RL-RIG: A Generative Spatial Reasoner via Intrinsic Reflection VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:36:35.289766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:33:09.627731Z digest=sha256:568471eb6515f1bdaf53acee4f2e5bf60c014c1cded365e32ad0660077d1d85e

Observation bab4e334-fae2-4c1e-9a71-35ae3a90e4cb · inbound

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework cites this paper.

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T20:19:15.914946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:19:15.914946Z digest=sha256:a9fb8d2cc80b6bffbaa065b3668a9a323ebf57eca6293eaf95cfe7d298f8ea81

Observation d24340bd-d641-4b58-9af2-91abae61c2c1 · inbound

Topo-R1: Detecting Topological Anomalies via Vision-Language Models cites this paper.

Topo-R1: Detecting Topological Anomalies via Vision-Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T11:45:32.863537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T11:41:27.021776Z digest=sha256:0359d2d7997acdd59a933f50d023bd7fd735f21ace2ff403718d35c505b0fc14

Observation d88e741c-b566-4753-a169-5bd6d602c2c9 · inbound

From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation cites this paper.

From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T00:15:23.806078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T00:15:23.806078Z digest=sha256:90ef45774d4d721d5445435ef4403bbeecea0210d323c3eb56807824031f0469

Observation 87762264-c57f-424c-9ee5-131cc2f8cd53 · inbound

Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking cites this paper.

Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T16:55:20.099628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:55:20.099628Z digest=sha256:8ae4e82499a60c69e612e97346f573dc196d608b9334eaffa9a701975fafb93a

Observation c2ccc4dd-40fd-4eb5-93e6-34465de90ef2 · inbound

Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs cites this paper.

Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:38:00.974574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:35:12.859669Z digest=sha256:1beef1217a56632ff287da4e83684b6d3e1e417744f7bfffb00df5011cc189d6

Observation d98bb4a2-87a3-49a4-8ddc-3a8e3e69a78d · inbound

Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models cites this paper.

Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:53:16.242557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T20:48:52.130130Z digest=sha256:4a7fb7ca4a93bb08804074a5acb9a27b828725ced1d21998b9670ee4057b6cc4

Observation 8e58e6d3-6608-4a82-912e-4955c72ecbe9 · inbound

Discovering Failure Modes in Vision-Language Models using RL cites this paper.

Discovering Failure Modes in Vision-Language Models using RL VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:21:25.761342Z digest=sha256:947376d1fd4aecfb53f26643b6b102343e8dacc1bdd146eee16bf4590da524ce

Observation 45ff02e8-4bc7-44eb-830e-861a206ba9b6 · inbound

CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models cites this paper.

CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:52:32.315964Z digest=sha256:f0b9f958d508b3162d3b09f6dad78f24793e24e63ef5a37209869ba883fff842

Observation 644fa6b0-d39a-40b9-8181-83ed4e26ef8d · inbound

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks cites this paper.

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:09:33.968186Z digest=sha256:e23da8d775c6d362fc4879b462714428743b570da72384df0a98851f695bda5c

Observation 558253d9-6193-4e14-8cd1-4bde4fbb394c · inbound

Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models cites this paper.

Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:35:21.514502Z digest=sha256:862d6a1c9007a3fc7ef94c4087fcd71dc8ceac2cae3a9ff8ca75c92cde32e0fa

Observation b54b29fe-c74b-488b-846b-361f34758cb1 · inbound

OOWM: Structuring Embodied Reasoning and Planning via Object-Oriented Programmatic World Modeling cites this paper.

OOWM: Structuring Embodied Reasoning and Planning via Object-Oriented Programmatic World Modeling VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:30:16.474205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:29:55.075825Z digest=sha256:063efe994dbbadf0f51101389cc90209429a28d471576d75092e94dca46fa630

Observation 3b667d81-2581-4c67-9edd-12839033cef5 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 193

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:0d641026eb7af9a47fa07ffb6c15c89a11d4413965ef040b4529274778e01d36

Observation c5a1fb6a-5f47-4201-af1b-736d1bcc01b9 · inbound

Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning cites this paper.

Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T07:27:15.270996Z digest=sha256:6d048da36802c7dfa86bd4e072fb5cfb45402d6035480c5c49272641ace0938c

Observation 7da63503-002e-4858-ac03-28b130b5c731 · inbound

One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models cites this paper.

One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 147

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T04:56:35.796962Z digest=sha256:6d72acdaed2cc2eeb05a972b1e94a3bc0620697b9f70526da4f19b4db2518acb

Observation 8a67ee26-dc25-4400-9898-6c88daf9116e · inbound

Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems cites this paper.

Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T23:27:54.704794Z digest=sha256:19b77c37f6efcd134c4610becd629474aebe17a3dc0019a29007fa40bdc18005

Observation 2e0ce69b-2b9b-47b9-9ed4-16a00ac1b5e8 · inbound

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding cites this paper.

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T12:26:01.568507Z digest=sha256:b11e1ec894a7f5408f0803131c1348f20b19169e7d5a3f195ab478e3b9cc6a87

Observation 556148bc-c795-41eb-b6a4-212a585ea028 · inbound

Can Multimodal Large Language Models Truly Understand Small Objects? cites this paper.

Can Multimodal Large Language Models Truly Understand Small Objects? VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T12:49:53.645987Z digest=sha256:58899a40eae95c8d4ef4b096d46c278714a6c4eac238f09aab210e23a3ba7271

Observation a7c99fd3-1d62-4345-bd59-091916a71229 · inbound

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection cites this paper.

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T04:46:16.497585Z digest=sha256:db73ebb2c00c9729376f6ef59c8d60efcfc989f2af9dcb3088a1da0da032f3d7

Observation d54c7844-3153-4215-be7a-09283c4ae2d5 · inbound

Improving Vision-language Models with Perception-centric Process Reward Models cites this paper.

Improving Vision-language Models with Perception-centric Process Reward Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T04:33:36.634359Z digest=sha256:70a985a08c97251d2ef8a767e6a43bf69ea1c7d0ca031ecbb0d09114cfef6332

Observation a443c9d2-ee01-4222-879c-4cf925ed77ce · inbound

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression cites this paper.

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 98

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T14:36:29.666730Z digest=sha256:c7b3c50af7454967336bc6f362e835a43997dd92254b4e9b1e2f47b42a0f382a

Observation 460e0320-59c6-49d9-9733-236fa7616836 · inbound

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression cites this paper.

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 98

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T04:52:09.685243Z digest=sha256:ec57a166d018314626784942f12ee4eedf578baa64b489d93dce4200ff933ba3

Observation a81389c6-a2b6-40a6-b5d9-93a8301ae4f9 · inbound

Perceptual Flow Network for Visually Grounded Reasoning cites this paper.

Perceptual Flow Network for Visually Grounded Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:9ae0d6df642f08c9aef6043ffb78450a129d0d242ac794bd0adba4248cba075d

Observation 28c62c2e-3fe9-432a-b5be-a4011a0207d0 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:d0d562b1069a4b769e30d3f311c2f909f77e9f9495aa33922be5f59cc9d4832b

Observation 658ab417-45af-4dca-a5bb-0218e6afd046 · inbound

GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning cites this paper.

GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T17:59:30.937349Z digest=sha256:249ed0fbf79eff334ca401c6e714dea83f200a5b2587b719210c34859715a3ba

Observation 67237bf4-4e75-4358-a145-afa3a8758ff2 · inbound

GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning cites this paper.

GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:45:12.113174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T00:37:16.179596Z digest=sha256:e2a8fec504002ba4d66477dc742c0fb9af8708dcbfc1b4db4774c4f3c51ed8f2

Observation ad232e79-e2fc-469c-8fa9-801c3359fa6f · inbound

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models cites this paper.

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T01:20:54.441367Z digest=sha256:5dd9814411551ab46d41f7e05abcf5fd8fd8461948d93df7e599175a1b4a333c

Observation da8d7f24-f5c8-4597-a479-d30a643b25c2 · inbound

RemoteZero: Geospatial Reasoning with Zero Human Annotations cites this paper.

RemoteZero: Geospatial Reasoning with Zero Human Annotations VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T18:24:46.030608Z digest=sha256:b8512e3e80098810e17eb2e772cfb7a7110b2e8e5bd1f93c7d9182b1b0f0aa3a

Observation 59973c7d-2181-407f-b537-ede3c94155a3 · inbound

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling cites this paper.

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:42:30.856741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T07:37:52.346280Z digest=sha256:eca58dd8a1fbc91840162e1c2f870d93f85fa354d83de97d80e8f11de5ac9330

Observation d0cdbd11-541b-46a3-8084-81a3914bf203 · inbound

Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning cites this paper.

Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:02:29.480442Z digest=sha256:d438a47ffc48e32452d0d3eaf67d08ec402920a9823e5a91c6b8f0dbf93e636a

Observation fcf3f8e9-7323-4a66-b103-bb7f5d2fbf20 · inbound

AffectGPT-RL: Revealing Roles of Reinforcement Learning in Open-Vocabulary Emotion Recognition cites this paper.

AffectGPT-RL: Revealing Roles of Reinforcement Learning in Open-Vocabulary Emotion Recognition VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T07:25:42.870223Z digest=sha256:6446ce2800e3774d62b049b5f9e29ce48b8dcea147a0be97edddb71f55486328

Observation 69960409-cda5-4e13-939c-a65f79365450 · inbound

LiteGUI: Distilling Compact GUI Agents with Reinforcement Learning cites this paper.

LiteGUI: Distilling Compact GUI Agents with Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T02:18:57.917353Z digest=sha256:b62ad43e8460d54c8ac7bb68d3b5e2084c10414855d48fd68a1e09a00df479dc

Observation f8acd791-5df8-4dd7-bbd6-c5a2498e6c76 · inbound

SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation cites this paper.

SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T03:53:46.809307Z digest=sha256:a782dda9875a6e71a759b3f9cb34445b94a484452a0a0fb84f7f35f1bb27791f

Observation 7c24febd-35b3-4a49-9703-83a62fa63925 · inbound

SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation cites this paper.

SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:53:02.231111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T21:52:08.637283Z digest=sha256:fad27a238b1b5fe4d5d7df0df370ece47d19dd0d7614e1efcd32e18028190965

Observation b56d0797-9ee0-43fc-8cae-06e47e8fcb9a · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:57:09.524296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:52:43.674969Z digest=sha256:2c33fde735261e688ab5a95123c6e2c43b4977dfd2bef9d23c34a4d54f8b95e3

Observation cfe4d501-f8f1-44c0-b010-d17b2a2fdd09 · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:29:28.677881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:28:37.680681Z digest=sha256:0eab8f803e609a9c79049fd14fc1d0000c5e96f01c6fd0f255b3c5aa5f356b08

Observation cc0f7559-f10b-461a-b8f7-da4e4af1109f · inbound

From Web to Pixels: Bringing Agentic Search into Visual Perception cites this paper.

From Web to Pixels: Bringing Agentic Search into Visual Perception VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.870026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T05:47:43.959052Z digest=sha256:d5bbf55a58ec1d7b6ea91e7ec1e85f93140bb6fd8417ebae63f362ed3b109d72

Observation 8612d913-230e-44a5-8cd4-73577d5900c8 · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 116

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:19:28.539146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:e463a37a100b76faa63a2395870d33779e5b5bd97a745cde773124fbdc8df631

Observation 2780cb40-bb28-4db9-a1fa-814241f6ab2b · inbound

Dual-Pathway Circuits of Object Hallucination in Vision-Language Models cites this paper.

Dual-Pathway Circuits of Object Hallucination in Vision-Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:19:23.972127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:18:41.987533Z digest=sha256:cc1b17d1f45359176bd5eb2afc52c8fde57d303f711a275361895df94ec3c467

Observation aa8cacb8-9afb-4e41-bea9-bf05ffde4c41 · inbound

CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves cites this paper.

CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T05:39:48.032650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T05:35:02.980473Z digest=sha256:38fa614d7777781a22734e0567f089f73f33b434c4b94ed310c50c932ab18fbf

Observation cc8e34e9-c739-44a8-b4e4-9ec6fcefc36f · inbound

CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves cites this paper.

CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:43:43.454906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T20:40:47.750678Z digest=sha256:0f5781d320ef8630b8d9f3250a0c8724f6cc1e7c2137dcf0e8e00ed6c361f460

Observation f748a5f6-7558-4d9c-a256-802773430a75 · inbound

Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination cites this paper.

Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T19:33:42.146995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:30:21.814764Z digest=sha256:23785268b924b9299323c232147497ed83444ea97150ebc297273a2cd11da4cc

Observation 3e34dc19-2157-4cb6-a67f-45f662c6b6cf · inbound

Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination cites this paper.

Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T19:55:01.633954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T19:49:05.008298Z digest=sha256:496adc92307aab979d3b4ee9ab447527514fafc753ac67ce7aee226ae74f1df9

Observation 0755c88d-71fc-4f4d-8870-a80a1415419a · inbound

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding cites this paper.

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:43:38.806515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T18:39:11.904941Z digest=sha256:4842a0f7f24bdae7d0cfb9f7267c83a6eda51e9c51a93b4faa6882542725b98a

Observation 6d2089e9-74ca-4f76-801f-e0c3b3565e8e · inbound

Segmentation, Detection and Explanation: A Unified Framework for CT Appearance Reasoning cites this paper.

Segmentation, Detection and Explanation: A Unified Framework for CT Appearance Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:23:37.672397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T18:20:38.720544Z digest=sha256:c1edfeb6a7955f7df767c9913f98dc3e64984db5d5296bb3972231f017664ef3

Observation 390cf3a5-ba9a-4ac8-bb4e-a3208d65369c · inbound

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation cites this paper.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.071277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:013930ec2bbc3aa79449c84ed868691c56485839a8fe95c7fc895a9f41ad9b95

Observation e0f8ecf4-68d3-45a7-baa9-88106c6b4874 · inbound

ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned Representation cites this paper.

ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned Representation VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:33:42.052215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:31:12.594807Z digest=sha256:a57bcedf8aaa6af45b70eb50b11687611f10d0a4708f32a2ff77ab1b027b74c8