Pith. sign in

Paper Citation Record · LEDGER

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

As of 13 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 16 inbound Pith citation observations for arXiv:2505.22019.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22019 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:24:54.668238Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:23:34.984840Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.135392Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69531200-78c5-409b-9075-705dee559482 · outbound

This paper cites write newline.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:48.648661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:48.648661Z digest=sha256:28aaa26fb9fe01cd0d381402c0c5a51e41524579a043a32da43167093cadeccd

Observation e8ccdf16-eff3-4b06-8bc0-390b31df102e · outbound

This paper cites Qwen2.5-VL Technical Report.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:48.786792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:48.786792Z digest=sha256:11f09cfddac4d1914c33800cfebe460cbe47bcf83a074ec0a662494f652966e5

Observation 268315af-74ce-4a4d-8877-217f37f1bfc5 · outbound

This paper cites Benchmarking large language models in retrieval-augmented generation.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Benchmarking large language models in retrieval-augmented generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:56.946878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T13:24:48.906125Z digest=sha256:3e5ab1923143a28c9965cc759993e85b4c6b0f986f3dee378e929c2884356ef1

Observation 5baac46d-a7d7-4111-b649-57ccaf6cddc8 · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than \ 3.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning R1-v: Reinforcing super generalization ability in vision-language models with less than \ 3

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:56.741210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T13:24:48.998331Z digest=sha256:2c4164817865fb7896290cc3c8eca71b32d8f9fd790c42c0123ec7efe75d9549

Observation cb69defb-b57e-41c1-bb83-ccf06e98986f · outbound

This paper cites ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.147441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.147441Z digest=sha256:3033d875b3b5d3d1a7698e78690887d4e6032fcaf318ff32265480deb319aea5

Observation 0af78db5-411e-4aa0-afe7-7ceefecb0d9d · outbound

This paper cites Mindsearch: Mimicking human minds elicits deep ai searcher.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Mindsearch: Mimicking human minds elicits deep ai searcher

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.352425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.352425Z digest=sha256:232976e0a8d63504de262f61dc69fe637bdfda5ae54e2aca92e0ef49c19eafb3

Observation 762efeb2-6d0c-4bec-97a0-d036d833bbeb · outbound

This paper cites Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.490774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.490774Z digest=sha256:65220e8f00de69316521b292593bd35d875d1d52647c675d67e7e2717fee2882

Observation e9790918-9649-4cc5-ab71-331c48752194 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.709222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.709222Z digest=sha256:144a1dbaedafec252b02dcd5156293dd7d7927712f0b8128dcd255b844266336

Observation d32f2860-0616-4982-b3a8-d69e8cd7924b · outbound

This paper cites M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.927205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.927205Z digest=sha256:b4c9496f67e5bc9afcbd1273d4714b437f1b40cb3d4b42d8f6bffe23afb100e9

Observation b8925cca-2976-4e05-b1cd-19f0c477861c · outbound

This paper cites PP-OCR: A Practical Ultra Lightweight OCR System.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning PP-OCR: A Practical Ultra Lightweight OCR System

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.065723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.065723Z digest=sha256:e903a8264582f64e539ffdc67b686722b4632b6f48e4c27247b4a47b8abf0328

Observation acb7fd87-8eb6-47a1-bd8b-4ed82d8b5d9e · outbound

This paper cites Colpali: Efficient document retrieval with vision language models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Colpali: Efficient document retrieval with vision language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:56.538027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T13:24:50.220936Z digest=sha256:b97c111ffcc16489fd0e45dbcf8dfa751afea6bdce7a2aa0a67985052daa0ab5

Observation b48724e1-3d34-4973-8f85-220a4d7e13b9 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.367366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.367366Z digest=sha256:87a2100301872d82d5492b302da3c0869491db0c02d174e842bc3a068429ada0

Observation 07fc34a4-9d1a-4da1-a76d-6a940277e99b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.460221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.460221Z digest=sha256:2e92586e0556f7afff66ff7caadbe3cf3216dba22b182ce91031cd264f960964

Observation 7ff22933-ee18-42ad-9a6e-3307aa72aa7c · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.559672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.559672Z digest=sha256:c7a0d886125b89253e0dd12da05541ee0c38eebf5970bd696479825340a33cf8

Observation 00eeeb99-9d5a-4aeb-9c23-ed590f381a76 · outbound

This paper cites OpenAI o1 System Card.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.683709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.683709Z digest=sha256:d666934f08b8505f08ebdb281652fedfa4541a9d5e143923bfc8375102e81145

Observation 00e16907-51c7-4867-a3b2-3f372a8b503a · outbound

This paper cites MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.840704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.840704Z digest=sha256:8b8b5c3cca690aa7cbe9043a211ba0df200ad2762810e3f10ea219fbf50b01dc

Observation f7548be0-6977-4d9b-8ed3-e540525ea30f · outbound

This paper cites DeepRetrieval: Hacking Real Search Engines and Retrievers with Large Language Models via Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning DeepRetrieval: Hacking Real Search Engines and Retrievers with Large Language Models via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:50.943846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:50.943846Z digest=sha256:843eaff82a858da05569fba5ada0cd617dd1ef1a7ba8c44c6dc9b72203d23469

Observation 0914aadd-44d0-486a-a3d4-d761414d506c · outbound

This paper cites Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.082163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.082163Z digest=sha256:e950ad3603cf8131a7929a45f2cec5aa3d3cf6f81339f5d126b4f65b75ff20fa

Observation 9ead5704-8345-4c87-852f-5f9e97d29a78 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.217474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.217474Z digest=sha256:3bc2eee3b938ea18ed1844e2634604c86e89ae09a7ea62b16b88d3872bdf864d

Observation 8e205582-82cd-4269-8724-f25b68dd7ad2 · outbound

This paper cites Reinforcement learning: A survey.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Reinforcement learning: A survey

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:56.351090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T13:24:51.389300Z digest=sha256:56a3becb76ef717bfb6a22aa5d3caef02bee7b6120e4561cb22b2eafd930fe20

Observation 6d2d3ac9-3765-48a5-8255-04acfd34cbbf · outbound

This paper cites NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.601363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.601363Z digest=sha256:ca6ee02969d2c9a9b736da077bcdf9f8d057a724f6158f77d58e420c5d556707

Observation c2d12a3a-869d-4955-84ef-bb3ddd92a9cb · outbound

This paper cites u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.700223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.700223Z digest=sha256:25e1765e78cdd097b532031fbc8ad6be88eb662b7b07418130e9c93d7d2a0de2

Observation b0219256-bd67-4df6-9a69-389dde6c5c69 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.824122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.824122Z digest=sha256:fb47d9720464dd245054f1e5fb261cf86825b0392aefcf3f57ced3a2159d8d5c

Observation 11459a76-2481-48f2-b128-0ef82e496bae · outbound

This paper cites Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.884343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.884343Z digest=sha256:10c721c3d7371d9a07e1672eacf3ea891433cc28340eaf1fbd6595274f4a3d22

Observation 8b9ceca8-29b6-4506-b59b-b18646e7e1e1 · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:51.963166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:51.963166Z digest=sha256:53376891f1893c45367e373aa5a0564e042ac50e2afbef67fc2e92847fed3d7c

Observation e8acbfd5-bd5a-4648-b7b9-5e14039ded2f · outbound

This paper cites Improved baselines with visual instruction tuning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Improved baselines with visual instruction tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.036227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.036227Z digest=sha256:8b66df2a21760aee7324353cc64457dfcac80d6d35593c2fea713c421372bed0

Observation e0f17fc1-8a65-485b-ae63-665bbe9a6844 · outbound

This paper cites LlamaIndex , 11 2022.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning LlamaIndex , 11 2022

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.146606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.146606Z digest=sha256:ed5f506c7e6a9fd334df0df976818aa23f85d036927860755ba85f247ec86d53

Observation 02dc48b8-0155-4e5a-a008-61cb42d6d12e · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.248859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.248859Z digest=sha256:d2f0c903b4f637df256918aa8e80f21d784832ddfeea32f361fc856a655da478

Observation c0ed3f2c-aa1f-435a-90bb-bee0c89944d7 · outbound

This paper cites MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.343132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.343132Z digest=sha256:32c83e502d13454386b1444424479dbad63e5bbb50541f80b6ecd53914fae814

Observation d634e7bc-32f8-43f8-a76e-bec59d2d6849 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.421805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.421805Z digest=sha256:8d337c53a2b348526fe0457e8cd344ab3d84c031f30633589cdbfa2e267dfaee

Observation 2c2ea2cf-8642-469e-924b-799c5c19984d · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Simpo: Simple preference optimization with a reference-free reward

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.518896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.518896Z digest=sha256:467ecc248be75b0b2db3ca295988596816aaefceb33bbc2df81af991f803b811

Observation 4c8c0d4e-b367-4663-93eb-1233a01ea3dd · outbound

This paper cites NV-Retriever: Improving text embedding models with effective hard-negative mining.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning NV-Retriever: Improving text embedding models with effective hard-negative mining

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.610720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.610720Z digest=sha256:4a04e210e057f9fbb9966fd4bff723983673a4e3c927dd803a8d907c2dc6296b

Observation 4dd90c39-b902-4ffc-b08b-f7206b3b0c2e · outbound

This paper cites Hello gpt-4o.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Hello gpt-4o

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.679894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.679894Z digest=sha256:a82f4a638b63a70f9914817613c5aa7437bd1fd4d951d5732b5c3fce44c7a048

Observation fcafe132-9e4f-4759-bc04-037d07ee81cb · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era, 2024.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Introducing gemini 2.0: our new ai model for the agentic era, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.763882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.763882Z digest=sha256:ba86b2e1af1def59de23a327bce495a3ed2dd90dde004808ac5e627bf2ed89d2

Observation d251e266-f212-41fb-bcc0-cba2aec026cf · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.861519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.861519Z digest=sha256:60dd5f2209fb6e40f9d5b9a01c56e455c70fba71ef1cd4bee27fc5254dbc17e8

Observation f33cf807-ebaa-485e-92ef-8f9286a003c8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:52.950386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:52.950386Z digest=sha256:4424ddc8c1a1725b9a82382913d9bd4d30bb9d9e41061d4d732b08fef039c7db

Observation 65ebfb42-2e5d-49bc-9dfe-29506b85d40c · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:56.100544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T13:24:53.041498Z digest=sha256:64656395a6b53ed38f5d5935080804487eeecc85be53146a108cc247f7df5889

Observation d74a9cfd-9ee9-48f7-b84d-946c85862326 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.171877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.171877Z digest=sha256:9481014cafe3ce71bddcecce512d1ff2f4adc3a6e48ff4156ac265a32fcc9c46

Observation b86812e1-643b-4b23-82d6-b42c127aa112 · outbound

This paper cites Reinforcement learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:55.809640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T13:24:53.264735Z digest=sha256:c304cfafc44317aa7bb2f5a2a58ca9b286ff03e9b800875f8a178dd82f718843

Observation ebfbe6e8-0a78-475e-9803-6fd86710633c · outbound

This paper cites Slidevqa: A dataset for document visual question answering on multiple images.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Slidevqa: A dataset for document visual question answering on multiple images

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:55.604920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T13:24:53.364738Z digest=sha256:9f44a320224d52af81be1c595351fd1bc2f11a3ce3295d5dd999c9a5b8f30988

Observation 056425ae-5ace-4508-a89e-5f31bb5c7e35 · outbound

This paper cites ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.442043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.442043Z digest=sha256:38863f84903701854f5a8fa481648f2ea0cfac3639804604b2119de56e98c94e

Observation 6ebaea24-4713-4d9a-a30b-d63e2cd79fc3 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.487266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.487266Z digest=sha256:a13c00e4223f6cbc7b007e1da8aef98689e94b9fda859bfa784bf09ea5c7b1b6

Observation 5bd4bdc1-02f8-4d2f-98c2-88b1d53c7328 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.586633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.586633Z digest=sha256:899f54268052f94c38acba1e082977d2e6b7966ca5943d1d9deca9b586a91a62

Observation f15666ee-bdbc-4691-9a14-5246d2f099a3 · outbound

This paper cites WebWalker: Benchmarking LLMs in Web Traversal.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning WebWalker: Benchmarking LLMs in Web Traversal

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.713826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.713826Z digest=sha256:57768d63276c5bb9780eaf1ac62ce92519351b57c4118917dedd3979423703f3

Observation e328ef64-c72b-47c8-8eda-6451b5932802 · outbound

This paper cites Unfolding the Headline: Iterative Self-Questioning for News Retrieval and Timeline Summarization.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Unfolding the Headline: Iterative Self-Questioning for News Retrieval and Timeline Summarization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:53.810076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:53.810076Z digest=sha256:d0811043943eb8a9c51c799d71766135c154af838668591bd9314554ca0ce214

Observation 04b2de5d-2172-4ee4-baca-17c71b3a49fa · outbound

This paper cites Rule: Reliable multimodal rag for factuality in medical vision language models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Rule: Reliable multimodal rag for factuality in medical vision language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:24:55.380239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T13:24:53.914447Z digest=sha256:d892ec9f43c1827c8ba1378280ec2d63255654a75ca58483cf8869788f5b6924

Observation 3ec533b4-47bb-425c-bc96-4f0808b4965f · outbound

This paper cites Qwen2.5 Technical Report.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Qwen2.5 Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.021458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.021458Z digest=sha256:5ee5cec3a96dae36f71762be0aab1e7613254f74951e818fef98968912096bf1

Observation 161fa363-a68a-42c3-8582-5de71460e329 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning React: Synergizing reasoning and acting in language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.132125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.132125Z digest=sha256:1f34b313733255fbf81ad3b3c153714c5879253e73f44af76c4339aca63705f0

Observation 512576cf-8bdd-4828-88db-3d113f7ce37f · outbound

This paper cites Perception-R1: Pioneering Perception Policy with Reinforcement Learning.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Perception-R1: Pioneering Perception Policy with Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.235680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.235680Z digest=sha256:061151a9459e250a5ad59252c2b7dd79ec34d5aa7274b700c709ca1bec783050

Observation 59e799c9-745e-4f80-8775-79913f385c70 · outbound

This paper cites Introducing Visual Perception Token into Multimodal Large Language Model.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning Introducing Visual Perception Token into Multimodal Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.345474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.345474Z digest=sha256:b9369ba3bdbaf43f6e13fb23d84b2b588511a107fafd3d022afe91acf9672143

Observation bdcf5e24-caf3-41f1-8888-177af150b37d · outbound

This paper cites VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.427481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.427481Z digest=sha256:954104b052f9b15367033dd621bf3b976d117ae01989464ffeb01a995a4e14ec

Observation 23622d0a-3561-49b8-9458-c3817c6b7f97 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.491170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.491170Z digest=sha256:fd3fe13b138368dcb0cfdb594fba8dfbe90a1c553bc7d30d574215e0617184ce

Observation 7d0bfe84-fb56-4288-a145-74e854cc139d · outbound

This paper cites , " * write output.state after.block = add.period write.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning , " * write output.state after.block = add.period write

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.566124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.566124Z digest=sha256:fbd8506db71ea0c35877f49ecf8d58996d2eb80d30f45d9f3804c9daa701e744

Observation 2d2679c0-1349-4812-9f51-95df8bf0e25d · outbound

This paper cites write newline.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:54.668238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:54.668238Z digest=sha256:a8912e8069e9ed86a90436c3f6e14d7fae01dcb4ae7185012e32bd1b46716049

Pith citing papers

Observation e4dc91ff-51d2-4752-aa5a-47c8c4b80e89 · inbound

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback cites this paper.

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T13:23:34.984840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:23:34.984840Z digest=sha256:040ccb0311afe2b36b2d2700c47626c831eff60eda89c5762dfe3b84227f8261

Observation bf6bdab8-5b24-4a69-9e98-363b25af0fbb · inbound

M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation cites this paper.

M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T22:52:00.547279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:52:00.547279Z digest=sha256:fbb23de8a079168bc03840c19412b7f7fce99912cadab8f36044953c65eada3c

Observation 1a8a7c1b-2d4e-46f7-b0c4-35065cf68ee2 · inbound

DeepEyesV2: Toward Agentic Multimodal Model cites this paper.

DeepEyesV2: Toward Agentic Multimodal Model VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:32:29.448501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T05:32:29.266583Z digest=sha256:ad5df4f8f4abfd77462568209de9e7c30d327fd8153d4e127d62a93524917256

Observation 2967d069-f60e-4f57-8581-64b9eed36b05 · inbound

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces cites this paper.

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:33:02.362467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T17:31:08.575993Z digest=sha256:4348ab2433702a69a26d607f64e9c897f1b48d07dd535378ecf53ecd7362cc67

Observation 97ad0c2a-e9dc-4d15-8928-993a8a37b856 · inbound

ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment cites this paper.

ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:15:58.159597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T17:43:15.630570Z digest=sha256:aaad7a71bdcfc238e432e077a63b8bc063ed7ab920f562c6f1cf7fa4a487da2f

Observation bec402f7-7a52-459e-b962-8d50ea8d82cf · inbound

VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning cites this paper.

VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:11:03.331597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T17:45:21.088097Z digest=sha256:600615cfec8663250d3e3eb3e5cbcf80caac32631d2c029ea33ac4765cda5970

Observation 36ea29d2-7175-46ed-95d2-4625e83a559c · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:10:26.498983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T13:09:24.304696Z digest=sha256:6d6b909a5072a4db137cacd1715aff711f0b83da87ede6f2313a85d362bf59e8

Observation ca24bd0d-f4ba-4938-9eea-6f75e42bbbf8 · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T16:18:24.497209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:18:24.497209Z digest=sha256:0dc6e065b3cabce7bd113ace54d60c8ed84a13015c4e7cab1b26b5b85e05915c

Observation 37e43759-ef2d-4700-bc61-27bd5aa2e355 · inbound

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing cites this paper.

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:46:49.213351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-07T16:18:35.484800Z digest=sha256:8b53070ec0ef3a002562a859b6d1d413ebdd0cb2bf11ea3f4830d5c8f8a33a44

Observation ce49c680-a71e-42ed-919f-fd2c54d9a638 · inbound

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory cites this paper.

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:13:40.493453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T19:11:15.761831Z digest=sha256:d616fc2d1d27cb10b99bd5bd8b7b7c2c51404543273bac7b01e95baa35eeed74

Observation 39c126c9-dc8c-43cd-8c4d-2adf8647af8b · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 215

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.136829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:b99cc5316e943c6a53e732d51132a811ade2f3ee3c8f9eed29514117c29923fd

Observation 6ad6b6c6-2c8d-42bc-8c53-fde35b53ed31 · inbound

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search cites this paper.

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:55:41.069237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-01T06:02:48.532478Z digest=sha256:d706b778cc1af4ea7ece06a3887a9f6efec6bb0426040c0e2ecfff4a2323a017

Observation c51fda0a-6a8b-47ae-ae27-03d8243df2b8 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 163

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:662fe35cb2f25a331cb55b232d2a21e8cf1b91cf1628e73266e6eddee0d08d19

Observation bb841e23-0e1c-448e-9519-cdee7432d478 · inbound

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis cites this paper.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:eed61db88f012f8996d7d9bb9e1a3b0aaf4dea739ff9ccabe2e095822ebe9c97

Observation b77f11b1-eb43-4aaa-a7ba-0da6aa9f7156 · inbound

HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering cites this paper.

HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T02:57:49.554753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T02:57:49.554753Z digest=sha256:21335f108e908f8b80587de0e4d70c1f99370f53dcfd898ef15dcb5fe05aa32d

Observation 07c32dab-92fe-471c-a56d-9a84f9aad946 · inbound

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent cites this paper.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:27.203082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:27.203082Z digest=sha256:d30e84cd712b36ace33ec5e8312c406fddd9591149054a993c7ea942db47b12f