Pith. sign in

Paper Citation Record · LEDGER

Perception-Aware Policy Optimization for Multimodal Reasoning

As of 7 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 39 inbound Pith citation observations for arXiv:2507.06448.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06448 v5

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T05:11:54.685897Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 39 of 39 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T08:15:57.618864Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact30
  • verified fuzzy2
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a3474fff-fe4f-41f0-9e9d-ccc27fa8be92 · outbound

This paper cites " * write output.state after.block = add.period write newline.

Perception-Aware Policy Optimization for Multimodal Reasoning " * write output.state after.block = add.period write newline

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T05:12:05.365873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:acd072d2c953c5913196aa743c00c814c3e1e7b4bbab677b74f040b7a1b6a052

Observation 359114f6-0ea0-474b-ae83-2668204d5b4f · outbound

This paper cites write newline.

Perception-Aware Policy Optimization for Multimodal Reasoning write newline

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T05:12:05.361811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:9ca0e49b64c31cfa0f63837fdda91ce42e828be3aa04e0ffa4d4ab9dca0de221

Observation c7a6e224-2324-42ff-930b-93ac80289f83 · outbound

This paper cites an unresolved cited work.

Perception-Aware Policy Optimization for Multimodal Reasoning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-19T05:12:05.370389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:1ecb21fa9555a52081b9c76d772a3b04c7a745185e33b7c709c5ea8de428624d

Observation 9d108898-0365-4fcc-978b-20562fbabe09 · outbound

This paper cites an unresolved cited work.

Perception-Aware Policy Optimization for Multimodal Reasoning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-19T05:12:05.358137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:76e6e5da51c83e4138b1c2c0f5907a9321ecab20010644e44978251ef56e55df

Observation aa9c559b-80fa-40ad-9b0e-07b32c2059ea · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Perception-Aware Policy Optimization for Multimodal Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.842916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:3c0be91c239e4a32f62c702fafb82c7b37a2245ece2a4aa20ef4522910f86814

Observation 58688159-628f-4801-a84b-83b277c74bac · outbound

This paper cites an unresolved cited work.

Perception-Aware Policy Optimization for Multimodal Reasoning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-19T05:12:05.393909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:b71aaee3e7c4b8953847cd3c84b2855c4214ea6e992b6c7e6a6e9e488359a54c

Observation 51a208cc-713f-419e-9032-46bfb049c5c2 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Perception-Aware Policy Optimization for Multimodal Reasoning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.819520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:25dac58e8d6db0dee28efd168bcbdc26b0313ab8f984bd49cb9fa06e540da66f

Observation 34cd7699-4101-4461-97f0-3892267b3726 · outbound

This paper cites T.; and Xu, X.

Perception-Aware Policy Optimization for Multimodal Reasoning T.; and Xu, X

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:04.898943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:43ad0bf70dc51dce8e83c4d51c33ee56b3e9f6e11486ba87c99631d62aa16b25

Observation f98ee5d2-74ed-49af-9b10-77eb4a1c9629 · outbound

This paper cites arXiv preprint arXiv:2506.09736 , year=.

Perception-Aware Policy Optimization for Multimodal Reasoning arXiv preprint arXiv:2506.09736 , year=

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:04.905784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:ec4816f8a9f11dedea9fcca9471d536578768095ce985b9c78f49e3b5e7057bd

Observation 410b8d1c-b9ef-439d-9c79-d6971078ea2f · outbound

This paper cites an unresolved cited work.

Perception-Aware Policy Optimization for Multimodal Reasoning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-19T05:12:05.388483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:8d4b4799f61bd0394bf652c874d82d9f2316978727c128a3a0ffb4eee5a5ef03

Observation bb4826de-a215-4f05-b0f8-2f7d09409a9e · outbound

This paper cites MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning.

Perception-Aware Policy Optimization for Multimodal Reasoning MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:04.799535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:abdf71532a6887a7a29cbc6645a07f75ddbf50770cd5c24737677b4ac2fbe4fc

Observation ffb4f5be-d707-4866-97ce-c3ef6144e1bf · outbound

This paper cites Noisyrollout: Reinforcing visual reasoning with data augmentation.

Perception-Aware Policy Optimization for Multimodal Reasoning Noisyrollout: Reinforcing visual reasoning with data augmentation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:04.831599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:f8a29704ca84abe9b04bb2a6afa48af674fb672d317a713d5747797790244532

Observation 7d976c68-be52-473c-a66f-980e932e76df · outbound

This paper cites VisionReasoner: Unified visual perception and reasoning via reinforcement learning.

Perception-Aware Policy Optimization for Multimodal Reasoning VisionReasoner: Unified visual perception and reasoning via reinforcement learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:04.815827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:122f066bf19a5cb38b70eee9553dbd23669049d19fca9c44a9cf8f619535b69a

Observation 73460f25-4ca9-468c-969b-9099993f6427 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Perception-Aware Policy Optimization for Multimodal Reasoning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.794677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:c5372c1acaa33394a6d06cd2996d63df36d77992f1384014fc7ab67798b9120a

Observation bd619933-98de-4453-99f9-7b4100ecbebb · outbound

This paper cites an unresolved cited work.

Perception-Aware Policy Optimization for Multimodal Reasoning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-19T05:12:05.385663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:c1c068e5442e2f6baa56adb6a69dd209553e41f5ebaf22acb1395299f19e332f

Observation 03a99d23-7a15-46c0-9792-07a312aee035 · outbound

This paper cites One RL to See Them All: Visual Triple Unified Reinforcement Learning.

Perception-Aware Policy Optimization for Multimodal Reasoning One RL to See Them All: Visual Triple Unified Reinforcement Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.838451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:cc1bea0aa83a732127e7fda6c956ba02f5bdda0079d6d0e44ab1a368b2b76734

Observation d72a3571-b1cd-4393-b4f1-333dcafb43b8 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Perception-Aware Policy Optimization for Multimodal Reasoning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.803695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:9140b04c61cd4aa06f72e6b5124d6e5b04cbf2a865a734101c4a4ba5e34e06db

Observation 827b5dc7-20b0-4833-bfd8-ba8355db7ba9 · outbound

This paper cites an unresolved cited work.

Perception-Aware Policy Optimization for Multimodal Reasoning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-19T05:12:05.396393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:ecb56a2e6991832fec350eb2b6afbc6a3f9ebc8d7864a594628c139e610f5e96

Observation d8de1717-616a-4f19-986b-1fd583b7780c · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Perception-Aware Policy Optimization for Multimodal Reasoning DINOv2: Learning Robust Visual Features without Supervision

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.866205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:dd170214c8bc48b998ed57a9666f1b2bd32973041baf64942fe5928f6d807462

Observation 1911500e-7472-4bcf-9a28-7987565d2d35 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Perception-Aware Policy Optimization for Multimodal Reasoning We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.847081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:c4a2a6373b4da863185e55fe6223e89335371950b4185e2ae0c1d5c69fb26f2e

Observation eeca7c3f-8da5-4eef-b9ef-73363b91c87f · outbound

This paper cites an unresolved cited work.

Perception-Aware Policy Optimization for Multimodal Reasoning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-19T05:12:05.377351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:0ab1957c55cf6dffb16f174dee1f73a86562f631e15d22e9eaa4dd1b0f30c552

Observation 2132d01b-cb54-44c6-82f6-09ce990bbddf · outbound

This paper cites an unresolved cited work.

Perception-Aware Policy Optimization for Multimodal Reasoning Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-19T05:12:05.382746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:611230b55b1dd9471e29a88e2aa2582bfd5a7ffa752575cc615049ede926aaa0

Observation bd50d1ed-3010-47a5-8dcb-f2e0c771f194 · outbound

This paper cites an unresolved cited work.

Perception-Aware Policy Optimization for Multimodal Reasoning Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-19T05:12:05.373467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:d911ca3c8b495e99d051f78e657e842996879f956cdf77d29970f97a5bcdb770

Observation f162e7b5-4581-4a42-af20-7e29102e4313 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Perception-Aware Policy Optimization for Multimodal Reasoning Proximal Policy Optimization Algorithms

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.854537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:db38a464f129fe6d555dc862a185e83be4acdafccdfd48d2f7e22206eb9ca6e5

Observation ff652852-c17e-43ea-9768-716aef33c163 · outbound

This paper cites an unresolved cited work.

Perception-Aware Policy Optimization for Multimodal Reasoning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-19T05:12:05.391352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:0313804129d291842f2334b63e3ebf8c25ae7ed9b06da8787df446b0fd152a9b

Observation 8498bdca-74c3-4d29-b333-5b79bece4580 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Perception-Aware Policy Optimization for Multimodal Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.811822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:9a6e70c003613e41b60a5892da3ce3fcb2efbb568483545808810a9c3df1df70

Observation b33e7b31-7f1a-48da-a873-cd18218f6cbf · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Perception-Aware Policy Optimization for Multimodal Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.869640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:45cfdee19da0e3e344cf4e13462950484c9c81be8d19942897bb315ea223d93c

Observation f5ed93e1-10c3-484d-805f-5073bfc110ec · outbound

This paper cites Satori-r1: Incentivizing multimodal reasoning through explicit visual anchoring.

Perception-Aware Policy Optimization for Multimodal Reasoning Satori-r1: Incentivizing multimodal reasoning through explicit visual anchoring

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:04.807856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:13b0e8f940d4dca7a320c5c7ac27c4c1eb9d896a59116b00c6cf67e3cfa58f9b

Observation d3897b44-a2e2-4d80-ac28-101316a47de4 · outbound

This paper cites Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning.

Perception-Aware Policy Optimization for Multimodal Reasoning Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.902376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:64bb8662fbfab43a304182621c3c59c8337c280481452a1670d2275e7ba52940

Observation b4364fa8-aa8f-482c-a81e-d79415ed2c5f · outbound

This paper cites Srpo: Enhancing multimodal llm reasoning via reflection-aware rein- forcement learning.arXiv preprint arXiv:2506.01713.

Perception-Aware Policy Optimization for Multimodal Reasoning Srpo: Enhancing multimodal llm reasoning via reflection-aware rein- forcement learning.arXiv preprint arXiv:2506.01713

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:04.886990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:27cf34be017bec8b4ed3cd8d69495c935ba50d9b8745d298ca86a4a4585bc959

Observation 678ba928-3566-4b5c-aa58-e4de9765feb6 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Perception-Aware Policy Optimization for Multimodal Reasoning VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.890706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:cd6c2be788a035ec72ffe292eda309516f3107dd42e626b6aadb1afc20055c82

Observation f68c5350-190b-4d9d-8bb9-663afbc9c051 · outbound

This paper cites Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning.

Perception-Aware Policy Optimization for Multimodal Reasoning Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:04.894974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:2c36c154af33ce8c942eb083f6f3c3af260ae701c7d50d1a617f301ddb71f2ab

Observation af990ec9-b8b4-44ed-bf44-15eb9a715caa · outbound

This paper cites Visionary-r1: Mitigating shortcuts in visual reasoning with reinforcement learning.

Perception-Aware Policy Optimization for Multimodal Reasoning Visionary-r1: Mitigating shortcuts in visual reasoning with reinforcement learning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:04.876965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:202091c7cf56133a39226577a9e00d43e31c1c5ed066c2bc4e46194ee24f3079

Observation 1ca4b3e0-ef2c-441c-b900-48f483a6cbbe · outbound

This paper cites Perception-R1: Advancing multimodal reasoning capabilities of MLLMs via visual perception reward.

Perception-Aware Policy Optimization for Multimodal Reasoning Perception-R1: Advancing multimodal reasoning capabilities of MLLMs via visual perception reward

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:04.880422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:767e699f135fb7466521363ef95cf5c0619bafe668eb52e2d9cfdda6704a17fb

Observation 173c6a40-503d-4c10-b562-64d19537027f · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

Perception-Aware Policy Optimization for Multimodal Reasoning LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.858085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:7b00a09ac015d52ab7cb110fc12c569e9d96daa3e8415de17ee20342e24b395a

Observation 8217ab54-f0bc-4ed4-be13-d2bacca3cbba · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Perception-Aware Policy Optimization for Multimodal Reasoning R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.861860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:4a956a54abe29b57366068474000b4214647451666026fb902508c12051e6e3a

Observation d435ef96-b13f-4d91-b360-be0bc04ef91e · outbound

This paper cites an unresolved cited work.

Perception-Aware Policy Optimization for Multimodal Reasoning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-19T05:12:05.379964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:8624fedce90353e331276c8566166eee8df6f68152775e026c7fdca6fad53c4c

Observation 6189f1c6-228d-434f-9e00-ccf452c3c943 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Perception-Aware Policy Optimization for Multimodal Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.883594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:681db7fbb9ea96cd676a21cd2673bccdd0248c9d0d966c3f625b8686d9e77263

Observation f5254e65-b890-4e9a-8c73-be5decc7e72f · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

Perception-Aware Policy Optimization for Multimodal Reasoning MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.823768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:5ccd849ec2a233b19dc30f38f1b568f8bb3bb432ab5b72ac8df9234de0cd11e4

Observation 64104d96-8d54-459a-8005-2951aebd1b1d · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

Perception-Aware Policy Optimization for Multimodal Reasoning MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.873225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:27c5103f4c49439f43c5e29142a8255456a7e18c1fecae2cdc6445874faf0505

Observation aed0ea4a-563a-456b-8f8f-9e4d7453fce8 · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

Perception-Aware Policy Optimization for Multimodal Reasoning DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.835044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:422fa4e9f280a9bd1c6711812e66941ed62f004926b1aa4a9e700aa4aa75adb8

Observation 31d75e9b-8420-4b1f-b5c2-3efea16f1745 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Perception-Aware Policy Optimization for Multimodal Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:04.827471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:e6fab990ff93b018a9324039fd7bcdceffbcbe2f093f13d378f46e7f70cdea05

Observation abdc0de2-7310-428f-85bd-6b68a72e6217 · outbound

This paper cites ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning.

Perception-Aware Policy Optimization for Multimodal Reasoning ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-09T03:07:59.867366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:11:54.685897Z digest=sha256:e76a2412d72bb3696122c262b6734ef289f470057e39f147355d89a021edaaa8

Pith citing papers

Observation e7d8987a-1e34-4486-b289-3db04d99566c · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.458569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:64d6854a1da21f20cd33870401229f2b0d55e67384d9306b100ade92888226ba

Observation f9dc0cc6-1063-4e4d-b0e1-a2eb60bb9f4b · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:57.618864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:57.618864Z digest=sha256:55ac3bf1480bba5bb5e7cfe5b9729a0a17d0fff71a41443886733ca224fd439c

Observation dc5aac8f-1b8b-458b-abec-5753b23ccdd4 · inbound

CPPO: Contrastive Perception Policy Optimization for VLM Agents cites this paper.

CPPO: Contrastive Perception Policy Optimization for VLM Agents Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T13:07:43.215122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:07:43.215122Z digest=sha256:965da107b7b6858ca71c0c8148ed8fbcadaf899e14505116b6694caf37e340fc

Observation 9ebbf5eb-e594-43d5-a72d-457eca42dedf · inbound

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding cites this paper.

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T13:02:00.155950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:02:00.155950Z digest=sha256:84050f8e4f810c393daadf14c8c12e29e5ae35e521ba7509cbe1423bbcc53b71

Observation 341f5a0e-27d1-46d1-bb01-95fb5bd16c37 · inbound

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding cites this paper.

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T06:36:33.925089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:36:33.925089Z digest=sha256:9b91e7aed0b3b5b32be94ce27c7f775ca748ec39bec94bdaba0eec98253952c1

Observation 82d0a0e6-e459-440b-9954-b295d9485faf · inbound

Forest Before Trees: Latent Superposition for Efficient Visual Reasoning cites this paper.

Forest Before Trees: Latent Superposition for Efficient Visual Reasoning Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:01:05.166099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T15:58:03.654150Z digest=sha256:3ebc9751f06598ce3b0679ca46550b5f2e083b6c445433fb3c1ad64e883b39a9

Observation bc9d2dd7-5bd4-4085-acbe-2f01483f87ae · inbound

Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning cites this paper.

Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:17:26.659379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T06:13:33.315525Z digest=sha256:f88201f59a036d1bedc9ff97e98e7e685ea3c63fe0a2ba8bb0feedfc138bb550

Observation 275b33b6-8eff-42e9-903b-cfa708b08b08 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:16:34.653831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:12:46.385646Z digest=sha256:2ddc021cc11f0e2548f40095ab4d3ff13881ea31bf38c22785f771726c3360e2

Observation 2a9a64f5-7f02-43bc-aea8-a10c3b96c477 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-22T10:31:25.357953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T10:30:06.829915Z digest=sha256:483ca07786ddf0935e09a2f68ffdf853e80111e38cdc14b1b9b772a57812f698

Observation 8bf1fa6a-ae1d-4042-8c52-3df0ba6d862e · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T22:00:10.239223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:00:10.239223Z digest=sha256:eb7a550627e09cf27497221a76102297d7c60bc9caabbc37261506fe7ec291a7

Observation 64d0fd5c-6c9a-43a2-8266-583407ae2783 · inbound

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators cites this paper.

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:58:28.528420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:57:47.657243Z digest=sha256:94ba16745d548b5eaf9b51850096ff2975c66b737842d60c01bc788b082d5aaa

Observation 2e6caec1-8066-4e13-936c-2856a4ff63b7 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:11:04.218805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:46:36.010169Z digest=sha256:68776c50f796b660bcbe0430fb3d63c53b51353b28a986e8f28516b2f5128d1c

Observation bed31903-6d60-458c-bdf1-19a0053bfdaf · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:01:26.702734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:38:06.774877Z digest=sha256:08498affe2feb074b3aa46922d6262e6aa7c872652447061c7c7476e5bdba1b8

Observation fa86347f-5917-44fd-8b0c-296fc21920c1 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:27:29.206516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:26:59.917150Z digest=sha256:7c4b22443720e35007366a860dde6a790f7c149c80af2221c0206a2bba52c0d0

Observation 53b754bb-7d08-4d98-833b-eabd2a17a703 · inbound

Q-DeepSight: Incentivizing Thinking with Images for Image Quality Assessment and Refinement cites this paper.

Q-DeepSight: Incentivizing Thinking with Images for Image Quality Assessment and Refinement Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-10T06:46:37.293875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T06:43:54.201604Z digest=sha256:ba2ce32778f940d6ca2836f4ec5ee169b8a5b376996d9e91118b243342b3d289

Observation c18877d0-e5e9-4621-b704-d2dc5b47f847 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:21:29.969986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:13b2f6eadea8ca0fb00bdd8bbaacc27c479d82474644f0f2276dd0b9969370b9

Observation 52f074b3-f4fe-4727-8cfa-6aefbe46dfa4 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:11:16.774088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:1e23d59e4c3647e14ab48107e9f6c9e28eaf787e7ee8adcd0dc85c7646587269

Observation 1be34a22-7c56-4d11-9965-435fddfb70df · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:02:40.675991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:76cc107bc7ffdaccacd867ff2b491eda464242b63806657c35d3ca6e702c79e1

Observation d9f2fb2f-1d95-4192-a7ec-45ebe732f83c · inbound

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models cites this paper.

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-12T11:01:30.624397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T01:20:54.441367Z digest=sha256:e736a0021baa1afc6577b64aad0ff2fa20c777c734d23085ed0141ba3abe813c

Observation 065c6ebb-74a3-43f5-9b64-e0dcacb1ecdb · inbound

Reinforcing Multimodal Reasoning Against Visual Degradation cites this paper.

Reinforcing Multimodal Reasoning Against Visual Degradation Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:06:25.242664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:37:21.451146Z digest=sha256:cd92aa9dd8d814f21aea1e0d79d34ed19a6bc5f9bbf6e12e648d5d7571248288

Observation 6b048127-87de-405c-ae5f-1eb83418d932 · inbound

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning cites this paper.

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T05:11:22.911223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T05:07:39.571227Z digest=sha256:26ccb4e641ea899257947525fd2ab7f960a02c653f880942dba82ed68bf9a78a

Observation ab8a3ff7-0a6f-49cf-b733-439b665071c0 · inbound

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning cites this paper.

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T07:47:31.624993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:43:50.633779Z digest=sha256:31cbfaf0681a29070e739795c87cd8b5049882e6378c704c352d7d2a64800a16

Observation 1cbdb3c0-b461-47b3-a991-d6213a5a1cbf · inbound

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning cites this paper.

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:41:26.448786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T02:24:12.349405Z digest=sha256:5c37bfb12871f8205e670bbc77628c258b296985ddd07e4a770559ec0b68ff74

Observation 730a824b-9a15-4357-8ae9-130108c9edc3 · inbound

Visual-Advantage On-Policy Distillation for Vision-Language Models cites this paper.

Visual-Advantage On-Policy Distillation for Vision-Language Models Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:24:42.923599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T07:24:22.355378Z digest=sha256:12eaf8e4bf765e3eff9487688fc791e4145ca2cb5c37b0a2367d12195bc75e0c

Observation da83562f-cc52-4b05-b104-82d54b740024 · inbound

Escape the Language Prior: Mitigating Late-Stage Modality Collapse in Audio Reasoning via Modality-Aware Policy Optimization cites this paper.

Escape the Language Prior: Mitigating Late-Stage Modality Collapse in Audio Reasoning via Modality-Aware Policy Optimization Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:47.577944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T17:45:10.339950Z digest=sha256:21a5aa6fcd8bafb28a6791738aa3cef5fdbcd6d1ef4aebedb15050d8efe89687

Observation d55c61b8-93b0-40b2-b23d-f28ea6276ca0 · inbound

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding cites this paper.

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T19:32:34.993488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T19:26:37.281563Z digest=sha256:87d63880daee9a1902682662ab8f74a57c37bd28768696f937586e6dd30f4a22

Observation 516f15e1-89d9-4ef8-a1be-bd009457c7a0 · inbound

Teaching the Way, Not the Answer: Privileged Tutoring Distillation for Multimodal Policy Optimization cites this paper.

Teaching the Way, Not the Answer: Privileged Tutoring Distillation for Multimodal Policy Optimization Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:07:18.029695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:40:02.133241Z digest=sha256:65d5b0f7daae491e0d1131ebcdc3aa6bac301c1caf437b131dddfca7eea7c3b2

Observation ffcadad2-aec4-47fc-afe0-0ef93208dae1 · inbound

DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning cites this paper.

DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-02T20:47:23.104445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T20:08:29.208550Z digest=sha256:e08667b90f2ac8e4fbbefe9aca30286297e6bf4ed544836568fa0a887b42078b

Observation bc345b3a-26ef-4a08-8500-5b42088179ed · inbound

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning cites this paper.

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 108

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:18:57.746659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T01:23:40.564561Z digest=sha256:a43a375886cce798e2af4f4ff795835174b9ed153625eb463bf86d381f162bad

Observation 2d32805b-602d-4a3a-8829-45d73cf98b2d · inbound

CFPO: Counterfactual Policy Optimization for Multimodal Reasoning cites this paper.

CFPO: Counterfactual Policy Optimization for Multimodal Reasoning Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:09:44.512609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T09:09:09.716984Z digest=sha256:ddca09e7d703a5800af4d2f821b929f1ea2d85b9b73296445343be78408746a1

Observation 054f1806-92bc-43be-a0e2-aea11df5e414 · inbound

VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct cites this paper.

VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:39:46.103753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:34:27.719022Z digest=sha256:523325cd6e5da9c69323b30c67189b958aacc31622bf5e523cad1b887fdb66fd

Observation 5a1459d4-fb35-448e-9015-3d5f92e44d7f · inbound

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning cites this paper.

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:20:06.728831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T21:23:10.051805Z digest=sha256:c013abb42ecd0a35779a27501a5cb8a2366038b48e6bbc64da1679a6ddcb1b64

Observation 55ac53a3-914e-4dbb-97b5-a3e494d341ec · inbound

Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs cites this paper.

Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:49:57.363070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:22:26.135526Z digest=sha256:82c0513cea1a36a57196e73db4d1f19b50b283082ab7843453a60960abf81bd7

Observation b10acfd0-27ca-4a6c-bef0-a342f495510a · inbound

RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning cites this paper.

RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T17:05:51.429152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T04:09:32.397341Z digest=sha256:84510405b2b3bcc50a4491e97befea0b8df8d0ed1674f9130165ddda93eedad0

Observation 38cc0a18-c7ba-4451-b6c6-be0006946a7b · inbound

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context cites this paper.

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T06:04:21.272222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T06:01:49.904752Z digest=sha256:8a3381e187700a406973425a6daedf59368d985503789a83558508730e755da4

Observation c0be25c6-b1c8-4810-ada9-a488ef14bfd1 · inbound

H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation cites this paper.

H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-12T09:29:14.105502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T09:29:14.105502Z digest=sha256:f48c1cd4d57d61f16397b9aa6532afeeabb5390dd83c8cdf08c49735ed93ec76

Observation 16c75267-7b7e-4d91-8939-5c284401d55e · inbound

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning cites this paper.

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T03:20:41.123647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:20:41.123647Z digest=sha256:73210c35b16c31164e7ff2b63ebb2f98ebe8dd8cddd5bc7971be387070dcb790

Observation c2a9dcd2-b824-48a5-bd32-77d62a6a058a · inbound

Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation cites this paper.

Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T06:36:36.304920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:36:36.304920Z digest=sha256:94319f85716f1b67678e4978781ade6d567e3439cbc5904a21160e66149d3f5a

Observation cf478290-8401-4027-90e6-4ac0fe3c2963 · inbound

MIRROR: Learning from the Other View for Multi-Modal Reasoning cites this paper.

MIRROR: Learning from the Other View for Multi-Modal Reasoning Perception-Aware Policy Optimization for Multimodal Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T07:15:04.823955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:15:04.823955Z digest=sha256:660666b97c30217e9ac08ce374f6716c98de0cacecce6b7ba4c0688fe0caf657