Pith. sign in

Paper Citation Record · LEDGER

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

As of 11 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 100 inbound Pith citation observations for arXiv:2504.05118.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.05118 v3

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T09:36:04.735688Z

measured 130 of 130 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 100 of 120 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:53.603007Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact15
  • verified fuzzy14
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9c526576-d64f-426d-a7b5-d38ad9a8d9a7 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-13T09:36:04.752542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:44141a4e95a12c9346bd4d67f4b72dfa40ae6e9eb29b20c6f701e0353660946b

Observation 688ae847-41cf-4bf9-99ad-4d51fd686aa5 · outbound

This paper cites Claude 3.5 sonnet.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Claude 3.5 sonnet

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:36:04.802369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:c2ba46cbde26d8bdaa64e5062b0206155cd8bc73a94b4adb1745b19053d42500

Observation 00603395-7d44-44e1-8647-37c0f2f2a5a2 · outbound

This paper cites Language models are few-shot learners.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Language models are few-shot learners

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:36:04.794506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:ea644fffe22b3ed05ddd17da9d1cd51cafb2da41d5a104c55213f364c10bf888

Observation bd1a6830-eb40-412a-8e10-317612742d39 · outbound

This paper cites Palm: Scaling language modeling with pathways.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Palm: Scaling language modeling with pathways

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:36:04.796897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:ad93d30c8c04fe16f0f27d4a6c4d92a7c10e47067ca2eee26419d2e2d67700c0

Observation 54bfdffd-2437-46c1-a8a4-0f7b05e304ba · outbound

This paper cites Gemini 2.0 flash thinking.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Gemini 2.0 flash thinking

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:36:04.798774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:d9cc9fb3f35c341678c8bc0375516c0f7f4e54ae46f512b9ac861daebf75fe3d

Observation 4f456f52-de83-49c3-8b8e-8e5b282cc5e4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T09:36:04.755391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:ab023734230968c54ef561521d24ad6bb105451bcf631d2fcd43d03bfa0400d1

Observation c3c91aa5-54f2-49a9-ab0b-555645f948ee · outbound

This paper cites Fletcher.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Fletcher

Reference 7

Resolution
verified exact
doi, observed 2026-05-13T09:36:04.749395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:a57db1c1861812d0358dffc14288cc9e0936a93fd13b8440e17a14815cef045f

Observation 58a0cc0a-0bd5-45a3-82b7-ab3fd0e67920 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T09:36:04.758229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:b5c6df0c218f95b21bb7642b609355ee331fb92f90164dcdcdae57bda9babd8b

Observation ef616619-62d7-456f-a74b-5ad3aa7cc012 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T09:36:04.760949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:6045ee54cb06d7669e589e33559ade26d84c14ccb0da4c81e6547ab448335ef9

Observation a6ae724f-7245-45ab-9641-f226e6dab474 · outbound

This paper cites Buy 4 REINFORCE samples, get a baseline for free! InDeep Reinforcement Learning Meets Structured Prediction, ICLR 2019 Workshop, New Orleans, Louisiana, United States, May 6.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Buy 4 REINFORCE samples, get a baseline for free! InDeep Reinforcement Learning Meets Structured Prediction, ICLR 2019 Workshop, New Orleans, Louisiana, United States, May 6

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:36:04.808364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:52d076a880c6a7e4a89421329f7a641b6bf38ab930390ecaf409c401a1a78846

Observation c5163c66-ef4e-4a53-bf4d-ca3bca05ad54 · outbound

This paper cites DeepSeek-V3 Technical Report.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks DeepSeek-V3 Technical Report

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-13T09:36:04.763589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:a8cef6b05a3158b035517b097072105f4fcb12fa407923cb865393d980452e2a

Observation c951f2f0-aba4-4e01-895c-8ea9ac7923d8 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Understanding R1-Zero-Like Training: A Critical Perspective

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-13T09:36:04.766765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:c31cf76a85958b7f11469e279faab9e007b7029e7ea053ce09c0d5a74a535c9c

Observation 921b92d7-9113-462c-8c72-13811903954b · outbound

This paper cites Real: Efficient rlhf training of large language models with parameter reallocation.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Real: Efficient rlhf training of large language models with parameter reallocation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:36:04.813996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:d3c9f80cbfc8b2c44b39242cfcc42d2dfaa21090f91776296d7f7b09a923f04e

Observation d9a9a1f3-3173-4c61-a174-1ed47c6594cc · outbound

This paper cites Self-imitation learning.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Self-imitation learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:36:04.815540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:71452a769fa091b91ad7595d001e64c9a847b384b6187a93cee4ecba20925e0f

Observation e5138e53-2eb9-4d0b-b972-f5e073b20d05 · outbound

This paper cites GPT-4 Technical Report.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks GPT-4 Technical Report

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-13T09:36:04.769079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:a6f54abbd4ae8bc7a9648e4f42a7f110ed711d19ab726d445e718fa066d647ff

Observation 58b95e6d-9afb-43e4-abdf-6665d147f98b · outbound

This paper cites Learning to reason with llms.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Learning to reason with llms

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:36:04.819038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:d7d16c69398457a35fd7c99c98d4a05df0749124a19cbef23f50c37e5872a721

Observation 2a726f58-9e07-4ca2-a1f1-5c8cdf11121b · outbound

This paper cites Training language models to follow instructions with human feedback.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Training language models to follow instructions with human feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:36:04.800565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:c76e881c2a8ef1b6422b6499b5ec9de4b88b91380c9ee59b7ce180cddc7d50ab

Observation c9d45b4f-945e-4ffc-b1f4-9762d8a36c91 · outbound

This paper cites Training language models to follow instructions with human feedback.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Training language models to follow instructions with human feedback

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:36:04.804265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:8b045ada058e794f6893802f66596b0ceab11fe003016e63117f721873e19c5d

Observation b3ead3f9-b96a-4ad2-91a3-9e7163ea00e8 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Qwq-32b: Embracing the power of reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:36:04.806179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:73157f8a7ff7fcbcbb938f6d2aeb6fdc0608472315503a9667c72bb72dcbf318

Observation cb693696-b6f7-4741-8fc1-9c4f2fab8ab0 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-13T09:36:04.771754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:1c0bdcf2714236ec854939296fc5a0d047387daf123b7359c178c70ccf232a70

Observation 4e99da7b-53a5-44bb-a92d-df802a851e38 · outbound

This paper cites Proximal Policy Optimization Algorithms.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Proximal Policy Optimization Algorithms

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-13T09:36:04.774430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:f724ae07fbde2e3c91f36a6d5050605591d95f730976ef69d995def3e413cc45

Observation 1183393e-47ab-47b0-b3c7-9b97004be1b7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-13T09:36:04.777032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:291b7b7acedca8b1af649e43a23cf8fbd5ee3a02040a3c2cbe2b33a33eb36e15

Observation c8d13674-9209-4112-805a-b9c8a65c2ace · outbound

This paper cites Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.780461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:74d5fcee639a57af136a4503bfd2b57800f02a2844c42ea64f91dfe55ba10320

Observation ca9b81a1-66ca-455b-8290-ccc2cbf5ec36 · outbound

This paper cites MIT press Cambridge.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks MIT press Cambridge

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:36:04.812195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:de69b0ccd9a3ac7d68a335e03b8f0df7fb582d727b61e851810d6583507eb833

Observation 8c62f649-2b1f-405b-890a-38a3d1fd38a5 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Gemini: A Family of Highly Capable Multimodal Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-13T09:36:04.782893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:0641e27ce1f03016a36c79fbd7160d74c629d4a8c43696ef41cd28c567c8478b

Observation 84ff15ce-ebef-48c2-9808-cad69e8cdd71 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-13T09:36:04.786222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:9a8a041622c2b09ce6280e5d73206ab0de5038b3941fb781208b2e3c7054d303

Observation f214a403-d89c-41fc-b98b-d1afd8cce16a · outbound

This paper cites Chi, Quoc V.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Chi, Quoc V

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:36:04.817487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:51dd5020dea17c8f9a17d1a95e081d1194375f482db6a31ff524648166824096

Observation 718a07f0-e92b-4450-aa93-1978f6489625 · outbound

This paper cites Grok 3 beta — the age of reasoning agents.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Grok 3 beta — the age of reasoning agents

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T09:36:04.810518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:195705ec2264cfeb6dc805ba08583e242cb3f5687ea7b98efaffef2b3e53bc7a

Observation 05d07834-058b-47ae-8bb5-ea8b5f0a98fa · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-13T09:36:04.788911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:986f4c1fcdae29334009ead88cb7727db022d92bd07e9c21c1a99e0a270f74d1

Observation 5c60ecf9-6bdd-416f-ad2c-0f8b063fc3fe · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.792060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:2f4110f46c7347c9a8b4fd879e976fd64eaeb560a52e2b10b5a4c01546996c0c

Pith citing papers

Observation 869bab39-fbf9-4e3b-b959-58f9c8efbeed · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 148

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:32:01.334345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:6b74fd54b2bb9c8351c41056b9345cf63a822f4d4f219ede0ca222c193009cf8

Observation 60307ffc-28ad-4d7c-a06a-facc7d2a605a · inbound

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning cites this paper.

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:46:56.959170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T18:46:11.571831Z digest=sha256:23c2e8ea38f39325a0b47bcff25d7e45708b5e9d8c169fe5de34a68d54a94f95

Observation 51145a8d-48b2-44fc-a1bb-22881df61304 · inbound

ToolRL: Reward is All Tool Learning Needs cites this paper.

ToolRL: Reward is All Tool Learning Needs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T00:26:48.575367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T00:26:48.291431Z digest=sha256:26e90726f9aa7d2a9e3bca2c87377707769db1b09edb7a520fdaddd09dc29a0f

Observation cee44b9b-7798-4947-ac7b-86376f501c6d · inbound

Reinforcement Learning for Reasoning in Large Language Models with One Training Example cites this paper.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:04.931329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:49631b5bebaeee96437ec56cce024032f7280c8ea37788b3455dbec4cab38157

Observation 1fdb0580-cf6d-4827-9a89-6ee24940d15b · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 171

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:3a86b6f7c5433688682fd5a2a65e7dcb5834e70ac33e1b96d331e4a97924669d

Observation f78eea0f-e438-4e9a-bad1-a1df752bc050 · inbound

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning cites this paper.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:53.603007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:53.603007Z digest=sha256:5c9b8d149ca7501c9bf340b6a655455470a98110183a418c4687364e3ee424fb

Observation aa196788-e30a-4150-803f-794700ddbea6 · inbound

An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents cites this paper.

An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:54.812847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:54.812847Z digest=sha256:684f2202759dad932beef90a7d2fba00d1947bda8f19a916d5cd559846df648c

Observation 4335d87f-8831-49c9-a802-ce15ccd97730 · inbound

R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning cites this paper.

R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:17.382981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:17.382981Z digest=sha256:917f6f71ebcfe03a2416c99a4d6e03340cb09fd76fe784e09bb57cf893cba98f

Observation 125f54f4-7faa-47e8-8054-6daaba203bc9 · inbound

Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective cites this paper.

Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 127

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:12.862798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:12.862798Z digest=sha256:58d63bef2f38e605781481c4bc8d22501dabe2a9d0c4101205290169c946c550

Observation e52cbf45-a0a6-4b75-9057-32083e2e5187 · inbound

Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles cites this paper.

Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:08.003933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:08.003933Z digest=sha256:58a96b4a77955682b95b35962c1bc04da95a4e140f686e52132c679733e5b6ba

Observation 3209e50a-0aff-4c99-8a20-2a8d9c5610d6 · inbound

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning cites this paper.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.447364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.447364Z digest=sha256:3c8e7850a23ddf17367fe01d01a3de6eb8f44170320036029d634826a05e67fb

Observation f0e7aab7-cfac-4690-8631-f363ae52bce6 · inbound

Skywork Open Reasoner 1 Technical Report cites this paper.

Skywork Open Reasoner 1 Technical Report VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-17T04:26:47.372535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T04:26:47.283983Z digest=sha256:4a9a5589c2c9ea1e75f515d3fd0290d0dc568bbaaacae18541d852c49a6e406b

Observation 0f4fdb04-2f99-40c7-9417-1ec1686aa36a · inbound

Decomposing Elements of Problem Solving: What "Math" Does RL Teach? cites this paper.

Decomposing Elements of Problem Solving: What "Math" Does RL Teach? VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:26.479274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:05:26.479274Z digest=sha256:09a61dc5c2e818779668f898415b3949d10f9c6fbb184bdc4c83e0c343ec99b5

Observation 4fd6d89d-55ff-46da-b1b7-fae958d9c605 · inbound

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning cites this paper.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.020162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:7eb2f8ae658036cabaa8f77ea87a72137f13d2030b3133ec0678910901b27295

Observation 32d60b21-8cf9-41d5-8188-815c95221139 · inbound

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment cites this paper.

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.802104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.802104Z digest=sha256:d22d450e911aaa0d2ca1456bee3c147047ebac900a4fc5cd0e0a291956630f7d

Observation 5816efab-176e-4812-b4af-4645f840aa09 · inbound

SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis cites this paper.

SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.789603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:35:47.789603Z digest=sha256:1552de8bb7955960726faee8a244606968ab0b943c27af6116258e22b8ee8070

Observation 1871f2f7-c535-4ef7-9dcc-a2ad22c02ed6 · inbound

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts cites this paper.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:57.031251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:57.031251Z digest=sha256:0e984152a1fc0cf8b91114d11463b59017a71b32daeffa229e526177d5c8868f

Observation f61d086a-d96b-4c33-a3fe-7d28d77515fd · inbound

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning cites this paper.

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T10:53:02.871689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T10:52:32.933138Z digest=sha256:90f9b2e7ab740b848787e012cbbc97d3f6bbc8b51ed22f6610fc921569b17a44

Observation 487a986d-c299-4028-900a-37dcdba052b3 · inbound

How Far Are We from Optimal Reasoning Efficiency? cites this paper.

How Far Are We from Optimal Reasoning Efficiency? VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:40.307416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:40.307416Z digest=sha256:1849d8e08d3477eaf9def28fc07f37c3704036752eb6f1d0a9080a8215e9ea5f

Observation 6e11b935-4602-4260-818a-7337de2eb813 · inbound

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning cites this paper.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.826337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.826337Z digest=sha256:2bfd8349ee4af1566ab5e2143981a88fd94840e7e03e403a19babdfbf2bec597

Observation f9d31dd3-77a0-4581-918f-ffdecd649ba1 · inbound

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning cites this paper.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.270146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.270146Z digest=sha256:f2d3376c8b9ea9dc6ddedf3afb19c6cb4a5183425ee86c173b9e133cace22375

Observation db25a27a-6821-4c88-840c-c455317f9b0e · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:38.584934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:38.584934Z digest=sha256:6bc38ddb0d54caaaac9e277ed9dd5b9df733d3413a500b09a5eaad081b836445

Observation 97bae04e-4157-4166-85dd-bbe2f8c258f4 · inbound

No Free Lunch: Rethinking Internal Feedback for LLM Reasoning cites this paper.

No Free Lunch: Rethinking Internal Feedback for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:26.517415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:26.517415Z digest=sha256:1040afa577936b1cbc8309d401fcc86c54245c827fc11fc97bbbfa5daae65ca1

Observation 5783c661-6638-4b03-8a33-e79151566a43 · inbound

Enhancing Large Language Models through Structured Reasoning cites this paper.

Enhancing Large Language Models through Structured Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:59:47.748775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:59:47.748775Z digest=sha256:f07cae88d8a8ac9dc00b848f5a401f026d843a4ab7b84e96c89124a576a95750

Observation 70018833-dd4f-4183-901f-6204e44b044d · inbound

Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model cites this paper.

Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:11.615470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:11.615470Z digest=sha256:2a5cc338737195c9daca65034a245bbcc83ff0639dff2e10ef0bb3a1fde63167

Observation 48d27688-5e9e-45c2-a2d9-51bd55bfa345 · inbound

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy cites this paper.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.855807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.855807Z digest=sha256:58f080478f96929041b922501dd32e130edad491d1721c0a6b46f4e67fa74d5d

Observation a892dcce-a0ce-4d20-ac7f-b1ce29ef605d · inbound

Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs cites this paper.

Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:12.002417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:12.002417Z digest=sha256:f00e27840c6d9a24f7c5266f36fb80e4ec9ad5e8357a283e990b8a8f4b5912bb

Observation b2d19da5-4da5-4381-9748-8ff4e7f0b63e · inbound

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization cites this paper.

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:21.180458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:14:21.180458Z digest=sha256:5e2030600d7d142e7e91750e63874bc60208394bb3edb7868c2ce567fbf5cac8

Observation 2bcd14f6-b579-4a79-ae0e-9ebfb33ca07c · inbound

First Return, Entropy-Eliciting Explore cites this paper.

First Return, Entropy-Eliciting Explore VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:14.726477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:14.726477Z digest=sha256:9880b312c252c4d3ccc113961b63949754204af5e5dca560ec9273e89a3afc11

Observation 6db4e52d-390b-4fe6-8ba8-7ddc2ba375ef · inbound

GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning cites this paper.

GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:18.637977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:48:18.637977Z digest=sha256:ff313f173a723e9554d6f3c79605d7ab168f42cd4a14f47c2ac14f96348ae153

Observation 0815638b-df74-481c-8fa8-bdd6e337cb8d · inbound

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization cites this paper.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.611272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.611272Z digest=sha256:f472490bbabe3dff28c626cdfa5899de70b67d14cfa6d5cd624e07665fea4ed8

Observation 3be5882a-233e-48b4-813e-08f3643fe0e1 · inbound

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR cites this paper.

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-21T23:24:26.245092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-21T23:20:45.685446Z digest=sha256:50cc2beffa56adc2e407edd91a036bd625e88c220405a3b7151a22c54e19c3b0

Observation 3f92bed6-70b5-4333-b3e0-f62e07a93083 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 298

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.341887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.341887Z digest=sha256:5884e6ffe5cfb4cdc13ebd1ba1d628b3d23b6cb6e9d2b0122c968db34d683f48

Observation 82ae6b57-19d1-4374-887c-9e3150957299 · inbound

Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving cites this paper.

Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T10:30:44.720568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:30:44.720568Z digest=sha256:991a3db43a875109dc0311e56b61c7c5f995469458f0b0650632f8f3b7cb89de

Observation 5c6ddd8e-ef9c-4744-9388-95e1c3ec7ba2 · inbound

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation cites this paper.

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T23:43:05.878469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:43:05.878469Z digest=sha256:323f9cbbac6577ec97a202c5a3ded98c6b1506319617aa2263d6594623b23836

Observation fdd19e39-4745-4c4a-990f-a61fd25109b0 · inbound

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models cites this paper.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:16.807447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:16.807447Z digest=sha256:b4573910ba60f5ca8ea3b5a4bede48ede9a042be6fbbb29aa71072be8ccb3d09

Observation 7b3fd920-5fa4-42b9-a1d5-0eb6394bca49 · inbound

An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal cites this paper.

An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:25:49.975190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:25:49.975190Z digest=sha256:f302b9fa9b62f526d775083911c312639fa0fb99829de5ed378ea022a1e82a6d

Observation 6eda87ac-b045-4ed7-bb6f-64f75a9649b6 · inbound

InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling cites this paper.

InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:34:23.995667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T22:33:09.674822Z digest=sha256:29234e5ef005eda1c53f7b9d5d84d4e877f93afe788c8fcf32c46ce02ca5adb3

Observation 9b3c42bb-6f90-47f4-a99a-ce216d03afba · inbound

QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting cites this paper.

QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 129

Resolution
unresolved
no resolver link, observed 2026-08-05T17:39:28.845867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:39:28.845867Z digest=sha256:3252384ee0c6c384e796915fcb5f4e9468d297930b036032e7dfd47db57b9553

Observation bcf8172e-f661-4089-b7e5-131ca1103dfb · inbound

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward cites this paper.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.959563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.959563Z digest=sha256:0de7f14a82eb780ec853b1c92b4c66b68375fd26fd01dcaeb60570458c5b4a1a

Observation cf81ad52-1075-4fd5-81a4-a154b849fd1d · inbound

DCPO: Dynamic Clipping Policy Optimization cites this paper.

DCPO: Dynamic Clipping Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T11:43:29.342020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:43:29.342020Z digest=sha256:03b31d4ad658cb56c1208e79d92a46af39e703b5fe369a4de2783ebf67e6e20c

Observation a4c1acac-b9ed-4ac5-a94a-3d44a3cb00b3 · inbound

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning cites this paper.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:13:58.985799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:9918de9ce646d9e82ee3e67181b58815344926760b4e8985c747a8b0d727b02d

Observation b2ab2c70-a02f-4218-9207-3bb352e10aca · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.013620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:36e022cd48b8eef4f8eb7e64d3d7b5f773159dd278bcdc986ea38bdd8c8c4795

Observation 97bdd282-cc34-4964-b295-0a99b8da6273 · inbound

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought cites this paper.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.947268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.947268Z digest=sha256:f5a68d6b5615da9465e1768062f8dc72e15484d6535d29448e23571153187ce0

Observation 9691f506-e6da-4df4-b0aa-c82e74f335aa · inbound

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning cites this paper.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.583503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.583503Z digest=sha256:ebf35acb4de95bed5bb41855cf3f1048e619cd9aa164c3a23bb85a842d361412

Observation 785be2f1-8e1f-4b8f-86da-ff8b432162eb · inbound

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models cites this paper.

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T18:53:02.795969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:53:02.795969Z digest=sha256:3e159eb9de1bbde5cbd72a92a88c04ffacae257d56b61d6e04e56196bb846c32

Observation 7b517741-e3b3-42c5-a97d-37e8906b71d3 · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:11:24.064514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:402483f6c8bbe46c1a2d8d6fd1e8f589dac4c0ad066f27650069c2d8349985d1

Observation aa34e41b-b028-4c4c-87b7-ebd7c5eb4527 · inbound

When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL cites this paper.

When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T20:30:35.546640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T20:29:25.874620Z digest=sha256:b6a28c848c7795218feda663fc4b9bd9d8e62ba2ddd35c6586f1a951e3896e63

Observation 77aaf1c0-0ea5-4a0e-b6c5-5f14e2f4fa33 · inbound

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization cites this paper.

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T09:48:10.988323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:48:10.988323Z digest=sha256:5cc44889592dd025d4bbc969c1b39e5700cb986855d5675121413306bb45b956

Observation 2216c8df-3317-471e-98f8-c4ba2496518a · inbound

The Art of Scaling Reinforcement Learning Compute for LLMs cites this paper.

The Art of Scaling Reinforcement Learning Compute for LLMs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:29:14.058755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:41c9054619a675aea9efcbaef05b07902c5edfa00f90479ba1af384e58a4733a

Observation 80bd88de-c989-4673-bf72-378a661194e1 · inbound

Sharpness-Guided Group Relative Policy Optimization via Probability Shaping cites this paper.

Sharpness-Guided Group Relative Policy Optimization via Probability Shaping VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:12:22.335258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T03:10:57.839146Z digest=sha256:6ca4ed57db11d970a4734d279ab1e74eafd46c449ac850bc1aa8846466476d61

Observation 894c1f32-963e-4af7-a059-60326dcbfa3a · inbound

From Ranking to Reasoning: Explainable Web API Recommendation via Semantic Reasoning cites this paper.

From Ranking to Reasoning: Explainable Web API Recommendation via Semantic Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T23:55:30.784895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:54fe09ea451681fc2fa9015e550d4b903116369cd97d99fd0b81f35fb726c0a4

Observation 99842e8b-521f-4a83-b48a-a42e6b4d92f8 · inbound

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation cites this paper.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:15:26.734331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:f7c2eca1ef830ec6589d7a8372918f302acf15d09cea5041cc5a7790cd9c3f01

Observation 11a595fe-f957-4204-b0ba-d491a1ebb9bf · inbound

Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance cites this paper.

Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T22:36:11.481330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:36:11.481330Z digest=sha256:2a51a3b11cdb002ec3359b413e06e06f1c4ac1eacdc31060f4a5712de1c979b0

Observation 3088faac-586a-4cd9-a193-07003b081fbb · inbound

Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention cites this paper.

Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T17:15:17.715125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:15:17.715125Z digest=sha256:ec848ad4a10e0b9603d82f7220f51ab281d940c4afd3f68dd12ce66308fc3b5d

Observation 42998865-32a2-4ce8-880b-7e501606d255 · inbound

Coupled Variational Reinforcement Learning for Language Model General Reasoning cites this paper.

Coupled Variational Reinforcement Learning for Language Model General Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.954210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.954210Z digest=sha256:c2bbacb572d57f7b5e99a76bfc3b21b4ed0445809fe47d76882dca46b7222ddb

Observation 90439b10-2e4c-44b2-aa78-c761068b8f09 · inbound

One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents cites this paper.

One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T14:20:31.329087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:20:31.329087Z digest=sha256:8284abb916e0064861e9d77e5cdca28b6e91a501f14919d0adf36801d6e169ed

Observation db471f3a-97b6-40fb-bb42-ceaf9af8a6a7 · inbound

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation cites this paper.

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T12:19:35.541027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:19:35.541027Z digest=sha256:d566a2c4bcbea7d61f09649a6564cde5f55cddc327c64d2cb6bb17313da89c57

Observation fd2c1ad8-9fa0-4e99-a721-4d17e4b3f26e · inbound

On the Non-decoupling of Supervised Fine-tuning and Reinforcement Learning in Post-training cites this paper.

On the Non-decoupling of Supervised Fine-tuning and Reinforcement Learning in Post-training VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:48:00.539725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T14:47:52.094353Z digest=sha256:1dc4d949ebd03f6c518cdfebecdf70acb152017f7e3b9fc069aa582f52c2d926

Observation e2bf7a88-6870-4603-8b9a-e0b6ea87f3dc · inbound

Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models cites this paper.

Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:54:30.992080Z digest=sha256:942b63893bc286eafb69dd7c7c6470797a566726c2d7fb8e33ac299011588811

Observation 449a277c-f700-4464-a621-dcb29aee9260 · inbound

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models cites this paper.

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.176606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T14:09:26.842696Z digest=sha256:3ad1c113133c8d7c08d4b0e748478028565c126bca6082c65344d375d4bc6de3

Observation 6a2394cc-416c-43cc-b30c-24eafab51e5d · inbound

Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing cites this paper.

Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:27:36.802347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T08:24:50.793194Z digest=sha256:639bab247e347b8ab080970fb88288eb385eb6b0c28273b4f98fc1f3c7003149

Observation 74b99aeb-af62-4ba0-bd79-f0d7da704245 · inbound

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks cites this paper.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T04:17:27.326104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:17:27.326104Z digest=sha256:6d0c08194ef83e8afe4510330d031c6b78169aa64912ebf8c2bc242c6af4076d

Observation 8e752be5-0dd7-4d4d-a87e-60e8e5562434 · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:14:11.058227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T13:13:13.293921Z digest=sha256:0874659f07fc64e40725c6a2344dd641749a806d760f84c54ea614be8f110695

Observation 64a40906-68d9-4469-b025-1701298f15cc · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:46.362138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:46.362138Z digest=sha256:8aa12df5753e373a7617eeeb087a24e561e2dbdc653f09c1299c6a68f7b060b5

Observation 504cdea3-76e5-4404-92c9-143aeb7376bb · inbound

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens cites this paper.

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:46:43.335437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T21:41:48.690125Z digest=sha256:64c2f4e1f172c6cc4343d43e6d0e69743e3a67f2259647d2a7d9cb5fd2d1328a

Observation 967afb4e-8c25-4313-b5f8-b9c7aa5998f5 · inbound

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens cites this paper.

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T22:52:04.254957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:52:04.254957Z digest=sha256:8729ea5155bf1fa6e66cf073f9f20d3f298a69b2923d5813772a6e586748ef70

Observation 191891a1-b127-4f5c-8b5d-3b0e56f91998 · inbound

User Simulator-Guided Multi-Turn Preference Optimization for Reasoning LLM-based Conversational Recommendation cites this paper.

User Simulator-Guided Multi-Turn Preference Optimization for Reasoning LLM-based Conversational Recommendation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-13T17:28:02.564997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T17:24:18.413112Z digest=sha256:74b441baa4a96fb8e29cc6184ddb076ffbdcca1782670478a77ea619197e347f

Observation b6a64307-21cd-4304-9cb3-6816ef54ec2d · inbound

Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning cites this paper.

Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T15:09:29.657563Z digest=sha256:f95cf8228eb1c9f5b2201d5fddaa049323572e2fd1b6977f67186baa666e10b2

Observation 923793dc-5f35-49b3-a7ac-3c8654fd9e19 · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T15:12:06.985609Z digest=sha256:5260fcb6b8c4e50b5812a411903ddc995624a1468e8c72a74c6b92d8342aaa59

Observation 411907de-312b-4c76-bd8f-1d76d9157ba8 · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T01:04:12.454268Z digest=sha256:78c83dfd9353b22747f73414f782f9cd5851f824ffff96c20dd4edff5c4d0a58

Observation 70ce5b5d-9328-4a41-bc12-fcf2b4694078 · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:7aa6d7022b0799a3502f035e79443c088bc6432c408ad80159948890b5ccb122

Observation 5ab556cc-f1a8-437f-8ab6-23c45982a90a · inbound

GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning cites this paper.

GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T00:14:29.531510Z digest=sha256:63239fc0d0b7b48bf46eaebf1a96b0e1c3714429379f9796c1cff95d07779a3d

Observation 4cdcb9db-4dd2-4cc3-9e98-8167b6bf4829 · inbound

V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization cites this paper.

V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T23:48:32.613988Z digest=sha256:1af9a44b80a748e84a61b2144c1e0f94d186f93c12758dacb739eeb93e7b59c5

Observation 5fde5381-1677-40f0-a0dc-53384c164c73 · inbound

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling cites this paper.

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-07T10:42:27.644514Z digest=sha256:736d1493e984408945755809b43086c451b9f0d59adb300b8e72605b80e37c03

Observation f0372826-3cbd-4bdf-87b8-63565a2a8406 · inbound

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling cites this paper.

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T15:28:14.099465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T15:28:14.099465Z digest=sha256:045abc7cf3f0b58cc62f28cd1059ac5b9dfd0c53d39cc70f5fbf8f05ec7306cd

Observation dc57751a-3ffd-4720-b218-b85a0efa5b07 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:1e5f8a20c5b58b9865b1549f653188c59855d84b46fa92ccf5ad54bc7785f536

Observation 4e618c80-c0f7-44ff-88bd-3904ba08c3ca · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:19b057ee2632cb35b2afa8b02684aea099d90f0addcb33c4194c07c4b87ac71d

Observation 9d369771-597d-419a-80d7-d53f367730a9 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 118

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:02:40.646826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:a610c272184dccc98eb365fedc858e3b57b7380ac89b138893ca2a75158a72bd

Observation 6b88b11e-4781-4203-9539-f00e70edb1cd · inbound

Segment-Aligned Policy Optimization for Multi-Modal Reasoning cites this paper.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:109543657eca8987c9f16ffe984bd336f1f36afb8550042ff1311760cb642b2b

Observation 8e95ec75-48e4-470a-a328-0090381acb83 · inbound

Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent cites this paper.

Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T19:34:22.546508Z digest=sha256:6db22cfeca813b150af2b222b00a49a8eeb24cf7f6feeee69634daa0fa40e1f9

Observation 61cfe08d-8965-4df7-86c5-18d40af3610e · inbound

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients cites this paper.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:52b02941d7eefb4845b386696b9dd4718c4bde5a83ef1a4c6fd92f1aa52d41ce

Observation 69d2a109-293f-484e-8705-1d389312eccd · inbound

Gradient Extrapolation-Based Policy Optimization cites this paper.

Gradient Extrapolation-Based Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-11T02:07:08.792030Z digest=sha256:ad3ee0bc9e0f5339cf7e02f133dbe15a62b55255ead5e397c91474469152d720

Observation 67bfa165-ca42-4238-80e1-a1e5357fe0a7 · inbound

Adaptive Negative Reinforcement for LLM Reasoning:Dynamically Balancing Correction and Diversity in RLVR cites this paper.

Adaptive Negative Reinforcement for LLM Reasoning:Dynamically Balancing Correction and Diversity in RLVR VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-11T01:46:46.055118Z digest=sha256:e781ed0cf9837c6bd7c53c78c62a3ac4d8ac14bd91a6bb836d6616429bdc73b3

Observation 5a6dba1b-5ed2-4efa-a3f6-fd00161f6faf · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 107

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:3de146bf4d751c8b5cd0db93f9aba08077e9e08425c48f9f6781f5eed592acbd

Observation 7dc51aac-f164-4d54-9d6c-0debb28bc28c · inbound

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control cites this paper.

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T00:50:42.836549Z digest=sha256:bae8048c7cb651096023a910d8802f648e169211b501335f3c1acf0fba01526b

Observation df6a9d0c-c0e2-445d-90c0-6dfc2db03596 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:c8ee6958472c95aa1b7a5823a25426729e34ac4001a04f1e0c76809392400420

Observation 9c5d05cb-84d2-4d8a-a027-80237918ffb6 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-19T18:07:42.410607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:ba00f45efd7bda112d50ee85d6a203c6f54bf3888b909a77493da5949a9bdf63

Observation df6eefd8-704e-413a-8445-6a5dd9b8bd83 · inbound

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits cites this paper.

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:24:03.186413Z digest=sha256:f4b2706c93c985dbacd38d41de8f6084d976cb0b44509efb6c589ea8e2ad9222

Observation 6edc0c2f-3e88-41b7-8a6a-1681d33bfdf7 · inbound

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs cites this paper.

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T02:44:33.143247Z digest=sha256:92a53251ea5b1654d34e3973a48a705bcd4d4d0f151e204697e029b158553eec

Observation 68621520-2191-4db4-8138-03e16bedd881 · inbound

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping cites this paper.

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:33:40.994346Z digest=sha256:74f75342708eb844d031930998d2f215b4841f33a81b0bd90b7f528344898b5f

Observation 2202c49e-bac5-4af8-b90b-363913ac368b · inbound

Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning cites this paper.

Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T01:19:49.761472Z digest=sha256:238dc0645bc7815b2b8b4220e7d66f781775b12912e158276b2c9306a9b1151f

Observation e84725e5-8a72-46f1-860c-cefa163fcbdf · inbound

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control cites this paper.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-13T07:30:41.399083Z digest=sha256:947356ac2226f7d3059a33e0e243d3d97d7bd301762435886010d6e1f26c0af7

Observation 313e936e-be72-4dfb-bbb5-d37f5199f9bc · inbound

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control cites this paper.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:05:06.381914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:5ca48126ee4eec2519920ffaf78019181e77e5b784bc235cf49842372b6ae502

Observation 1ceaf6ca-ea55-4278-9bc6-143d28c75a6c · inbound

AIS: Adaptive Importance Sampling for Quantized RL cites this paper.

AIS: Adaptive Importance Sampling for Quantized RL VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:14:52.654351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T03:13:14.384567Z digest=sha256:210d5230de0f20f748561eafb7da3de88b8af4dd4ce793a2b837636d4048bf98

Observation d207b28e-0eb3-451d-b031-739af31a52ac · inbound

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs cites this paper.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.806981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:d80443db8512f9e492b698db6aa80d271bcf126b39e80749f40dbe7ca61002ec

Observation dbf033de-1706-4492-8b27-2e756aea105c · inbound

Self-Supervised On-Policy Distillation for Reasoning Language Models cites this paper.

Self-Supervised On-Policy Distillation for Reasoning Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T14:43:22.093104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T14:42:55.368104Z digest=sha256:e6c885b6d9f2c4324dcb027ae2fe769a15f85d51b365228e2884db9db4134f5d

Observation cf985c6e-659f-47c6-8b7e-349cbab13ea0 · inbound

PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR cites this paper.

PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-21T02:29:25.316981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T02:24:48.872065Z digest=sha256:e8d69887b557da850fdb61e6530bda9c84dae80bf1de47a906763f5332526650

Observation 5aead835-031d-4dfa-9f94-aac20548c730 · inbound

Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals cites this paper.

Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:51:15.656925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T07:50:44.907952Z digest=sha256:780265f0469bd08ca60e28e0fd6ce7d33b74b00e72daf19ee510221835236e90

Observation 1a170f90-bce1-40cf-8f98-67734c5206f0 · inbound

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments cites this paper.

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-06-29T16:53:40.530817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T16:51:36.524194Z digest=sha256:56b8017f1a5ec907ec12e28c708f94f10ad9d88139c6ea3e39f74a1003a5af2f