Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T09:36:04.735688Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 100 inbound Pith citation observations for arXiv:2504.05118.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T09:36:04.735688Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:53.603007Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
30 of 30 outbound references displayed
External citation measurements
2
pith, observed 2026-08-05T02:28:24.338817Z
Observation 9c526576-d64f-426d-a7b5-d38ad9a8d9a7 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 688ae847-41cf-4bf9-99ad-4d51fd686aa5 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Claude 3.5 sonnet
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 00603395-7d44-44e1-8647-37c0f2f2a5a2 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Language models are few-shot learners
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bd1a6830-eb40-412a-8e10-317612742d39 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Palm: Scaling language modeling with pathways
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 54bfdffd-2437-46c1-a8a4-0f7b05e304ba · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Gemini 2.0 flash thinking
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4f456f52-de83-49c3-8b8e-8e5b282cc5e4 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c3c91aa5-54f2-49a9-ab0b-555645f948ee · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Fletcher
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 58a0cc0a-0bd5-45a3-82b7-ab3fd0e67920 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ef616619-62d7-456f-a74b-5ad3aa7cc012 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a6ae724f-7245-45ab-9641-f226e6dab474 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Buy 4 REINFORCE samples, get a baseline for free! InDeep Reinforcement Learning Meets Structured Prediction, ICLR 2019 Workshop, New Orleans, Louisiana, United States, May 6
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c5163c66-ef4e-4a53-bf4d-ca3bca05ad54 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks DeepSeek-V3 Technical Report
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c951f2f0-aba4-4e01-895c-8ea9ac7923d8 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Understanding R1-Zero-Like Training: A Critical Perspective
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 921b92d7-9113-462c-8c72-13811903954b · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Real: Efficient rlhf training of large language models with parameter reallocation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d9a9a1f3-3173-4c61-a174-1ed47c6594cc · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Self-imitation learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e5138e53-2eb9-4d0b-b972-f5e073b20d05 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks GPT-4 Technical Report
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 58b95e6d-9afb-43e4-abdf-6665d147f98b · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Learning to reason with llms
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2a726f58-9e07-4ca2-a1f1-5c8cdf11121b · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Training language models to follow instructions with human feedback
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c9d45b4f-945e-4ffc-b1f4-9762d8a36c91 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Training language models to follow instructions with human feedback
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b3ead3f9-b96a-4ad2-91a3-9e7163ea00e8 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Qwq-32b: Embracing the power of reinforcement learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cb693696-b6f7-4741-8fc1-9c4f2fab8ab0 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4e99da7b-53a5-44bb-a92d-df802a851e38 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Proximal Policy Optimization Algorithms
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1183393e-47ab-47b0-b3c7-9b97004be1b7 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c8d13674-9209-4112-805a-b9c8a65c2ace · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ca9b81a1-66ca-455b-8290-ccc2cbf5ec36 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks MIT press Cambridge
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8c62f649-2b1f-405b-890a-38a3d1fd38a5 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Gemini: A Family of Highly Capable Multimodal Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 84ff15ce-ebef-48c2-9808-cad69e8cdd71 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f214a403-d89c-41fc-b98b-d1afd8cce16a · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Chi, Quoc V
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 718a07f0-e92b-4450-aa93-1978f6489625 · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks Grok 3 beta — the age of reasoning agents
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 05d07834-058b-47ae-8bb5-ea8b5f0a98fa · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5c60ecf9-6bdd-416f-ad2c-0f8b063fc3fe · outbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 869bab39-fbf9-4e3b-b959-58f9c8efbeed · inbound
Reinforcement Learning from Human Feedback VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 148
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 60307ffc-28ad-4d7c-a06a-facc7d2a605a · inbound
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 51145a8d-48b2-44fc-a1bb-22881df61304 · inbound
ToolRL: Reward is All Tool Learning Needs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cee44b9b-7798-4947-ac7b-86376f501c6d · inbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1fdb0580-cf6d-4827-9a89-6ee24940d15b · inbound
Seed1.5-VL Technical Report VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 171
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f78eea0f-e438-4e9a-bad1-a1df752bc050 · inbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa196788-e30a-4150-803f-794700ddbea6 · inbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4335d87f-8831-49c9-a802-ce15ccd97730 · inbound
R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 125f54f4-7faa-47e8-8054-6daaba203bc9 · inbound
Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 127
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e52cbf45-a0a6-4b75-9057-32083e2e5187 · inbound
Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3209e50a-0aff-4c99-8a20-2a8d9c5610d6 · inbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0e7aab7-cfac-4690-8631-f363ae52bce6 · inbound
Skywork Open Reasoner 1 Technical Report VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0f4fdb04-2f99-40c7-9417-1ec1686aa36a · inbound
Decomposing Elements of Problem Solving: What "Math" Does RL Teach? VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fd6d89d-55ff-46da-b1b7-fae958d9c605 · inbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 32d60b21-8cf9-41d5-8188-815c95221139 · inbound
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5816efab-176e-4812-b4af-4645f840aa09 · inbound
SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1871f2f7-c535-4ef7-9dcc-a2ad22c02ed6 · inbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f61d086a-d96b-4c33-a3fe-7d28d77515fd · inbound
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 487a986d-c299-4028-900a-37dcdba052b3 · inbound
How Far Are We from Optimal Reasoning Efficiency? VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e11b935-4602-4260-818a-7337de2eb813 · inbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9d31dd3-77a0-4581-918f-ffdecd649ba1 · inbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db25a27a-6821-4c88-840c-c455317f9b0e · inbound
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97bae04e-4157-4166-85dd-bbe2f8c258f4 · inbound
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5783c661-6638-4b03-8a33-e79151566a43 · inbound
Enhancing Large Language Models through Structured Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70018833-dd4f-4183-901f-6204e44b044d · inbound
Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48d27688-5e9e-45c2-a2d9-51bd55bfa345 · inbound
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a892dcce-a0ce-4d20-ac7f-b1ce29ef605d · inbound
Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2d19da5-4da5-4381-9748-8ff4e7f0b63e · inbound
CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bcd14f6-b579-4a79-ae0e-9ebfb33ca07c · inbound
First Return, Entropy-Eliciting Explore VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db4e52d-390b-4fe6-8ba8-7ddc2ba375ef · inbound
GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0815638b-df74-481c-8fa8-bdd6e337cb8d · inbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be5882a-233e-48b4-813e-08f3643fe0e1 · inbound
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3f92bed6-70b5-4333-b3e0-f62e07a93083 · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 298
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82ae6b57-19d1-4374-887c-9e3150957299 · inbound
Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c6ddd8e-ef9c-4744-9388-95e1c3ec7ba2 · inbound
Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdd19e39-4745-4c4a-990f-a61fd25109b0 · inbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b3fd920-5fa4-42b9-a1d5-0eb6394bca49 · inbound
An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eda87ac-b045-4ed7-bb6f-64f75a9649b6 · inbound
InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9b3c42bb-6f90-47f4-a99a-ce216d03afba · inbound
QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 129
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcf8172e-f661-4089-b7e5-131ca1103dfb · inbound
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf81ad52-1075-4fd5-81a4-a154b849fd1d · inbound
DCPO: Dynamic Clipping Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4c1acac-b9ed-4ac5-a94a-3d44a3cb00b3 · inbound
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b2ab2c70-a02f-4218-9207-3bb352e10aca · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 97bdd282-cc34-4964-b295-0a99b8da6273 · inbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9691f506-e6da-4df4-b0aa-c82e74f335aa · inbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 785be2f1-8e1f-4b8f-86da-ff8b432162eb · inbound
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b517741-e3b3-42c5-a97d-37e8906b71d3 · inbound
Multiplayer Nash Preference Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aa34e41b-b028-4c4c-87b7-ebd7c5eb4527 · inbound
When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 77aaf1c0-0ea5-4a0e-b6c5-5f14e2f4fa33 · inbound
Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2216c8df-3317-471e-98f8-c4ba2496518a · inbound
The Art of Scaling Reinforcement Learning Compute for LLMs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 80bd88de-c989-4673-bf72-378a661194e1 · inbound
Sharpness-Guided Group Relative Policy Optimization via Probability Shaping VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 894c1f32-963e-4af7-a059-60326dcbfa3a · inbound
From Ranking to Reasoning: Explainable Web API Recommendation via Semantic Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 99842e8b-521f-4a83-b48a-a42e6b4d92f8 · inbound
MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 11a595fe-f957-4204-b0ba-d491a1ebb9bf · inbound
Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3088faac-586a-4cd9-a193-07003b081fbb · inbound
Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42998865-32a2-4ce8-880b-7e501606d255 · inbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90439b10-2e4c-44b2-aa78-c761068b8f09 · inbound
One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db471f3a-97b6-40fb-bb42-ceaf9af8a6a7 · inbound
Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd2c1ad8-9fa0-4e99-a721-4d17e4b3f26e · inbound
On the Non-decoupling of Supervised Fine-tuning and Reinforcement Learning in Post-training VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e2bf7a88-6870-4603-8b9a-e0b6ea87f3dc · inbound
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 449a277c-f700-4464-a621-dcb29aee9260 · inbound
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6a2394cc-416c-43cc-b30c-24eafab51e5d · inbound
Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 74b99aeb-af62-4ba0-bd79-f0d7da704245 · inbound
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e752be5-0dd7-4d4d-a87e-60e8e5562434 · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 64a40906-68d9-4469-b025-1701298f15cc · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 504cdea3-76e5-4404-92c9-143aeb7376bb · inbound
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 967afb4e-8c25-4313-b5f8-b9c7aa5998f5 · inbound
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 191891a1-b127-4f5c-8b5d-3b0e56f91998 · inbound
User Simulator-Guided Multi-Turn Preference Optimization for Reasoning LLM-based Conversational Recommendation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b6a64307-21cd-4304-9cb3-6816ef54ec2d · inbound
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 923793dc-5f35-49b3-a7ac-3c8654fd9e19 · inbound
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 411907de-312b-4c76-bd8f-1d76d9157ba8 · inbound
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 70ce5b5d-9328-4a41-bc12-fcf2b4694078 · inbound
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5ab556cc-f1a8-437f-8ab6-23c45982a90a · inbound
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4cdcb9db-4dd2-4cc3-9e98-8167b6bf4829 · inbound
V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5fde5381-1677-40f0-a0dc-53384c164c73 · inbound
Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f0372826-3cbd-4bdf-87b8-63565a2a8406 · inbound
Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc57751a-3ffd-4720-b218-b85a0efa5b07 · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4e618c80-c0f7-44ff-88bd-3904ba08c3ca · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9d369771-597d-419a-80d7-d53f367730a9 · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6b88b11e-4781-4203-9539-f00e70edb1cd · inbound
Segment-Aligned Policy Optimization for Multi-Modal Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8e95ec75-48e4-470a-a328-0090381acb83 · inbound
Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 61cfe08d-8965-4df7-86c5-18d40af3610e · inbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 69d2a109-293f-484e-8705-1d389312eccd · inbound
Gradient Extrapolation-Based Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 67bfa165-ca42-4238-80e1-a1e5357fe0a7 · inbound
Adaptive Negative Reinforcement for LLM Reasoning:Dynamically Balancing Correction and Diversity in RLVR VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5a6dba1b-5ed2-4efa-a3f6-fd00161f6faf · inbound
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7dc51aac-f164-4d54-9d6c-0debb28bc28c · inbound
HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation df6a9d0c-c0e2-445d-90c0-6dfc2db03596 · inbound
AIPO: Learning to Reason from Active Interaction VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9c5d05cb-84d2-4d8a-a027-80237918ffb6 · inbound
AIPO: Learning to Reason from Active Interaction VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation df6eefd8-704e-413a-8445-6a5dd9b8bd83 · inbound
The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6edc0c2f-3e88-41b7-8a6a-1681d33bfdf7 · inbound
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 68621520-2191-4db4-8138-03e16bedd881 · inbound
Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2202c49e-bac5-4af8-b90b-363913ac368b · inbound
Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e84725e5-8a72-46f1-860c-cefa163fcbdf · inbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 313e936e-be72-4dfb-bbb5-d37f5199f9bc · inbound
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1ceaf6ca-ea55-4278-9bc6-143d28c75a6c · inbound
AIS: Adaptive Importance Sampling for Quantized RL VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d207b28e-0eb3-451d-b031-739af31a52ac · inbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dbf033de-1706-4492-8b27-2e756aea105c · inbound
Self-Supervised On-Policy Distillation for Reasoning Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cf985c6e-659f-47c6-8b7e-349cbab13ea0 · inbound
PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5aead835-031d-4dfa-9f94-aac20548c730 · inbound
Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1a170f90-bce1-40cf-8f98-67734c5206f0 · inbound
Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.