Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T23:33:10.824995Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 100 inbound Pith citation observations for arXiv:2503.14476.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T23:33:10.824995Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:44:35.961914Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
41 of 41 outbound references displayed
External citation measurements
7
pith, observed 2026-08-05T02:28:24.338817Z
Observation bf69f2fb-4194-4775-b396-3b04c8c958fa · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Learning to reason with llms
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 82c25e79-80e3-4b0f-b478-56ec6b11807a · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a62947e7-0eac-41a5-8f6b-cdaa2a572af1 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale GPT-4 Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7e2e3413-76a2-4acc-b0fe-7b40996b56d7 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Claude 3.5 sonnet
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d2d6d9f1-71d9-4df4-aeee-716c0f95e2b2 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Language models are few-shot learners
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 781fd30a-94ab-44fe-8414-b59507835446 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Palm: Scaling language modeling with pathways
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f2ddb11d-d609-41ff-9aab-a80de5b2b064 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale DeepSeek-V3 Technical Report
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 45bee8c6-9b58-4c89-ba51-2e1c93245934 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Grok 3 beta — the age of reasoning agents
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 73206345-4efe-4019-9a7e-a589635c6f66 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Gemini 2.0 flash thinking
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b8e585ba-663c-4daa-91ff-39515fadc3e4 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Qwq-32b: Embracing the power of reinforcement learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ed0712cc-0800-4619-b28c-bdfa039ab602 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 325e40a2-7203-4c23-a302-e01e0e0d2c30 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Qwen2.5 Technical Report
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f33de9cf-eec2-4453-83db-1f09bebf594e · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale An Empirical Study on Eliciting and Improving R1-like Reasoning Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ac6bb475-0f2f-4fbb-b2b8-202823f44509 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Open-reasoner- zero: An open source approach to scaling reinforcement learning on the base model.https://github.com/ Open-Reasoner-Zero/Open-Reasoner-Zero
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2ab05f3a-37d7-484c-9821-675e734778c7 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 472c55df-44bb-47d7-bd87-d01c7ce34dbe · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Process Reinforcement through Implicit Rewards
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c4faa081-3ba1-411f-b4ef-6dbde192dc74 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Token-Supervised Value Models for Enhancing Mathematical Problem-Solving Capabilities of Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 531ef94a-12a6-4835-b0ef-ade283579f4a · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 753f2d48-65e1-43e0-b6fa-b36884c6ae38 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8bee5c55-19b4-44cf-9ad3-9094548446d2 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale HybridFlow: A Flexible and Efficient RLHF Framework
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fef29dde-0eac-4c94-8e8d-9833432dab40 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Proximal Policy Optimization Algorithms
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f96911f5-35ca-4ec7-8304-23d64f07c98c · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale High-dimensional continuous control using generalized advantage estimation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6304fbeb-1beb-4138-8a50-af1a54035d54 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Training language models to follow instructions with human feedback
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f3c48e26-a5c1-499b-9a92-0ad379b0609b · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Concrete problems in ai safety
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 976fcbb8-c2c8-41c0-8565-87a155f023f8 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Reinforcement learning with a corrupted reward channel
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 55f8327b-5fd0-4297-b244-9ec0bff2eb58 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Specification gaming: the flip side of ai ingenuity
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2bad6074-0ed0-46ae-bd29-02b1d33a1ce6 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b090dab5-ae1f-4438-9355-0cf49c4b0c52 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Scaling laws for reward model overoptimization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7f0532fb-c2af-498b-95e4-5e53f5f0cb4b · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Reward hacking in reinforcement learning.lilianweng.github.io, Nov 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7f3f4976-6270-472f-a995-54cca691ac76 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Generative language modeling for automated theorem proving
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f776b224-262a-4ee5-ad22-abedf1d49b85 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Solving olympiad geometry without human demonstrations
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9671c03c-7442-49fb-a3af-3a94e31bcfb5 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Alphageometry: An olympiad-level ai system for geometry
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0c30b8fa-a4e9-4f5f-b6c1-e8f621a071bf · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Ai achieves silver-medal standard solving international mathematical olympiad problems
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8c529730-701f-48a3-8912-a09bac9d6ae4 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Coderl: Mastering code generation through pretrained models and deep reinforcement learning.Advances in Neural Information Processing Systems, 35:21314–21328
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation abec4026-a124-4eb8-bfe9-ddb15ca94714 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Reflexion: Language agents with verbal reinforcement learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 36a7154a-69cb-4bb7-9447-d47ad0bde6ba · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Teaching large language models to self-debug
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e4f32f3a-dbf4-4208-8944-6ac7fc227847 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Rlef: Grounding code llms in execution feedback with reinforcement learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fc980eca-952f-4e1e-b465-ddf65d8fdb91 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 14fda047-4c20-458a-a442-4e2da1a09b2c · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Decoupled weight decay regularization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c51a8b68-f83a-4f56-99fa-7e5af95f863d · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale First, note that the answer consists of an integer part and a square root term
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c0e2ce93-1364-4eca-9281-97cb2b63dad3 · outbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale Let B be the set of residents who own a set of golf clubs
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 734e70fe-fa22-492f-b07f-a61db578ec65 · inbound
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5ba38af3-159d-467c-9c25-3b7f99c5bbf9 · inbound
Learning to Reason at the Frontier of Learnability DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 88fb637f-2325-443c-82ff-053f31c17ad3 · inbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b801f8f0-ec35-4899-9195-21f9995690c3 · inbound
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 05d07834-058b-47ae-8bb5-ea8b5f0a98fa · inbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9cc0c66f-1531-4f89-8ac8-883264aa5fc7 · inbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9a6a2c45-d65a-4049-aece-f5cc7120eacd · inbound
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 43a80739-3434-4af7-8f5c-87ab5e5b45ba · inbound
ToolRL: Reward is All Tool Learning Needs DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ce80f86d-eeb6-4241-9a09-21e0f4e91a77 · inbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 26fc22d3-9293-4848-b07a-8d6b20e8bb00 · inbound
Seed1.5-VL Technical Report DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 166
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c7cb80a5-9253-417f-a78f-1b3c45744716 · inbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c9ae3b39-ce40-4366-bf22-bba134d03aec · inbound
Group-in-Group Policy Optimization for LLM Agent Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b778e018-1636-4f23-891f-dbaaaa36b509 · inbound
The Hallucination Tax of Reinforcement Finetuning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa17cf45-45ea-48df-8151-bfd00c7d1077 · inbound
General-Reasoner: Advancing LLM Reasoning Across All Domains DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6fdcaf4-4f08-469b-9515-792345129f24 · inbound
Reward Reasoning Model DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a25740f0-4537-4d6f-8f54-d60862085901 · inbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33aab401-0b34-4dc2-9521-4f7dd023735e · inbound
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d17c69fb-0447-4821-94ee-8b7e972aff5c · inbound
lmgame-Bench: How Good are LLMs at Playing Games? DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41384598-14da-41dc-ae07-b941ab3cd98d · inbound
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 474c2c7b-1cc0-44be-99c4-f81dfbc43091 · inbound
NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9423a99-3bc6-4f52-a8cb-ee85a07db547 · inbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 510d0643-00d6-4fef-aef8-1ab9ad7221ec · inbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ee74034-adfa-4c42-8c6b-7a8371bd8dbd · inbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d5116c2-3996-4605-802c-33106f2444d1 · inbound
O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f137afa9-0a97-4d25-9cc9-51b11a59b0c6 · inbound
R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10cbe391-988a-4570-9c01-550f63ae962e · inbound
DeepRec: Towards a Deep Dive Into the Item Space with Large Language Model Based Recommendation DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e629553a-ec71-4f4c-aeb8-be8a9c61d419 · inbound
R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 249c4d65-1e62-49b1-8920-e501a572bab4 · inbound
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62fb29e7-735c-4b34-9cc0-4231cb58ef34 · inbound
Plan-R1: Safe and Feasible Trajectory Planning as Language Modeling DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd620298-cab5-4451-a2aa-84df904a8cb4 · inbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63cbc4b7-953c-499d-986d-d3fc73d92757 · inbound
Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b71ca5b-2ab2-46d5-8779-29b6180e8744 · inbound
Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a19cde46-1023-442b-802e-b09dc1152cda · inbound
Formally Solving Answer-Construction Problems in Lean DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9a22cf9-07fd-4372-a117-a775ac9d9bc1 · inbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad890374-dd81-4fc5-9731-dd33c7d49399 · inbound
VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1e7a4dd1-a25f-4783-b71d-bac8905ba860 · inbound
On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69316f9c-050b-4ccf-8937-93ae51d826a2 · inbound
Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 50d49ff3-dcaf-46aa-922e-d6ae27d2001d · inbound
SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a53ce49d-145b-4e6a-ac43-cc76e727053d · inbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df8213ea-708b-44fc-8778-90e3ff9b27be · inbound
SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ac65f12-04af-4f88-9429-c02619ba8415 · inbound
Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 125
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 798cd88d-0aea-4720-b162-2f587ba6f6d6 · inbound
Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b333d74e-723b-4e2f-9f79-a3f8e37df4f0 · inbound
Lifelong Safety Alignment for Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5e259d3-2194-40a6-bdbb-67cc08680aa4 · inbound
MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capability DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e79e282-ae10-462a-b77d-59ed721a99a5 · inbound
Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5f5ff59-2f3b-4638-b652-9df06b60a614 · inbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfbbdad7-3e5b-4f76-a2f8-96a7f26b2e10 · inbound
Reinforcing General Reasoning without Verifiers DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89a3db25-1141-493b-a92f-89c2556b3b23 · inbound
TabReason: A Reinforcement Learning-Enhanced Reasoning LLM for Explainable Tabular Data Prediction DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ed4233d-21bb-4557-800f-f04794c7317e · inbound
Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dd26810-09e8-4b80-a6fc-02256ad6fec6 · inbound
Skywork Open Reasoner 1 Technical Report DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f8a98493-4dc1-47da-a516-9671f54832fd · inbound
Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d53205b8-a3e2-48cf-ad19-68ecbb65810e · inbound
RAG-Zeval: Towards Robust and Interpretable Evaluation on RAG Responses through End-to-End Rule-Guided Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6460a496-1546-47db-89c7-7bfca8d0211a · inbound
EvolveSearch: An Iterative Self-Evolving Search Agent DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 820e3921-aa3e-4902-b365-a3e0e076264f · inbound
WebDancer: Towards Autonomous Information Seeking Agency DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa45b29e-9039-478e-b74b-7f210b2826c2 · inbound
Accelerating RLHF Training with Reward Variance Increase DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d43a438e-6464-4aa2-9aca-a28846594327 · inbound
Are Reasoning Models More Prone to Hallucination? DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 681ebc7e-91ea-4c3c-b865-518350f00f3b · inbound
Grounded Reinforcement Learning for Visual Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2fa84e31-7281-42fd-bf91-ebeed07c9b66 · inbound
ZeroGUI: Automating Online GUI Learning at Zero Human Cost DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c2a96cf-38b2-41b9-9ea6-59fc39a26ca3 · inbound
Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d8616b4-cc9f-478c-83fe-fa1433f61ea9 · inbound
MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e61d0b-a2c8-4729-9c60-20162a21aa35 · inbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 97176b73-2444-4d0e-8d7e-1cff15f52cb7 · inbound
Towards Effective Code-Integrated Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 605d6a0f-b46e-45a6-b1fd-25588ab8c0ca · inbound
Reinforcing Video Reasoning with Focused Thinking DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9d0d6fc-7e1c-441a-8cf2-ed4ba973b7c1 · inbound
ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3538b41-65a8-40cd-9408-f3867b7e9582 · inbound
RLAE: Reinforcement Learning-Assisted Ensemble for LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9af23f9f-4d9b-47fb-8a1b-8f4bbf3eea20 · inbound
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2a6f081-263a-452c-8fcf-e02f9268efb1 · inbound
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea6bd0d-5eff-42cd-8422-36b9b3d4ecad · inbound
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ada27f0-ae0a-47d7-ac19-3c065fd19407 · inbound
Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4d6ed220-c3b6-406f-a283-3abb183608da · inbound
SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 209e468b-62bb-4f5a-b75b-ea1eb0702bd0 · inbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24986511-57a8-42cf-965a-fad7904e2c78 · inbound
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 378849d3-2eeb-4410-9d09-7f42e5337343 · inbound
Improving LLM-Generated Code Quality with GRPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2ab2182-d128-4808-a244-e300fc9dbb07 · inbound
Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d3da582-05a5-4454-8db3-fe030f03c885 · inbound
StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdf9bd70-ff73-41df-b817-ebc7834f7af2 · inbound
Seed-Coder: Let the Code Model Curate Data for Itself DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae02c8c2-824d-4188-926e-b4aa1678da7f · inbound
Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98430df0-679c-475b-be78-8a36753bf931 · inbound
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71e6ada2-e435-4c2c-ac6f-b27145cbf544 · inbound
QiMeng: Fully Automated Hardware and Software Design for Processor Chip DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5038f393-df94-4407-8ad9-d8331a54f79a · inbound
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 381b6018-32b7-4864-a7f0-3657153c080d · inbound
CodeContests+: High-Quality Test Case Generation for Competitive Programming DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aa1046d-820a-47b7-93e7-aa2b8b71956e · inbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6322f138-1f81-47ad-8802-2bdb1a21490b · inbound
How Far Are We from Optimal Reasoning Efficiency? DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32bd10b1-7c45-47e1-9654-29027edc5667 · inbound
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0505c5db-18fa-4fd8-bd05-48cf231bc2a9 · inbound
MiniCPM4: Ultra-Efficient LLMs on End Devices DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1605fb19-019b-4012-bcce-2f1da7f47d0e · inbound
SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99f5089e-12c0-4f1e-b071-e2dbd9917abd · inbound
Reinforcement Pre-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d56dc75e-cda6-408c-84b6-5867ff745e4d · inbound
A Survey on Large Language Models for Mathematical Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56f7a275-2495-411c-a8cd-15faa612de94 · inbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42b9d5fd-007a-4d0f-8532-9d229f7347a2 · inbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b61a26d-890a-4211-83ab-ee5887694417 · inbound
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfc66114-dd6a-40db-8a4e-b2fb90b34213 · inbound
RePO: Replay-Enhanced Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2908b548-4c0e-41ef-a3f6-4d729ffb6861 · inbound
CoRT: Code-integrated Reasoning within Thinking DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02b7edf3-4df0-4b51-9275-cd2534582e1c · inbound
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df998082-8d39-4586-bd26-e31f6501b6fd · inbound
Magistral DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 907799d7-17df-4531-9005-ab4ee4715382 · inbound
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30ab0835-abef-46be-821d-04b704d77905 · inbound
Schema-R1: A reasoning training approach for schema linking in Text-to-SQL Task DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fe48f52-8d98-45e2-ae22-e99dab76d1b2 · inbound
Continual Learning for Generative AI: From LLMs to MLLMs and Beyond DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 232
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b93c90ad-17e2-412f-aad4-b3c5f5b4155f · inbound
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c92f8f91-3a7a-4a5d-94f8-4032ea539520 · inbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.