Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T11:43:29.701438Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 17 inbound Pith citation observations for arXiv:2509.02333.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T11:43:29.701438Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T07:57:26.363737Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
15 of 15 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 612402b8-1257-4662-aa19-493ee8a8300c · outbound
DCPO: Dynamic Clipping Policy Optimization Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06950766-b439-4f42-8812-ce89b7b1a72c · outbound
DCPO: Dynamic Clipping Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81639376-3c9f-4a41-a978-1420a4578330 · outbound
DCPO: Dynamic Clipping Policy Optimization HybridFlow: A Flexible and Efficient RLHF Framework
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4463601e-c8c2-427e-808f-e1b7ca51064d · outbound
DCPO: Dynamic Clipping Policy Optimization Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12433ec2-b841-44aa-aaf2-3a277c7b855c · outbound
DCPO: Dynamic Clipping Policy Optimization Qwen2.5 Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf81ad52-1075-4fd5-81a4-a154b849fd1d · outbound
DCPO: Dynamic Clipping Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 937f706c-1370-41df-be2a-827bf3bbdb27 · outbound
DCPO: Dynamic Clipping Policy Optimization Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6ea7d39-d2b7-4c29-9ad7-5f9cbcc86825 · outbound
DCPO: Dynamic Clipping Policy Optimization Group Sequence Policy Optimization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fafd8e02-71a0-49ac-b601-3dcc6dbe5b31 · outbound
DCPO: Dynamic Clipping Policy Optimization θParameters of the actor model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f2b6d977-e482-4e16-8d7d-b47931e50ad9 · outbound
DCPO: Dynamic Clipping Policy Optimization Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f2ff9d1d-6452-4efb-be0b-1dcbb29f38a6 · outbound
DCPO: Dynamic Clipping Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9734dab3-faec-4543-8bea-8ddb9313634f · outbound
DCPO: Dynamic Clipping Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ae4ce3d-0029-4c60-8867-dc1e9a2f3892 · outbound
DCPO: Dynamic Clipping Policy Optimization Proximal Policy Optimization Algorithms
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ade20952-dfcf-4917-a25d-374db064673f · outbound
DCPO: Dynamic Clipping Policy Optimization DeepSeek-V3 Technical Report
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a7e4936-c230-4a9a-9461-d3d8e254ab2b · outbound
DCPO: Dynamic Clipping Policy Optimization Measuring Mathematical Problem Solving With the MATH Dataset
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 959b34f9-ed7b-458b-b93e-ebb4e6625e74 · inbound
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation DCPO: Dynamic Clipping Policy Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 619592de-df22-4efa-9f53-ca814126aaed · inbound
SSPO: Subsentence-level Policy Optimization DCPO: Dynamic Clipping Policy Optimization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 77a4282b-d569-4f04-92b1-ad826df520c1 · inbound
Chart Specification: Structural Representations for Incentivizing VLM Reasoning in Chart-to-Code Generation DCPO: Dynamic Clipping Policy Optimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e98b6496-b2d0-4e43-ad84-d51363ff9d06 · inbound
Policy Improvement Reinforcement Learning DCPO: Dynamic Clipping Policy Optimization
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e22525b4-c3cc-40b8-8ddd-17e9aa3afd86 · inbound
Design Conditions for Intra-Group Learning of Sequence-Level Rewards: Token Gradient Cancellation DCPO: Dynamic Clipping Policy Optimization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation aaf16670-f46d-4f31-8aea-15efdc23cb14 · inbound
Design Conditions for Intra-Group Learning of Sequence-Level Rewards: Token Gradient Cancellation DCPO: Dynamic Clipping Policy Optimization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e50610b-be2f-4a28-85eb-833774fab59a · inbound
MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models DCPO: Dynamic Clipping Policy Optimization
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 71fbcbc2-a948-4e15-a760-343c73547c9e · inbound
Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance DCPO: Dynamic Clipping Policy Optimization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f63ea4b3-fee7-4416-9b64-cae9c7f20040 · inbound
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization DCPO: Dynamic Clipping Policy Optimization
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a6d9ae2b-fa2d-4991-95a4-32a88213f45b · inbound
Revisiting DAgger in the Era of LLM-Agents DCPO: Dynamic Clipping Policy Optimization
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1907e884-6752-419b-ad78-3e2e61c988e7 · inbound
Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation DCPO: Dynamic Clipping Policy Optimization
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 10a2d3fe-40a5-4d0e-a7dc-0966a032f248 · inbound
Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation DCPO: Dynamic Clipping Policy Optimization
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ac7a7c67-77f7-46e8-9875-4e3aaeeb8c81 · inbound
Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs DCPO: Dynamic Clipping Policy Optimization
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b91c7737-5557-43fa-bb00-d7f3f9c38e9a · inbound
Baichuan-M4: A Clinical-Grade Medical Agent System for Continuous Care DCPO: Dynamic Clipping Policy Optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 46c84091-9a12-405a-a5fa-59c35417e406 · inbound
GUI-AC: Enhancing Continual Learning in GUI Agents DCPO: Dynamic Clipping Policy Optimization
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation acc6b601-86f6-4862-9b33-f954d69f85b7 · inbound
GUI-AC: Enhancing Continual Learning in GUI Agents DCPO: Dynamic Clipping Policy Optimization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11ddaa8c-629e-4672-85dd-f8a103c228e5 · inbound
Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning DCPO: Dynamic Clipping Policy Optimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.