Pith. sign in

Paper Citation Record · LEDGER

DCPO: Dynamic Clipping Policy Optimization

As of 18 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 17 inbound Pith citation observations for arXiv:2509.02333.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.02333 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:43:29.701438Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:57:26.363737Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 612402b8-1257-4662-aa19-493ee8a8300c · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

DCPO: Dynamic Clipping Policy Optimization Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T11:43:28.500127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:43:28.500127Z digest=sha256:07c3b4065070872d2f12b678c33c6353aeafa5d3dc36c486cee2f49290c2a2d7

Observation 06950766-b439-4f42-8812-ce89b7b1a72c · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

DCPO: Dynamic Clipping Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T11:43:28.762835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:43:28.762835Z digest=sha256:a5eea81aa56fa07e03c26d818d17eae8c3972616a53855b1444a145528059465

Observation 81639376-3c9f-4a41-a978-1420a4578330 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

DCPO: Dynamic Clipping Policy Optimization HybridFlow: A Flexible and Efficient RLHF Framework

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T11:43:28.966325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:43:28.966325Z digest=sha256:f3d3727f07017b3d4a1c3396e977dd5e8564a676d58f852b0f5bfb0c65c0c923

Observation 4463601e-c8c2-427e-808f-e1b7ca51064d · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

DCPO: Dynamic Clipping Policy Optimization Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T11:43:29.052172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:43:29.052172Z digest=sha256:7b1a2573223fc899e63feacb0fa934d67a4e7c885d8414d0a5376dd9e51a1932

Observation 12433ec2-b841-44aa-aaf2-3a277c7b855c · outbound

This paper cites Qwen2.5 Technical Report.

DCPO: Dynamic Clipping Policy Optimization Qwen2.5 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T11:43:29.133703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:43:29.133703Z digest=sha256:cb7caa01fc64b323d279c9040077df595e9fbd2067cb51440330a35ff9229f55

Observation cf81ad52-1075-4fd5-81a4-a154b849fd1d · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

DCPO: Dynamic Clipping Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T11:43:29.342020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:43:29.342020Z digest=sha256:eb93dc3d7bdfc519ff5c6223918ded3dfe293b1f977d45280124cbf024d2bc1c

Observation 937f706c-1370-41df-be2a-827bf3bbdb27 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

DCPO: Dynamic Clipping Policy Optimization Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T11:43:29.437153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:43:29.437153Z digest=sha256:55df462fd92a125c7d3582908514eb448afcd29b30de6e58b873adff337ec077

Observation d6ea7d39-d2b7-4c29-9ad7-5f9cbcc86825 · outbound

This paper cites Group Sequence Policy Optimization.

DCPO: Dynamic Clipping Policy Optimization Group Sequence Policy Optimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T11:43:29.538938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:43:29.538938Z digest=sha256:502cf8c84b7d48e228103587bdacda6cc32072e1cc353b48e215409776481266

Observation fafd8e02-71a0-49ac-b601-3dcc6dbe5b31 · outbound

This paper cites θParameters of the actor model.

DCPO: Dynamic Clipping Policy Optimization θParameters of the actor model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:43:30.684866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T11:43:29.603547Z digest=sha256:20e6182699546cc11e08f68a3af56b3898e48f5e0be39b019997054e499ba502

Observation f2b6d977-e482-4e16-8d7d-b47931e50ad9 · outbound

This paper cites an unresolved cited work.

DCPO: Dynamic Clipping Policy Optimization Unresolved cited work

Reference 15

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T11:43:30.468680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T11:43:29.701438Z digest=sha256:f3faafea9ae2fc57a17f11d159d661bf2bbf9c731cc8d14f29810adb318bf711

Observation f2ff9d1d-6452-4efb-be0b-1dcbb29f38a6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DCPO: Dynamic Clipping Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T11:43:28.901811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:43:28.901811Z digest=sha256:02f971cee089f53feb38c17ccad5c55f6461fab64a12db9966bc8feeeb729aa2

Observation 9734dab3-faec-4543-8bea-8ddb9313634f · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

DCPO: Dynamic Clipping Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T11:43:29.221424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:43:29.221424Z digest=sha256:efacbf6fd37c3084e84ff0aa88fee3ec0aa313e1b2d808c874caf84ccf6634ba

Observation 5ae4ce3d-0029-4c60-8867-dc1e9a2f3892 · outbound

This paper cites Proximal Policy Optimization Algorithms.

DCPO: Dynamic Clipping Policy Optimization Proximal Policy Optimization Algorithms

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T11:43:28.827284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:43:28.827284Z digest=sha256:cc8ce7f13e4216b356ffae3f7c13b778c7dafb4e7d5a2de3116326db728bffb9

Observation ade20952-dfcf-4917-a25d-374db064673f · outbound

This paper cites DeepSeek-V3 Technical Report.

DCPO: Dynamic Clipping Policy Optimization DeepSeek-V3 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T11:43:28.584291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:43:28.584291Z digest=sha256:f02cb1554f5bba7637c9b86f6b1c6d286e723e93afe768e69e24b1503d1203e7

Observation 1a7e4936-c230-4a9a-9461-d3d8e254ab2b · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

DCPO: Dynamic Clipping Policy Optimization Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T11:43:28.710266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:43:28.710266Z digest=sha256:14ea4ad2a93e5f8ba8ef97af3bc1dc2bc51598ad56f9152f8f2eeab87f085aec

Pith citing papers

Observation 959b34f9-ed7b-458b-b93e-ebb4e6625e74 · inbound

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation cites this paper.

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation DCPO: Dynamic Clipping Policy Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T07:57:26.363737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:57:26.363737Z digest=sha256:af459538aa74a871733802736a59783eda00208a233e0a8125ebfd087632a808

Observation 619592de-df22-4efa-9f53-ca814126aaed · inbound

SSPO: Subsentence-level Policy Optimization cites this paper.

SSPO: Subsentence-level Policy Optimization DCPO: Dynamic Clipping Policy Optimization

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:20:34.252024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T01:19:45.247967Z digest=sha256:f0c9a448b520a0575a41f34107ea81274cc1f10f477ac0bfebed36a9e2a10583

Observation 77a4282b-d569-4f04-92b1-ad826df520c1 · inbound

Chart Specification: Structural Representations for Incentivizing VLM Reasoning in Chart-to-Code Generation cites this paper.

Chart Specification: Structural Representations for Incentivizing VLM Reasoning in Chart-to-Code Generation DCPO: Dynamic Clipping Policy Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T06:07:01.461595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:07:01.461595Z digest=sha256:8042538656e90d76078e47c1ec6dd98dc3a433e629c41221c66aefd9bff47773

Observation e98b6496-b2d0-4e43-ad84-d51363ff9d06 · inbound

Policy Improvement Reinforcement Learning cites this paper.

Policy Improvement Reinforcement Learning DCPO: Dynamic Clipping Policy Optimization

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:48:22.789866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T22:47:05.132020Z digest=sha256:6660ab1837b0fe2cf84312e8bec55bed92a14650ddda69bce2232624834bf050

Observation e22525b4-c3cc-40b8-8ddd-17e9aa3afd86 · inbound

Design Conditions for Intra-Group Learning of Sequence-Level Rewards: Token Gradient Cancellation cites this paper.

Design Conditions for Intra-Group Learning of Sequence-Level Rewards: Token Gradient Cancellation DCPO: Dynamic Clipping Policy Optimization

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:43:07.466439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T18:41:26.965577Z digest=sha256:02696bfc2e22148e087e64e28025ffecad1c7aad36615d4a0ceb456327b29534

Observation aaf16670-f46d-4f31-8aea-15efdc23cb14 · inbound

Design Conditions for Intra-Group Learning of Sequence-Level Rewards: Token Gradient Cancellation cites this paper.

Design Conditions for Intra-Group Learning of Sequence-Level Rewards: Token Gradient Cancellation DCPO: Dynamic Clipping Policy Optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T13:07:31.208017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:07:31.208017Z digest=sha256:798b517dd68bdfb6995024e735ebee84341ddb2b9fb445b633529dfc3a2fca52

Observation 0e50610b-be2f-4a28-85eb-833774fab59a · inbound

MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models cites this paper.

MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models DCPO: Dynamic Clipping Policy Optimization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:06:52.794873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T07:05:45.081955Z digest=sha256:6b864b75e33a8c0111de037c651f81e9aa474ad88ed99c021d0663636fc705b7

Observation 71fbcbc2-a948-4e15-a760-343c73547c9e · inbound

Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance cites this paper.

Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance DCPO: Dynamic Clipping Policy Optimization

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:12.993056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T08:05:36.206946Z digest=sha256:42dcc986f5cf2d23a6d6cf8bd1e160f6c36ba47a44765bdf0f2cd002bfdc2818

Observation f63ea4b3-fee7-4416-9b64-cae9c7f20040 · inbound

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization cites this paper.

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization DCPO: Dynamic Clipping Policy Optimization

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:42:02.915611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T01:41:55.435903Z digest=sha256:188af7c76b37b352202a65d76b627af2b0e76a9096cf5d5f4871daba97a66a82

Observation a6d9ae2b-fa2d-4991-95a4-32a88213f45b · inbound

Revisiting DAgger in the Era of LLM-Agents cites this paper.

Revisiting DAgger in the Era of LLM-Agents DCPO: Dynamic Clipping Policy Optimization

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:57:53.349495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T19:56:06.762156Z digest=sha256:1e2c0140f6eddd76bac9b23e3917bd5c707800286351a796dcda34a48dfea369

Observation 1907e884-6752-419b-ad78-3e2e61c988e7 · inbound

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation cites this paper.

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation DCPO: Dynamic Clipping Policy Optimization

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:09:41.324572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-21T06:08:55.571828Z digest=sha256:5fb58bc4ca06461c8510c1941487511ba10f7ed0dff30198b1c4c4744df3c71e

Observation 10a2d3fe-40a5-4d0e-a7dc-0966a032f248 · inbound

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation cites this paper.

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation DCPO: Dynamic Clipping Policy Optimization

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:15:47.324889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-30T17:15:53.286925Z digest=sha256:bac639ada2abe0a02acdae74fa7ad882f1671375f2eaaf99ffd0e74e88f3ee21

Observation ac7a7c67-77f7-46e8-9875-4e3aaeeb8c81 · inbound

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs cites this paper.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs DCPO: Dynamic Clipping Policy Optimization

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:13:26.575383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:e0f4418530566c40a6bf1d243384e059e6fb305cea238dc2cc82b0832bec913d

Observation b91c7737-5557-43fa-bb00-d7f3f9c38e9a · inbound

Baichuan-M4: A Clinical-Grade Medical Agent System for Continuous Care cites this paper.

Baichuan-M4: A Clinical-Grade Medical Agent System for Continuous Care DCPO: Dynamic Clipping Policy Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:01:07.442871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T16:58:54.293859Z digest=sha256:08f20c95164b2d5c6eeb918c7fc8ee0db91b2459ca83162dff6ca928f6fb43b2

Observation 46c84091-9a12-405a-a5fa-59c35417e406 · inbound

GUI-AC: Enhancing Continual Learning in GUI Agents cites this paper.

GUI-AC: Enhancing Continual Learning in GUI Agents DCPO: Dynamic Clipping Policy Optimization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:27:36.576093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T13:56:09.049753Z digest=sha256:0d0096caa02d48f322934b967ece4b386c20d64f021d3ea1adcc8e9c5f04875e

Observation acc6b601-86f6-4862-9b33-f954d69f85b7 · inbound

GUI-AC: Enhancing Continual Learning in GUI Agents cites this paper.

GUI-AC: Enhancing Continual Learning in GUI Agents DCPO: Dynamic Clipping Policy Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T14:27:05.589465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:27:05.589465Z digest=sha256:8f436f09ae120733cea8c5abca61bdf87c0544cdcc82b04db3da3ff49bed1301

Observation 11ddaa8c-629e-4672-85dd-f8a103c228e5 · inbound

Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning cites this paper.

Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning DCPO: Dynamic Clipping Policy Optimization

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-09T18:26:26.298788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-09T18:19:12.180499Z digest=sha256:bc2c81149e2cac476a8f0fa8afd9a9c7403cdc33c58f462dafdca33bcf940ee8