Pith. sign in

Paper Citation Record · LEDGER

What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 44 inbound Pith citation observations for arXiv:2503.01491.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.01491 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:33:19.906869Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 753f2d48-65e1-43e0-b6fa-b36884c6ae38 · inbound

DAPO: An Open-Source LLM Reinforcement Learning System at Scale cites this paper.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:35:13.471029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:bfab6a8b9003ed5a2a59f61227b59f6a71780f2ea3829d0036e9b394cc81bcfc

Observation 5c60ecf9-6bdd-416f-ad2c-0f8b063fc3fe · inbound

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks cites this paper.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.792060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:cf5ea4c7aeeffb852e9884536d43bc291bc4a2bbb309cedcc56e5b7a028dbaaf

Observation b87ec276-4947-4f84-ab2a-7014f9912e0b · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.129490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:0d1bb4186583b14086392c361c2ce03067ad3130481a82dbed30a574d354a3bb

Observation e4056bbd-5d2a-4514-81dd-aafef7c0700f · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 149

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:19.906869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:19.906869Z digest=sha256:055a5ff64c0ab1b9e90223a6a897e1eafdab557a20378eb9979fe1f1f2b1f0f7

Observation fac2e47b-4999-44bc-8193-6a25d9126a7b · inbound

Reinforcement Learning for Reasoning in Large Language Models with One Training Example cites this paper.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:04.900537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:903173bb7e37bf3dfa8f211b326f34afba4fa778f9454e1bcda39812d17748c2

Observation bc970383-16f6-4198-8620-b50c5f323077 · inbound

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models cites this paper.

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 137

Resolution
unresolved
no resolver link, observed 2026-08-16T05:12:18.680956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:12:18.680956Z digest=sha256:4aa81830ed1708ffe3d4d24d0d5cfa565d62e78022f70dc02f25d2d1cff65725

Observation 8a248983-7160-41f9-bc5d-cbb54a625ba1 · inbound

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models cites this paper.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.692497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.692497Z digest=sha256:a9e0b4687de57fbde1ce945c2d9d8ed5df098084456f32ed7d12ef9d127f556f

Observation cdbef30b-f367-4bef-a478-724a06886ebe · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 168

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.841675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:4ab5833c8cb946d9c9ade8780d6754e1d980a23372a2b081e7dc847fd93aa630

Observation 30caf4b3-60a0-444d-8f18-cbce39e87a96 · inbound

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning cites this paper.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:59.582127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:59.582127Z digest=sha256:17f2a15531dd5aac631aec0dad66ae303d63edbcfe183891e79a9994b0a3eb7d

Observation 8a89e05c-8fec-4938-b143-5e4d83dab744 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.361455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.361455Z digest=sha256:2b40ba8518ebc647fd8a916283fc794cbdf2c5613ab929209b265bc1a01bffff

Observation 7bac26cf-a7c9-457e-b4af-6384c8746ff6 · inbound

Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles cites this paper.

Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:08.843880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:08.843880Z digest=sha256:75fc5e572cfecb85b973bb1b6c41cfa3898903a72a0d0008058249b031d537af

Observation 41185ef1-a4a3-4ab5-8244-c3389629e608 · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:38.437982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:38.437982Z digest=sha256:c73f8330c8dde72b1bf212d311e8c7b3135c8baf961e1c79c5359fc6d7e304e2

Observation 0359e4e7-fb51-4003-a55b-23be9a60d1b9 · inbound

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling cites this paper.

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:40:46.430743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-21T23:39:39.018498Z digest=sha256:5c44cc536d8224db35cc308fe5d723563d8406cb1ffa43eb72ab0c59cbdd7ee8

Observation d3e2fe4e-2353-4b39-b05f-cf96aee9919c · inbound

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization cites this paper.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.604562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.604562Z digest=sha256:d3d6c24d27ff3484e77c2e98a434ad0f248e69914611e49c4684acbd5ccc2efd

Observation 65780489-24fc-4086-8b40-c06b37a4b90a · inbound

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning cites this paper.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:58.979808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:f9b3a0e0499194a84353d58c628a77ac5a4c2bbd3d48b0894847c216a4c214d9

Observation 5c9667f0-812e-4e67-8d41-0e8f7003c117 · inbound

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought cites this paper.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.938067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.938067Z digest=sha256:2f9c42f670f1c27c4ea0d8244ced08daf750aa62ed92520e70394a637b5d2fd6

Observation 9a039030-0780-4570-af85-a240679e1ace · inbound

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework cites this paper.

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T05:45:04.288900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:45:04.288900Z digest=sha256:274289a92681ea144822bdf5a8eab58b887b2f45a851e2a0e65ad25c45d1670a

Observation 0fa64f27-2935-4dae-ba5a-1f772792bbb3 · inbound

ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards cites this paper.

ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:21:18.679910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T11:20:02.118579Z digest=sha256:b0173c2f684a98cd0307c2c6aa65d10b67e5a3ff3b972b62bb640b55e35cfa54

Observation ac9027f1-458d-4505-a4d9-19450ba8c1d2 · inbound

The Art of Scaling Reinforcement Learning Compute for LLMs cites this paper.

The Art of Scaling Reinforcement Learning Compute for LLMs What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:29:14.050525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:b90879c181e257357a0ba7b1752f874a7784dc7959389092bce635fd15a6f500

Observation 04809f1f-fea8-4b58-a2c3-8888db5b6d5a · inbound

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation cites this paper.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:15:26.747000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:65d2a44fb58e79d7e8ca3539efd090452ba09f2cdaaa06e478fa550d46dc5c7d

Observation 20b51108-69a6-4faf-9107-233b10865420 · inbound

Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention cites this paper.

Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T17:15:17.712961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:15:17.712961Z digest=sha256:41821ee23b4ddfded81ba471f9776bec1ecd6f8e57179f8bf47a9f9925950d1c

Observation 31ef4ebc-f0b0-47b0-985e-68fa4d6f064b · inbound

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models cites this paper.

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.030895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T14:09:26.842696Z digest=sha256:5df8f4f116b734d34efddf5b2682f2a4d53ec7e7ad2074610effdbc0634d6837

Observation 7951a39a-8e29-4105-9e88-6df39b8ef26c · inbound

Stabilizing Policy Optimization via Logits Convexity cites this paper.

Stabilizing Policy Optimization via Logits Convexity What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:08.739032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:08.739032Z digest=sha256:9eb990318af4b1a05d8c0918dcfbc310cbfa3cd54bd1565a6b94507c0950d2e9

Observation 66db204c-ed65-47a9-92ef-c982aa62aa79 · inbound

Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning cites this paper.

Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.345214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:09:29.657563Z digest=sha256:cd4b3a08b26151a22ebc461da82593152f705cc7b630c708339c56789f170beb

Observation 3bb9ead0-a0e3-4ebe-9f8d-7d3751944c49 · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:01:08.079209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:12:06.985609Z digest=sha256:b451d66f1a2b6a2ee10b9f761b422e261d720dd6b1cabd8706221d492ee59b07

Observation 94558f63-35ae-47fd-8d26-779ff6493a39 · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:54.566752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T01:04:12.454268Z digest=sha256:dd2fdc77d84887b1e36fecf21673b0181c1c534e9e7a57ecde0c24e39aa7b806

Observation bc66f473-1bf1-4c82-93a1-2f496ac34c44 · inbound

Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization cites this paper.

Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:56:01.737351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T15:45:15.962965Z digest=sha256:dc799edb14902df4dda3a5a9c871a66fa1198fa39baa9ef6091d0fa7f567e998

Observation f2228a7e-ff07-46f9-9d1f-8490b81e1d2f · inbound

Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization cites this paper.

Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T20:58:19.104482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:58:19.104482Z digest=sha256:2f819bd87695fb0624f6b5bcdf8d28b9649028a6bdd933d6d6f53b6eb72166ae

Observation 2fe4635e-d5d1-4b5f-99aa-f632a0fb50c4 · inbound

Segment-Aligned Policy Optimization for Multi-Modal Reasoning cites this paper.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:51:07.895095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:8496fc7b3cb5d5fd304c9e8d88c7edcdaedd509125a5fff41fe899920af619fc

Observation e12b29cd-60a4-4113-aa62-cab3f6236696 · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:55.031255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:5468322d2aab2c39e9628428966f02d73ef3c8514ffa80bdcdc9ed03316b4c32

Observation b3fcf1df-66e4-4b00-a07b-b03b4210df44 · inbound

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs cites this paper.

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:46:18.933292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T02:44:33.143247Z digest=sha256:aaaeda3aa5d1d98c8f244b4fe6889566e89747cbc82a3d8526f038c72418adb9

Observation 26dab740-55b7-4829-91d6-748d3a6f615c · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:12:22.747759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T06:08:28.855671Z digest=sha256:1aac15be1c91c2efab314e14490e44f84c712fc2f744f883fb0380daa02665a6

Observation 1394ac26-6e52-457a-a261-5452c7f99f99 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T10:01:22.994765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:a1e36d19b72c4ce8e59d0ed22d6f8668106bb95942f75293ec0625beed9643bc

Observation b9cbef82-41a3-4267-abd8-159305b64d18 · inbound

Extreme Region Policy Distillation cites this paper.

Extreme Region Policy Distillation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:24:02.202182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T23:14:56.223606Z digest=sha256:a517921cc37c27b7e0d86c954215a3179ae0b0792aef60b16953d8e0d494b4ad

Observation 3e59a4e0-5516-4789-bd87-45ea1d5982c1 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 207

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.277692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:d396bdd4c934e644a8ab9a9dcc35ad0369ddc32241fbac7ffd063ab08db96ee9

Observation efc2402c-5688-441b-b8ca-debbd664132b · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:13.368996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:33377941e6bb2105bbafd9ebd27fb09d1ed391f8e0436d8963bf106aa3c8088d

Observation bf00ac9e-5b34-4693-8441-de58a793eace · inbound

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling cites this paper.

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:56:30.041691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T10:25:10.559953Z digest=sha256:560e0b70aa1a456e173f6a9371ec6464912e12e4ba4a037e6751709ef4fee04a

Observation 04f6656b-bf9f-42e5-9081-5c27fb4c390a · inbound

VIMPO: Value-Implicit Policy Optimization for LLMs cites this paper.

VIMPO: Value-Implicit Policy Optimization for LLMs What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:30.690053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T17:59:29.066879Z digest=sha256:25b10c4a29964de8f15cba7f9a6a14d35b60d75927e138ac8a37fec21ddc4333

Observation 906b2229-b53a-49da-92d2-0d6d5f83132c · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 254

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.705178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:da6db3173333618691f8c85e255677fea3f8834c20703ba724aa2705fecbf5e6

Observation 6bbf5e80-5a18-4823-a0c3-2e3b10152ac7 · inbound

CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents cites this paper.

CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-07T14:03:48.691291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-07T13:57:33.821970Z digest=sha256:f198cfcc0db1706747f4623935cc8db21b60a2907904fa20195670872b4730b3

Observation b0504468-7755-4dd9-a45f-4f2ae2ccb475 · inbound

CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents cites this paper.

CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T14:03:48.463323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-07T13:57:33.821970Z digest=sha256:0f232e6768bf5e98d620bf9e210eb3cffc03f22d33c1a80e45644c014a75657d

Observation eef9deb5-9289-4a84-8567-56f20e255c95 · inbound

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models cites this paper.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.706606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.706606Z digest=sha256:1ef8b24591eea003278f4ff3a0c69a13b3777911e7f30731a760973f811b5b65

Observation 47b2cca1-ef7f-429d-a4b9-b4f81fd44a93 · inbound

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning cites this paper.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.895805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.895805Z digest=sha256:280cd3a049b2dcc3eadf1c026b23361e29b83f0cffc1252a40dca8c235f41445

Observation af1ee5bb-ac15-47a5-a867-87075d555589 · inbound

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning cites this paper.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.992560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.992560Z digest=sha256:60e19907b4bf1bcfeca6414ac19320acb8379740ed0ee48181f8b17095b90761