Pith. sign in

Paper Citation Record · LEDGER

VinePPO: Refining Credit Assignment in RL Training of LLMs

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 67 inbound Pith citation observations for arXiv:2410.01679.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.01679 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 67 of 67 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:07:08.079012Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T11:16:11.296273Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d46ce63e-8996-4551-81a3-11bee3ec1ad3 · inbound

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning cites this paper.

Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:42:19.201696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T01:42:19.004468Z digest=sha256:6a3ffc842708a5457f53bb1a94c982f7f987682a8ffb5059d372092ba3e7117e

Observation 09915755-9dcb-4e1b-b0af-9d6ef12953fd · inbound

A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios cites this paper.

A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T21:58:46.573945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:58:46.573945Z digest=sha256:97aa39265476e393b487350bf600b8360af31c1a35495b15cea89458dea434e3

Observation 24be95d8-0663-42f6-9192-92e5cd642d3e · inbound

SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs cites this paper.

SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:57:44.301141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:57:44.301141Z digest=sha256:748fe3fd70725a7f5a678902274ec149dfcbae808f08902f6ef074cacf83c6fc

Observation d43bdd78-07f9-409b-8162-5646d56c8996 · inbound

AlphaPO: Reward Shape Matters for LLM Alignment cites this paper.

AlphaPO: Reward Shape Matters for LLM Alignment VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:51:08.144783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:51:08.144783Z digest=sha256:d7ce3d01fbd33fde1ba6377baed83db2ab1617d6e8c4f91b8570e08c21c90dfe

Observation 2c82506e-c364-4e71-897b-b8ac5ab0a83f · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T21:20:59.221163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:30450b9c05da1d2c7b96f4a725d5d5460c898ecf168aa82c92f335065abf189a

Observation 7264ae2e-8654-42d1-9a4e-1d75ffc76cea · inbound

T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling cites this paper.

T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T18:05:43.046044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:05:43.046044Z digest=sha256:e7ebf1e7c446bff12218ce35574b566b6f314f13331f30081f9ac9f1cacfd792

Observation 660afd43-a456-4ade-93de-089f31c37342 · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:23:30.899521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:7835da63b91a3eb3f47692c80c30f9729df9f88def95d75c4d86e58bfbf49c7b

Observation d8060d98-83f4-4402-b46b-a9de15137134 · inbound

Reinforcement Learning for Long-Horizon Interactive LLM Agents cites this paper.

Reinforcement Learning for Long-Horizon Interactive LLM Agents VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T14:56:00.220169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:56:00.220169Z digest=sha256:2116542a8d9622e7bdf1f42d1c8aced53ef7c9c42c60cdabe15f2ae89204d7e3

Observation f8d070c7-893d-42df-86e7-3ed7693068de · inbound

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning cites this paper.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.010658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.010658Z digest=sha256:e71d11332726abdf154b688cfc07af2fd53ea4feb574edc049f349acad5c16fc

Observation 66fa54c1-421c-4938-8538-0c19b137cf03 · inbound

Learning to Reason at the Frontier of Learnability cites this paper.

Learning to Reason at the Frontier of Learnability VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:42:26.025966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:f54f90ed117b63b7d500dfc65bba502d4171652abc836eb18771249770cb657d

Observation 531ef94a-12a6-4835-b0ef-ade283579f4a · inbound

DAPO: An Open-Source LLM Reinforcement Learning System at Scale cites this paper.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T23:35:13.436228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:6c6a4c3cc9a5d0d79d51af6b000e62c88c6d17104edfb0a04da9e22346bf7ea3

Observation 9884ecf5-e565-4cd1-8f51-55f1320569f9 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 165

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T19:32:01.369168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:86d4487cd5fa65f8c59c6253fd7545f309e791cf8a6f2a04f780bcfde177c7c6

Observation 679cdec8-d62f-4978-b54c-89a068006e56 · inbound

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning cites this paper.

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:46:56.989883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T18:46:11.571831Z digest=sha256:60d40c295cd3eebf26f90ca4a30accefbee64ca0720053f9bdba3c6112b9e727

Observation 5c437fa1-e43b-4fe2-9ac4-1f9243aeb264 · inbound

Reinforcement Learning for Reasoning in Large Language Models with One Training Example cites this paper.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:51:04.893040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:4f8eb97f11484b201da54a07144694df6ce4551f991b787278ca0420e173088b

Observation e3dc08e0-b3c8-4c0a-8ac5-5e1252dd3826 · inbound

Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations cites this paper.

Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:08.079012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:08.079012Z digest=sha256:9dbdd7d0ace85a0fe4b98ee899e3ce096f94d4c5847d639c35e584034b23cafd

Observation 2fd4e5a9-2fff-427d-9069-f95664080fc9 · inbound

Group-in-Group Policy Optimization for LLM Agent Training cites this paper.

Group-in-Group Policy Optimization for LLM Agent Training VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:15:08.648122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T09:15:08.193357Z digest=sha256:1501741f4e3d207c659a1df19d4b8e00d98090d64775f103db96879df38fb47a

Observation 63b5de8a-0e45-4078-b03c-0c1fcbee1a0e · inbound

Effective Reinforcement Learning for Reasoning in Language Models cites this paper.

Effective Reinforcement Learning for Reasoning in Language Models VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:16.694006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:16.694006Z digest=sha256:e57cf4bb5e9dfc6b6ecb2aacec7f1508e17449644fac633e9ce6c1444e1293c7

Observation 5139bbd4-11cc-4431-be00-5e305202204e · inbound

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training cites this paper.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.663718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.663718Z digest=sha256:b2204d0e4f22968664abda5ee78ba920138a87c800b0c63e931ca9de5113a68f

Observation 743dc8c4-c296-4e3d-837d-dbfc0f4d43cc · inbound

How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning cites this paper.

How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:21.867757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:21.867757Z digest=sha256:747aa57314bb42425040c85fdbf2fc7967f4444fd604fe3a18e9b3461eb67e33

Observation aedd8de2-cdae-4679-b01c-fc73610148f5 · inbound

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts cites this paper.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:54.702600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:54.702600Z digest=sha256:e389d7249ab0325265f24ec2219d29e505a9cfbc78a6d8f376bbf24332ed93f4

Observation 229e0672-ac94-41c0-aad3-7d737231e735 · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.026404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.026404Z digest=sha256:a1241a6890a5e018a59abeae0964ff5de4c64638670f011fb774a418cf8a6e0d

Observation 3157642f-77c6-4d89-837c-c2ca43862907 · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:47.060414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:47.060414Z digest=sha256:9fa24f39d42ed515498e99fbecf86516bd9ae43f0ed6698fad0843683c651932

Observation ecaff3a5-cd06-4143-839f-ef6f18d77e92 · inbound

Intent Factored Generation: Unleashing the Diversity in Your Language Model cites this paper.

Intent Factored Generation: Unleashing the Diversity in Your Language Model VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:47:42.410649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:47:42.410649Z digest=sha256:50099703c889bca4c4827c3bdcbb6d2de8e7409ba65b16d2087d066107b152db

Observation ebe55ede-7678-4f2c-8c1e-c65490ac9874 · inbound

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search cites this paper.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:21.838257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:21.838257Z digest=sha256:bf76b3b2aaa35714b5f309c090e9a8f2599c151565c6e44d284b898d1318e026

Observation abef9125-2b0d-482e-9369-b428b999344d · inbound

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning cites this paper.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.793524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.793524Z digest=sha256:fab5baed5168b1440991bfd8456a95dc24502a9eab6a0ce7120563d4be2863de

Observation 7659bb44-44de-4f05-b2e4-a30d560d3f6d · inbound

AdapThink: Adaptive Thinking Preferences for Reasoning Language Model cites this paper.

AdapThink: Adaptive Thinking Preferences for Reasoning Language Model VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:37.524741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:37.524741Z digest=sha256:164e7d4da5c05148a1ace5d059755694b269bb154a01afd2fa1327892277f8ce

Observation 7a228f38-abdb-453c-813b-0609cfd6acc2 · inbound

Enhancing Large Language Models through Structured Reasoning cites this paper.

Enhancing Large Language Models through Structured Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:59:47.529281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:59:47.529281Z digest=sha256:1516a196d9ad9a0a1e7102a2a2d313ac472f67ddbddf449f9932433f4c54f031

Observation 4fcc9aa5-d6d3-420c-90c1-532a17c17fe1 · inbound

First Return, Entropy-Eliciting Explore cites this paper.

First Return, Entropy-Eliciting Explore VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:12.630333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:12.630333Z digest=sha256:9132a6b14acdcd477abbbc98db7d5be0aff3c67708d7b6f2d50013531ba684cd

Observation eed7fb36-b81f-4e42-afde-45576bac2ccc · inbound

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning cites this paper.

Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:49:03.216906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:49:03.216906Z digest=sha256:6fb1b4c67be6ed4774ff5997e3b9985b2369950a478e529e01a9cb104bbfaffd

Observation b07b515e-c171-45cd-8835-0d1c08964fb0 · inbound

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR cites this paper.

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T16:45:52.821880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:45:52.821880Z digest=sha256:f5dbaf4f65f46c488453d47116b124b55f7378067f44297cca56b74677146013

Observation 6179bf92-2916-48d4-8521-503ca9af2743 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T19:21:48.734746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:60d1d1b0198efa0670e17c683176dceb58ee601682c7cdf05e1875342eb8e3df

Observation 258df9e2-bacb-48fb-a5fc-371f2e2d24b5 · inbound

GPO: Learning from Critical Steps to Improve LLM Reasoning cites this paper.

GPO: Learning from Critical Steps to Improve LLM Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:55:56.199088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:55:56.199088Z digest=sha256:82a782f3c0c62358e44038da144ee5f1a39f5bc73783bdf89e4b3304ab6561ce

Observation 1fb923ae-a3f1-4e42-be94-8fce0d68c6c1 · inbound

Polychromic Objectives for Reinforcement Learning cites this paper.

Polychromic Objectives for Reinforcement Learning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:56:19.879623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:6fca9bfbe313d47f65ba084e221eed77e77b7a3406f27145f3c5688bc83bf399

Observation 054fc3b3-4d4f-442f-83f4-f9d546c5bdd7 · inbound

Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing cites this paper.

Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:27:36.792411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T08:24:50.793194Z digest=sha256:dd6443315d764670c5885ccc9f6736f54c5024fb447eebacc83079bc61375da2

Observation 859fe84f-2850-4502-9cc9-6efb93a0ed50 · inbound

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks cites this paper.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T04:17:27.143818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:17:27.143818Z digest=sha256:2c95763bb3c7063ee616dc46a2665505841d7e15a760672d36be219bf472f9bf

Observation 292aef98-a97d-4cc6-84ce-9f2450eb9a4d · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.791107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.791107Z digest=sha256:07c84734377f44324fff472cee4564eb94e388f3807367fef090342051d7f2bb

Observation fdafd0cb-8359-4589-80fa-a54d0b1c4e9f · inbound

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization cites this paper.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.456107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.456107Z digest=sha256:9bfecef690e7d14b0f0e5872dd9f02b247410c9bc3aef73c7232fc06d64e1bf6

Observation 5f1fdd7b-4fd3-4395-bb15-b8b36d29d35f · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:01:06.836055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:12:06.985609Z digest=sha256:dee25d70fcbd275cdabe9e42a735dd8370eed65c4b48c82c62f758944c87201d

Observation 3e47f047-efc4-40a2-a348-a2ad4ba492e5 · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:54.588405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T01:04:12.454268Z digest=sha256:56dfd1fa20b1c33b605a62d82abb02fe69f8b76c38fca4d21a494fb7f497010c

Observation 3e1195a2-9c35-4c76-8bb9-831c3d69e442 · inbound

GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning cites this paper.

GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:14:46.346210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T00:14:29.531510Z digest=sha256:661bde011bb41cd73e74855b013f49e861baa63f94ea04dc4378842a6076c04b

Observation dbe3be7e-f002-44ec-b718-ec4a4f173752 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:29.789199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:b292fd3d6b7f4cccf03a76b87aeb1d98a5d3972cd1c4b441f43c94284e845265

Observation 60df7934-c0f0-41a7-8f8c-79b03742e514 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:17.149688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:ff7c22737906746145970ddb4213cc8ac8236abf44a025aa9c893b8d5fdb1745

Observation c42d6553-ecef-4eb9-b9a6-0a35e43a5894 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:02:40.725434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:5a7773596a867bfa7822c53c190415516cb622ca87b82016ba9f00ca5a1a4ed1

Observation b3006c93-5a7d-400d-ac68-89784e893364 · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:00:55.014461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:c8cd7a996c03f582441115713c92e6b0d9e52cf80f5023784b50f205923b4eb2

Observation 4c6eab82-1882-4d0a-ba4e-38c3655da520 · inbound

Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy cites this paper.

Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:48:28.628790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T01:46:24.724553Z digest=sha256:99cb7bcd5e5385c32a5a207ab14c0750c35eb39fb8a3ab3b472a9d9ed8a7cd4b

Observation 285b56dd-05cd-480c-9a83-30f501a7bdb8 · inbound

Learning from Language Feedback via Variational Policy Distillation cites this paper.

Learning from Language Feedback via Variational Policy Distillation VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:39:00.324286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T20:34:36.764090Z digest=sha256:9a006a0492c59dc0d9a7eceb8408da1625e682425d46938451a9876379e3de91

Observation a054cb16-c7c7-4cb8-960d-33c1654713d8 · inbound

LamPO: A Lambda Style Policy Optimization for Reasoning Language Models cites this paper.

LamPO: A Lambda Style Policy Optimization for Reasoning Language Models VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:13:58.415659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T05:11:35.785227Z digest=sha256:ad7f4b22fdacf06789b95e7448c6b7f675aca2d7f097eaf0e3343fb0d3ee5449

Observation cdcf3e34-b79e-4657-b783-19a8cfee7570 · inbound

DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards cites this paper.

DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:29:40.147973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-21T05:24:46.545570Z digest=sha256:affc81e956c338c1355d8468d791a042e539c65d36b09c5942e8e72e87f17f1a

Observation 2000b893-ccfe-487d-8945-36f0bcf0c6b4 · inbound

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use cites this paper.

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.592676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T13:49:58.677299Z digest=sha256:5ed973f9f230f6c953bfa6b7259e35ccc2ede5116904685fd55b0fd2fbd49056

Observation 313d9a8a-2f6c-4ca2-bfb1-5f815fd7ac08 · inbound

ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate cites this paper.

ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:16:00.836667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T22:54:53.415067Z digest=sha256:dfa8309043f163e35f85efbe7a395b93b1a65362f9bac508de71d14daef51629

Observation c4c6f982-3795-4de8-ac5b-69e43e76e27d · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:16:13.394437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:7465088e5e4a9ca571059ce64d04768003ec84bd46dff539661c86e9132416b5

Observation f062efd6-2eb7-483f-ab81-f3370396598f · inbound

Self-Distilled Policy Gradient cites this paper.

Self-Distilled Policy Gradient VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:56:27.891368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T11:24:11.878439Z digest=sha256:c6962c1c0b7feeec4d9077857b6c3787030245d4bd7b58696288ffef649d9b3c

Observation 026f94e9-4d8e-4b7c-a18b-9b97c165eef9 · inbound

ProcessThinker: Enhancing Multi-modal Large Language Models Reasoning via Rollout-based Process Reward cites this paper.

ProcessThinker: Enhancing Multi-modal Large Language Models Reasoning via Rollout-based Process Reward VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:10:07.051662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-04T20:09:16.222780Z digest=sha256:94abe5628c0f043267c05fec11d9ec9fc8d72251dfc91a283cbc4930012baa04

Observation c80068eb-114e-436d-abe5-6da4aa2a9fea · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.569959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:c4bed106abc9baa79c92e9002bf648530fa1b51b933edcb87b041baef14d5e67

Observation 80f80eab-c7f2-4122-a141-22f610ac46ad · inbound

CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning cites this paper.

CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:54:22.302365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-30T07:48:29.333893Z digest=sha256:c3fa90f379aac3ea1d8e5b1f2d80ee8ce912bec4dd8eacbeaf0e52bb64755ee8

Observation 887dfdeb-cc6e-4321-a082-49098d3643e3 · inbound

ECHO: Learning Epistemically Adaptive Language Agents with Turn-Level Credit cites this paper.

ECHO: Learning Epistemically Adaptive Language Agents with Turn-Level Credit VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T16:44:56.587708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T04:36:35.213066Z digest=sha256:ad72167760a829d95661db55a48c07b79f0d65d8fa18af5e42fa171ae5d10f50

Observation 1f894e4c-bd00-49c7-b473-7880171e54e1 · inbound

Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG cites this paper.

Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T05:47:14.970921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:47:14.970921Z digest=sha256:e2ba410d6bd200bfbf4fe8b04ab44b60f10eeb681c89693e8f3044e4437e6b16

Observation 4e61ab1e-c568-4a6d-9cbf-62b4391d7674 · inbound

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy cites this paper.

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T04:43:45.592808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:43:45.592808Z digest=sha256:4485a47c2a143b23cbace184732b9a14b8e393e4c5f79bcc9c71426bd7e0d6f5

Observation f71504bb-3335-4a32-adbd-5d17756fe24e · inbound

Weak-to-Strong Generalization via Direct On-Policy Distillation cites this paper.

Weak-to-Strong Generalization via Direct On-Policy Distillation VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 105

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T12:33:45.165685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-07T12:31:42.224094Z digest=sha256:30627e75efb3913da911c40d0aebabd5dec4c8061861d09ce1d6174533d8a6ce

Observation e68e1632-e745-4550-a34c-fc003f8ded03 · inbound

Weak-to-Strong Generalization via Direct On-Policy Distillation cites this paper.

Weak-to-Strong Generalization via Direct On-Policy Distillation VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 101

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:bd0fae746acd5c95f678db4b4e02c195bd628bdb63b0b98bc030ba1ddbe019ee

Observation cb0895ee-fa12-415a-95cf-58c0b8c7ecdc · inbound

RLVP: Penalize the Path, Reward the Outcome cites this paper.

RLVP: Penalize the Path, Reward the Outcome VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T11:16:11.298001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:ee6f48a87df0f85c3aee0c01cda39f29dae0516cdabbaea5a8c91a277d62be7c

Observation 0152b37e-a63f-4c93-ba80-b264a7d4efdc · inbound

Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning cites this paper.

Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T04:40:38.083853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:40:38.083853Z digest=sha256:bd12b595609560758770eef13b574f3f70830310e9da150443aef5d2a0d11c34

Observation 3494809f-a207-4435-950f-a4ad60214071 · inbound

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning cites this paper.

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-31T01:31:10.823817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T01:31:10.823817Z digest=sha256:3fc1c451030cf3e1d5aae7ff7973e92b094e768ae22662b1a014d44a9a6c3924

Observation 61fc87b8-3b93-4468-85a1-b746303ea32b · inbound

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning cites this paper.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:11.667302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:11.667302Z digest=sha256:a60813ef341351834f6f8659fcfeaa0e00197c3c8ff394552d88dcfaad000f8c

Observation ee7c5b74-6577-4cb7-87cc-c9ea7e5aa1ab · inbound

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models cites this paper.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.338599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.338599Z digest=sha256:4995bc7b4fce5a86c97796f4cfe723fd73e3f4ff8e215bf2e1466baf6fb413ba

Observation e006c770-568d-42b6-a465-2885c785e0ac · inbound

ChronoVision: Temporal Reasoning via Latent State Reconstruction cites this paper.

ChronoVision: Temporal Reasoning via Latent State Reconstruction VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 237

Resolution
unresolved
no resolver link, observed 2026-08-07T01:02:19.185524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:02:19.185524Z digest=sha256:cb48fcd52b2ce6622e41baf1090df15a238c5cd0ab5e47930b3f6685acbc0477

Observation 6f773d29-ae78-4dd9-b3f4-7b028565420b · inbound

Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents cites this paper.

Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T19:29:14.819960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:29:14.819960Z digest=sha256:dab617948cd807c843c9e23ec10c8bdf5be58c5a8b2eb08c563dd3442640ca46