Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 67 inbound Pith citation observations for arXiv:2410.01679.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:07:08.079012Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T11:16:11.296273Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation d46ce63e-8996-4551-81a3-11bee3ec1ad3 · inbound
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 09915755-9dcb-4e1b-b0af-9d6ef12953fd · inbound
A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24be95d8-0663-42f6-9192-92e5cd642d3e · inbound
SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d43bdd78-07f9-409b-8162-5646d56c8996 · inbound
AlphaPO: Reward Shape Matters for LLM Alignment VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c82506e-c364-4e71-897b-b8ac5ab0a83f · inbound
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7264ae2e-8654-42d1-9a4e-1d75ffc76cea · inbound
T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 660afd43-a456-4ade-93de-089f31c37342 · inbound
Process Reinforcement through Implicit Rewards VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d8060d98-83f4-4402-b46b-a9de15137134 · inbound
Reinforcement Learning for Long-Horizon Interactive LLM Agents VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8d070c7-893d-42df-86e7-3ed7693068de · inbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66fa54c1-421c-4938-8538-0c19b137cf03 · inbound
Learning to Reason at the Frontier of Learnability VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 531ef94a-12a6-4835-b0ef-ade283579f4a · inbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9884ecf5-e565-4cd1-8f51-55f1320569f9 · inbound
Reinforcement Learning from Human Feedback VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 165
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 679cdec8-d62f-4978-b54c-89a068006e56 · inbound
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5c437fa1-e43b-4fe2-9ac4-1f9243aeb264 · inbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e3dc08e0-b3c8-4c0a-8ac5-5e1252dd3826 · inbound
Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fd4e5a9-2fff-427d-9069-f95664080fc9 · inbound
Group-in-Group Policy Optimization for LLM Agent Training VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 63b5de8a-0e45-4078-b03c-0c1fcbee1a0e · inbound
Effective Reinforcement Learning for Reasoning in Language Models VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5139bbd4-11cc-4431-be00-5e305202204e · inbound
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 743dc8c4-c296-4e3d-837d-dbfc0f4d43cc · inbound
How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aedd8de2-cdae-4679-b01c-fc73610148f5 · inbound
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 229e0672-ac94-41c0-aad3-7d737231e735 · inbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3157642f-77c6-4d89-837c-c2ca43862907 · inbound
A Survey on Large Language Models for Mathematical Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecaff3a5-cd06-4143-839f-ef6f18d77e92 · inbound
Intent Factored Generation: Unleashing the Diversity in Your Language Model VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebe55ede-7678-4f2c-8c1e-c65490ac9874 · inbound
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abef9125-2b0d-482e-9369-b428b999344d · inbound
DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7659bb44-44de-4f05-b2e4-a30d560d3f6d · inbound
AdapThink: Adaptive Thinking Preferences for Reasoning Language Model VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a228f38-abdb-453c-813b-0609cfd6acc2 · inbound
Enhancing Large Language Models through Structured Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fcc9aa5-d6d3-420c-90c1-532a17c17fe1 · inbound
First Return, Entropy-Eliciting Explore VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eed7fb36-b81f-4e42-afde-45576bac2ccc · inbound
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b07b515e-c171-45cd-8835-0d1c08964fb0 · inbound
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6179bf92-2916-48d4-8521-503ca9af2743 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 258df9e2-bacb-48fb-a5fc-371f2e2d24b5 · inbound
GPO: Learning from Critical Steps to Improve LLM Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fb923ae-a3f1-4e42-be94-8fce0d68c6c1 · inbound
Polychromic Objectives for Reinforcement Learning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 054fc3b3-4d4f-442f-83f4-f9d546c5bdd7 · inbound
Beyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional Pairing VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 859fe84f-2850-4502-9cc9-6efb93a0ed50 · inbound
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 292aef98-a97d-4cc6-84ce-9f2450eb9a4d · inbound
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdafd0cb-8359-4589-80fa-a54d0b1c4e9f · inbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f1fdd7b-4fd3-4395-bb15-b8b36d29d35f · inbound
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3e47f047-efc4-40a2-a348-a2ad4ba492e5 · inbound
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3e1195a2-9c35-4c76-8bb9-831c3d69e442 · inbound
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dbe3be7e-f002-44ec-b718-ec4a4f173752 · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 60df7934-c0f0-41a7-8f8c-79b03742e514 · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c42d6553-ecef-4eb9-b9a6-0a35e43a5894 · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b3006c93-5a7d-400d-ac68-89784e893364 · inbound
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4c6eab82-1882-4d0a-ba4e-38c3655da520 · inbound
Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 285b56dd-05cd-480c-9a83-30f501a7bdb8 · inbound
Learning from Language Feedback via Variational Policy Distillation VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a054cb16-c7c7-4cb8-960d-33c1654713d8 · inbound
LamPO: A Lambda Style Policy Optimization for Reasoning Language Models VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cdcf3e34-b79e-4657-b783-19a8cfee7570 · inbound
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2000b893-ccfe-487d-8945-36f0bcf0c6b4 · inbound
Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 313d9a8a-2f6c-4ca2-bfb1-5f815fd7ac08 · inbound
ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c4c6f982-3795-4de8-ac5b-69e43e76e27d · inbound
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f062efd6-2eb7-483f-ab81-f3370396598f · inbound
Self-Distilled Policy Gradient VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 026f94e9-4d8e-4b7c-a18b-9b97c165eef9 · inbound
ProcessThinker: Enhancing Multi-modal Large Language Models Reasoning via Rollout-based Process Reward VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c80068eb-114e-436d-abe5-6da4aa2a9fea · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 80f80eab-c7f2-4122-a141-22f610ac46ad · inbound
CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 887dfdeb-cc6e-4321-a082-49098d3643e3 · inbound
ECHO: Learning Epistemically Adaptive Language Agents with Turn-Level Credit VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1f894e4c-bd00-49c7-b473-7880171e54e1 · inbound
Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e61ab1e-c568-4a6d-9cbf-62b4391d7674 · inbound
ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f71504bb-3335-4a32-adbd-5d17756fe24e · inbound
Weak-to-Strong Generalization via Direct On-Policy Distillation VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e68e1632-e745-4550-a34c-fc003f8ded03 · inbound
Weak-to-Strong Generalization via Direct On-Policy Distillation VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb0895ee-fa12-415a-95cf-58c0b8c7ecdc · inbound
RLVP: Penalize the Path, Reward the Outcome VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0152b37e-a63f-4c93-ba80-b264a7d4efdc · inbound
Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3494809f-a207-4435-950f-a4ad60214071 · inbound
Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61fc87b8-3b93-4468-85a1-b746303ea32b · inbound
Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee7c5b74-6577-4cb7-87cc-c9ea7e5aa1ab · inbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e006c770-568d-42b6-a465-2885c785e0ac · inbound
ChronoVision: Temporal Reasoning via Latent State Reconstruction VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 237
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f773d29-ae78-4dd9-b3f4-7b028565420b · inbound
Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.