Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T01:03:50.900234Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2608.03068.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T01:03:50.900234Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 198e8273-1157-42ea-87a4-f1de19bf9c72 · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning DeepSeek-V3 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f49a1024-cbde-4c97-a3d1-15074d109f82 · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6307409-bfff-417d-96c1-a50c2b545310 · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Mihir Prabhudesai, Lili Chen, Alex Ippoliti, Katerina Fragkiadaki, Hao Liu, and Deepak Pathak
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7e09ead-8a90-4a9e-95fc-49075a756c03 · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Maximizing Confidence Alone Improves Reasoning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bb7b8b2-740b-4747-ab40-041cd0009fed · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c48d44f1-d719-44e6-8d60-4d963273726d · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Offline Regularised Reinforcement Learning for Large Language Models Alignment
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7ab1b69-45a5-4627-893c-40d3a1179c0b · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06828e25-d770-4794-91aa-04f2467a6c68 · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb9eb6a6-0b0a-470f-be7b-34ee3c3bd211 · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 656f73ed-543a-4fa8-9343-6d6f4884cfd9 · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 04e254a3-3d8f-4309-b802-dd0885ef1db6 · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 234fcb07-92af-4807-996e-7e34e791be50 · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47b2cca1-ef7f-429d-a4b9-b4f81fd44a93 · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f64dcd2a-c888-4b80-90f1-c759daa5af2f · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53510328-38fc-4371-ab95-f428ae84eb73 · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Proximal Policy Optimization Algorithms
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f05c8805-548d-48d0-8a4b-158208d944e1 · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 471fe5c0-bde5-4e7e-a082-1e7ca1a9f9f6 · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da67af90-fa16-46a0-a5c0-349c22f72ca6 · outbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.