Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T19:44:02.317303Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2605.26784.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T19:44:02.317303Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 25e016ad-833b-4ba4-bb27-138a7328cc14 · outbound
Ratio-Variance Regularized Policy Optimization Constitutional AI: Harmlessness from AI Feedback
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cf08b0b9-2f6f-4d24-a527-dda933894760 · outbound
Ratio-Variance Regularized Policy Optimization Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 34485d47-4743-4400-a09d-9c895fef8c0d · outbound
Ratio-Variance Regularized Policy Optimization Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0a50c1d7-d55d-4e07-b5e7-742036fe99c8 · outbound
Ratio-Variance Regularized Policy Optimization Emergence of Locomotion Behaviours in Rich Environments
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fe5c8c06-98f3-45dc-b182-0eb65282cfb3 · outbound
Ratio-Variance Regularized Policy Optimization SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 80e2cbf4-dd6f-4560-b3c0-e0898b2de06e · outbound
Ratio-Variance Regularized Policy Optimization What can rl bring to vla generalization? an empirical study.arXiv preprint arXiv:2505.19789
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df1dc87e-2b68-453a-aad5-bb4965f5b013 · outbound
Ratio-Variance Regularized Policy Optimization VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4d5ce8c-2441-4c57-ac10-1e8ddc904bb4 · outbound
Ratio-Variance Regularized Policy Optimization Differentiable Trust Region Layers for Deep Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 42fb7b7b-e288-4bc2-a5d4-6af7262093e9 · outbound
Ratio-Variance Regularized Policy Optimization Wasserstein Policy Optimization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation afaaa5d0-713c-4d07-891a-575637848102 · outbound
Ratio-Variance Regularized Policy Optimization Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a287af2-64b6-43f1-8687-db7b7326d15c · outbound
Ratio-Variance Regularized Policy Optimization Proximal Policy Optimization Algorithms
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b7b3b1a9-39e8-4a27-8d9a-6691c4cc940a · outbound
Ratio-Variance Regularized Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f108616f-79d7-43ac-adf5-b3c5c05d03dc · outbound
Ratio-Variance Regularized Policy Optimization V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7d7a8cf5-ff55-4844-9fda-020b42df42bb · outbound
Ratio-Variance Regularized Policy Optimization Provably Convergent Policy Optimization via Metric-aware Trust Region Methods
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b294cf96-df86-4e2e-be66-3d63b81b54c9 · outbound
Ratio-Variance Regularized Policy Optimization Klear-reasoner: Advancing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 453e5ccb-c092-45a6-8f8d-1bdab7c52581 · outbound
Ratio-Variance Regularized Policy Optimization Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1dede566-7515-4ec2-825d-a010d3d250db · outbound
Ratio-Variance Regularized Policy Optimization DeepMind Control Suite
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8635219d-654b-4759-8ec3-c548d38a8919 · outbound
Ratio-Variance Regularized Policy Optimization Simple Policy Optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 59e794fb-451c-4f4a-9e8a-f300173a5c84 · outbound
Ratio-Variance Regularized Policy Optimization Qwen3 Technical Report
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6f95295a-bbf0-4acb-80a1-458f6d9178eb · outbound
Ratio-Variance Regularized Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 48afe975-e9c1-4f7a-822a-ab10846bf9df · outbound
Ratio-Variance Regularized Policy Optimization Guaranteed Trust Region Optimization via Two-Phase KL Penalization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e88cb61b-9b0b-444b-90b8-b4ca75de5e4c · outbound
Ratio-Variance Regularized Policy Optimization A Stochastic Trust-Region Framework for Policy Optimization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e0dd0bc3-4c10-4463-9dab-39c0008eda06 · outbound
Ratio-Variance Regularized Policy Optimization R2VPO utilizes adaptive dual updates for LLMs and fixed dual factors for robotics tasks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a3b7f2f-6888-4db5-81bb-9759af692699 · outbound
Ratio-Variance Regularized Policy Optimization Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d416ea4d-258f-41ba-9805-266904275ad5 · outbound
Ratio-Variance Regularized Policy Optimization Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9edf4df3-956f-4861-9c1f-508211cae22a · outbound
Ratio-Variance Regularized Policy Optimization Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.