Pith. sign in

Paper Citation Record · LEDGER

Ratio-Variance Regularized Policy Optimization

As of 5 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2605.26784.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.26784 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T19:44:02.317303Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact18
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 25e016ad-833b-4ba4-bb27-138a7328cc14 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Ratio-Variance Regularized Policy Optimization Constitutional AI: Harmlessness from AI Feedback

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.956092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:32c4e965d87be1e09c290754811fbab32faa48e5952dd79906afe6a4002ffcaa

Observation cf08b0b9-2f6f-4d24-a527-dda933894760 · outbound

This paper cites Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition.

Ratio-Variance Regularized Policy Optimization Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.959578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:3b5c7ce32205840931e51e0663bec74cfdcf65e6b2414f7d0dd9942524eb0973

Observation 34485d47-4743-4400-a09d-9c895fef8c0d · outbound

This paper cites Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO.

Ratio-Variance Regularized Policy Optimization Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:53:55.985787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:a0464472b832a8ad76dfbe6b4baf29a3129006e72f0e974bb4296cca4b2609cf

Observation 0a50c1d7-d55d-4e07-b5e7-742036fe99c8 · outbound

This paper cites Emergence of Locomotion Behaviours in Rich Environments.

Ratio-Variance Regularized Policy Optimization Emergence of Locomotion Behaviours in Rich Environments

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.971665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:39afb6502b32940e80d784fe1b4948014e30b5b97452836224fd6c0eb48bf348

Observation fe5c8c06-98f3-45dc-b182-0eb65282cfb3 · outbound

This paper cites SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints.

Ratio-Variance Regularized Policy Optimization SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.977140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:97b13aa50f17dcbefa87766c3ff0eb739c5e7ea578ef5b57e9c8216f626386ea

Observation 80e2cbf4-dd6f-4560-b3c0-e0898b2de06e · outbound

This paper cites What can rl bring to vla generalization? an empirical study.arXiv preprint arXiv:2505.19789.

Ratio-Variance Regularized Policy Optimization What can rl bring to vla generalization? an empirical study.arXiv preprint arXiv:2505.19789

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.971507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:8bc6670bc6c242edd87d8f3e3167fb90a98bb04ec25b0446495f30740ca230fe

Observation df1dc87e-2b68-453a-aad5-bb4965f5b013 · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

Ratio-Variance Regularized Policy Optimization VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.988110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:62223a0fd00d0d65b5d94de60e5bd696c738a5560729e9498e044282ad9be0b5

Observation d4d5ce8c-2441-4c57-ac10-1e8ddc904bb4 · outbound

This paper cites Differentiable Trust Region Layers for Deep Reinforcement Learning.

Ratio-Variance Regularized Policy Optimization Differentiable Trust Region Layers for Deep Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.990855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:57549b78af260e48395387bc7ca45aecdd122c094e8eb3326177c01fabad22e7

Observation 42fb7b7b-e288-4bc2-a5d4-6af7262093e9 · outbound

This paper cites Wasserstein Policy Optimization.

Ratio-Variance Regularized Policy Optimization Wasserstein Policy Optimization

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:53:55.983387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:ade5598d26967af7c8a5764f42b2f5b2920c2adaafbfa7a8e7acee264031d3a2

Observation afaaa5d0-713c-4d07-891a-575637848102 · outbound

This paper cites Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs.

Ratio-Variance Regularized Policy Optimization Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.944800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:3460ee2f999127fa2f7fda34821607a716454ed97d060d22dd374bd92f03505d

Observation 9a287af2-64b6-43f1-8687-db7b7326d15c · outbound

This paper cites Proximal Policy Optimization Algorithms.

Ratio-Variance Regularized Policy Optimization Proximal Policy Optimization Algorithms

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.939103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:b79777f37ad1b11c05273b57e2112cdfda129349506f95a69e0027875ce4c012

Observation b7b3b1a9-39e8-4a27-8d9a-6691c4cc940a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Ratio-Variance Regularized Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.949640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:6fef470a5c98c64a6282c037407164bc2ac99484febe0562bd81aefca6415ddb

Observation f108616f-79d7-43ac-adf5-b3c5c05d03dc · outbound

This paper cites V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control.

Ratio-Variance Regularized Policy Optimization V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:53:55.936626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:54b4b8215bccd05eb449a6d19ad2f25d61b29a2f8f97e4af1e729886932fc311

Observation 7d7a8cf5-ff55-4844-9fda-020b42df42bb · outbound

This paper cites Provably Convergent Policy Optimization via Metric-aware Trust Region Methods.

Ratio-Variance Regularized Policy Optimization Provably Convergent Policy Optimization via Metric-aware Trust Region Methods

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.943707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:3023b5b7085ed743e4361f84f89337eb65d08d62c2f02c1bd16bebd9a113209b

Observation b294cf96-df86-4e2e-be66-3d63b81b54c9 · outbound

This paper cites Klear-reasoner: Advancing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629.

Ratio-Variance Regularized Policy Optimization Klear-reasoner: Advancing reasoning capability via gradient-preserving clipping policy optimization.arXiv preprint arXiv:2508.07629

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.953529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:9776392d638b9c4d1ad03b6ffceec4a93cf652afcced055b50c581c59784c38e

Observation 453e5ccb-c092-45a6-8f8d-1bdab7c52581 · outbound

This paper cites Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models.

Ratio-Variance Regularized Policy Optimization Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T19:53:55.921721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:c2f96d5b5e955f365b4adee57d9921f9b053d52297624a40364ed80a64363d8b

Observation 1dede566-7515-4ec2-825d-a010d3d250db · outbound

This paper cites DeepMind Control Suite.

Ratio-Variance Regularized Policy Optimization DeepMind Control Suite

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.952076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:8b75cbb39e35b178381d0d8e281bdd58029b2280bd3c8a6b82eb2a398d9c91a6

Observation 8635219d-654b-4759-8ec3-c548d38a8919 · outbound

This paper cites Simple Policy Optimization.

Ratio-Variance Regularized Policy Optimization Simple Policy Optimization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.985392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:4bfee6998f631fcd34a54bbd90babcec1a914a4fbc543b190e894023e0bf9145

Observation 59e794fb-451c-4f4a-9e8a-f300173a5c84 · outbound

This paper cites Qwen3 Technical Report.

Ratio-Variance Regularized Policy Optimization Qwen3 Technical Report

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.974209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:f3f1ee70940f8b57d2658960b85d47cc9593df909e65a697cd72a220b4807789

Observation 6f95295a-bbf0-4acb-80a1-458f6d9178eb · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Ratio-Variance Regularized Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:55.976909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:e6bf46e8898e0aa747e7037736db68764ff339537d318993cf47f2cec4f57857

Observation 48afe975-e9c1-4f7a-822a-ab10846bf9df · outbound

This paper cites Guaranteed Trust Region Optimization via Two-Phase KL Penalization.

Ratio-Variance Regularized Policy Optimization Guaranteed Trust Region Optimization via Two-Phase KL Penalization

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.962399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:94a59b92b18df8547ed156d081c04a6f8e8c9525e4999501acf4fe2e82b711b6

Observation e88cb61b-9b0b-444b-90b8-b4ca75de5e4c · outbound

This paper cites A Stochastic Trust-Region Framework for Policy Optimization.

Ratio-Variance Regularized Policy Optimization A Stochastic Trust-Region Framework for Policy Optimization

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.979899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:75333020a49c829050746b6a687a8e0b6980abd7f4e565e8f46e8eff1a15daad

Observation e0dd0bc3-4c10-4463-9dab-39c0008eda06 · outbound

This paper cites R2VPO utilizes adaptive dual updates for LLMs and fixed dual factors for robotics tasks.

Ratio-Variance Regularized Policy Optimization R2VPO utilizes adaptive dual updates for LLMs and fixed dual factors for robotics tasks

Reference 23

Resolution
malformed identifier
no resolver link, observed 2026-06-29T19:44:02.317303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:ce7e0f46bdf6c2c9dc2780f6c16ac64b05d5850c6044500dc8f65212624073d5

Observation 3a3b7f2f-6888-4db5-81bb-9759af692699 · outbound

This paper cites an unresolved cited work.

Ratio-Variance Regularized Policy Optimization Unresolved cited work

Reference 24

Resolution
malformed identifier
arxiv_id, observed 2026-06-29T19:53:55.982603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:a374d9e0a8951ae1bd26abd7e40664c039ac3ef35414a3229cea2f2af4a77e7a

Observation d416ea4d-258f-41ba-9805-266904275ad5 · outbound

This paper cites an unresolved cited work.

Ratio-Variance Regularized Policy Optimization Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-29T19:44:02.317303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:a99526560e4fc5fe6aaeae0eb93c196cbde2b309adc39b3235b8e534888fa043

Observation 9edf4df3-956f-4861-9c1f-508211cae22a · outbound

This paper cites an unresolved cited work.

Ratio-Variance Regularized Policy Optimization Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-29T19:44:02.317303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T19:44:02.317303Z digest=sha256:81295759e05ab8b5b35e4639af95bd04c8ca9bb51d0653cc855590630ad4b17e

Pith citing papers

No inbound Pith citation observations are available.