Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T22:45:45.812699Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 5 of 5 outbound references and 0 inbound Pith citation observations for arXiv:2605.30888.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T22:45:45.812699Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
5 of 5 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bc6448ca-57ab-472a-82b4-ec82c98070ac · outbound
The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement Process Reinforcement through Implicit Rewards
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c461e07b-b884-4024-a723-128acf495ea6 · outbound
The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab73f003-ea14-4d3d-8ca0-675ce7b261c7 · outbound
The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6d6ef053-c41a-4f6d-8500-5acf34162c40 · outbound
The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement Proximal Policy Optimization Algorithms
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3b820d05-c515-4f02-8ea9-312c88f200bf · outbound
The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement Group Sequence Policy Optimization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
No inbound Pith citation observations are available.