Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2504.04524.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T06:47:16.873682Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:39:42.243396Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 8a53ea4e-2d10-467b-868c-15af3475e4ae · inbound
Learning to Reason under Off-Policy Guidance Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 18f62dfc-7d9f-4739-b673-f9e314a89751 · inbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d354e495-9260-4fa8-a6ce-32bd312f9be2 · inbound
Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fccb151c-0cc3-4f93-83c2-cf182470efc2 · inbound
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 19fedf20-ab9c-4e66-b1ab-2428440dc814 · inbound
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7ef419e3-2fbe-40db-8337-6893563729c0 · inbound
Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.