Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2407.16574.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:31.365581Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T13:14:10.944002Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 70da07b1-0cf1-44f1-ba5c-0c3613c1b6ab · inbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b795ecb2-01f1-4d2c-9ace-c13414b4b50b · inbound
SGPO: Self-Generated Preference Optimization based on Self-Improver TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf5c47bb-8e85-47b3-a615-562ec3d23ee3 · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9426f27c-1d05-4acf-8f71-d3236dfb0862 · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69356ac8-eeec-4bed-9ebf-82ed49ed2f7a · inbound
Stabilizing Policy Optimization via Logits Convexity TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.