Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2503.19612.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:53:38.658756Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:59:40.139689Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation f0367503-491f-4ba0-b81e-918e5a46db32 · inbound
On a few pitfalls in KL divergence gradient estimation for RL RL-finetuning LLMs from on- and off-policy data with a single algorithm
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 013e6b96-1d08-4581-a879-967f60f4dbef · inbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model RL-finetuning LLMs from on- and off-policy data with a single algorithm
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90f94ae6-a3cc-47aa-b88f-c1b0f8f19fcf · inbound
Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization RL-finetuning LLMs from on- and off-policy data with a single algorithm
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd77c9ef-0340-4704-963b-cd137d466dbe · inbound
Trust Region On-Policy Distillation RL-finetuning LLMs from on- and off-policy data with a single algorithm
Reference 258
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b6d480f-b1fb-4cfd-9e9c-3c2fa01e489d · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning RL-finetuning LLMs from on- and off-policy data with a single algorithm
Reference 198
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.