Pith. sign in

Paper Citation Record · LEDGER

RL-finetuning LLMs from on- and off-policy data with a single algorithm

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2503.19612.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.19612 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:53:38.658756Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:40.139689Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f0367503-491f-4ba0-b81e-918e5a46db32 · inbound

On a few pitfalls in KL divergence gradient estimation for RL cites this paper.

On a few pitfalls in KL divergence gradient estimation for RL RL-finetuning LLMs from on- and off-policy data with a single algorithm

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:38.658756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:38.658756Z digest=sha256:85822fec12425c9e40c623fd717cf01f80bb4527f6fcf70ac1dd04207856a8ad

Observation 013e6b96-1d08-4581-a879-967f60f4dbef · inbound

Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model cites this paper.

Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model RL-finetuning LLMs from on- and off-policy data with a single algorithm

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:38.706318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:57:38.706318Z digest=sha256:87c6cba18df13fead48df672d049890f13446de939211d9e16b4b3ea529bfd28

Observation 90f94ae6-a3cc-47aa-b88f-c1b0f8f19fcf · inbound

Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization cites this paper.

Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization RL-finetuning LLMs from on- and off-policy data with a single algorithm

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:39.656738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T18:17:43.956278Z digest=sha256:88909ce522770c42a12e92faea397d2c36810ed659a5a5089a4110a80baade8f

Observation cd77c9ef-0340-4704-963b-cd137d466dbe · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation RL-finetuning LLMs from on- and off-policy data with a single algorithm

Reference 258

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.658389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:6e0d214718648351500a37ffd3e6668a4155528586d507f14322aa3d8ae0873b

Observation 8b6d480f-b1fb-4cfd-9e9c-3c2fa01e489d · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning RL-finetuning LLMs from on- and off-policy data with a single algorithm

Reference 198

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.142556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:5658704c3bfd28693c16d6bb436bdf221394057595c7cb076fe4dd41644090ed