Pith. sign in

Paper Citation Record · LEDGER

RL-finetuning LLMs from on- and off-policy data with a single algorithm

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2503.19612.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.19612 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:43:44.491352Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:40.139689Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 15d0c076-1e86-41e2-9d49-123c6e2a08f0 · inbound

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models cites this paper.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models RL-finetuning LLMs from on- and off-policy data with a single algorithm

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.491352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.491352Z digest=sha256:27643b87d35201f544fa04f37fb4d59f730a9a9f072355a0a95da7d3e5a49ab8

Observation d4e4e0e5-f2b5-4ddd-bbf6-eee1db7c913d · inbound

ShiQ: Bringing back Bellman to LLMs cites this paper.

ShiQ: Bringing back Bellman to LLMs RL-finetuning LLMs from on- and off-policy data with a single algorithm

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:57.608629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:57.608629Z digest=sha256:126fc9ba8f54068940d3aa10cb972bf2e37a77be9a853035f8d25efdbc04cb3a

Observation f0367503-491f-4ba0-b81e-918e5a46db32 · inbound

On a few pitfalls in KL divergence gradient estimation for RL cites this paper.

On a few pitfalls in KL divergence gradient estimation for RL RL-finetuning LLMs from on- and off-policy data with a single algorithm

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:38.658756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:38.658756Z digest=sha256:dcb7c1bfb09e1564fb2e68e40d6308228860bbcbb6814b29ad93d612bfee9a03

Observation 013e6b96-1d08-4581-a879-967f60f4dbef · inbound

Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model cites this paper.

Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model RL-finetuning LLMs from on- and off-policy data with a single algorithm

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:38.706318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:57:38.706318Z digest=sha256:41ed334b06b2050567cd1fd8b6a1926b85632e59296fdca61d15d8adb33fa56d

Observation 90f94ae6-a3cc-47aa-b88f-c1b0f8f19fcf · inbound

Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization cites this paper.

Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization RL-finetuning LLMs from on- and off-policy data with a single algorithm

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:39.656738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-08T18:17:43.956278Z digest=sha256:72036fe8233acb6add9077c2a99500f16ec9621cecafb747d506ffa609b7c6c0

Observation cd77c9ef-0340-4704-963b-cd137d466dbe · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation RL-finetuning LLMs from on- and off-policy data with a single algorithm

Reference 258

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.658389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:a3cd27713d767da67ec25fab81cdbaf82cb586986e089bfefdd26e4cc2bbdf12

Observation 8b6d480f-b1fb-4cfd-9e9c-3c2fa01e489d · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning RL-finetuning LLMs from on- and off-policy data with a single algorithm

Reference 198

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.142556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:90345f471e57f479aa26d0ae4613a5689630bc00d197ec0feef23e35c62122c6