Pith. sign in

Paper Citation Record · LEDGER

Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2303.17396.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.17396 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T11:28:31.931515Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5edb3c11-8fc5-4400-b5af-2af1766ec643 · inbound

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling cites this paper.

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:40:46.202696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T23:39:39.018498Z digest=sha256:2d21a9634da67802ed04e5db0b44d1dfad1e136f044e94dc41e6ce1ed2069fe7

Observation ea6b5855-b0e0-4111-a622-dda54e4e03f6 · inbound

Reinforcement Learning with Action Chunking cites this paper.

Reinforcement Learning with Action Chunking Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:22:06.420671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T05:18:20.960945Z digest=sha256:14ca24eb084dff3cb96adde95f2f0d6946b41f8601a90d30dd468bbdcb096f70

Observation d2898d64-945d-4361-81bc-da44e88c9c50 · inbound

COOPO: Cyclic Offline-Online Policy Optimization Algorithm cites this paper.

COOPO: Cyclic Offline-Online Policy Optimization Algorithm Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:13:18.014008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T13:11:16.568415Z digest=sha256:e513eb98bad5bd752486f84d4f5da3fdd65cd8d43605dfdca6e7560cba04eb2d

Observation f41da937-d31d-4bf2-a611-c38b899f92fa · inbound

Improving Robotic Generalist Policies via Flow Reversal Steering cites this paper.

Improving Robotic Generalist Policies via Flow Reversal Steering Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:48:35.713169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:20:19.209180Z digest=sha256:b017711a7fe15b1c9114457a774b67b562b49e280608765fa1fb14151b61d818

Observation 7fa6845a-7bf4-4fe6-a75b-8ce07cb021ef · inbound

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors cites this paper.

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:30:07.881241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T21:15:07.735336Z digest=sha256:9e616cddcf8159b1c34b58a852dd4c73d62b178609c482d7d218f989f206788f

Observation ca81e43d-d058-45b3-b7a0-2de550c049aa · inbound

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning cites this paper.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:bfb0677c9a7c80c759fd033ed2298b1039e5c1d5f82255c672515f515afba5f9

Observation 87b90e74-70a7-421b-87d2-1abcf9db6afc · inbound

Evaluating Fuzz Testing for Reinforcement Learning Agents cites this paper.

Evaluating Fuzz Testing for Reinforcement Learning Agents Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T11:28:31.931515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:28:31.931515Z digest=sha256:6a33ac77a02222e5cf78690bc06c3c67babd87408e0e4c8eede46a91c32b1173