Pith. sign in

Paper Citation Record · LEDGER

On-Policy Replay for Continual Supervised Fine-Tuning

As of 21 August 2026, this Paper Citation Record lists 3 of 3 outbound references and 1 inbound Pith citation observation for arXiv:2605.29495.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.29495 v1

Coverage vector

measured 3 of 3 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T09:05:33.034464Z

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T00:38:41.160077Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

3 of 3 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fb37a9e9-4231-4a3e-9305-faa28754926a · outbound

This paper cites GeRe: Towards Efficient Anti-Forgetting in Continual Learning of LLM via General Samples Replay.

On-Policy Replay for Continual Supervised Fine-Tuning GeRe: Towards Efficient Anti-Forgetting in Continual Learning of LLM via General Samples Replay

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T09:13:16.302628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T09:05:33.034464Z digest=sha256:254b0c084e9a102614e5312e3eab5e4c541327cf4df04db0a61b07d989a8055c

Observation 2a250fa2-bb35-4187-92c4-f137bbba3b37 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

On-Policy Replay for Continual Supervised Fine-Tuning Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T09:13:16.299992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T09:05:33.034464Z digest=sha256:049eb9d0bdee1cf4dac82ba9babd70b5ef9390887b53e6c45f8c9cd781f291dc

Observation c741a415-7a1b-4a47-a9f7-aeb649339e89 · outbound

This paper cites Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning.

On-Policy Replay for Continual Supervised Fine-Tuning Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T09:13:16.305324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T09:05:33.034464Z digest=sha256:9930dee27128dc8b946eb894ea4d0a4eff1c0154cccd4a23ce7cb158dab818e4

Pith citing papers

Observation d06789e2-a277-43d9-bd79-ce7a9ea4405e · inbound

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models cites this paper.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models On-Policy Replay for Continual Supervised Fine-Tuning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:41.160077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:41.160077Z digest=sha256:0ff0c95a3de0a4d7dda5b06d40194e565d1919fd5e44c513c0bc11aa2b556dcf