Pith. sign in

Paper Citation Record · LEDGER

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation

As of 9 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2608.04788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04788 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:07.147879Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c1ad2c2e-7897-4a58-936b-c7abc4d9cffe · outbound

This paper cites On-Policy Delta Distillation.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation On-Policy Delta Distillation

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.666528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:05.523794Z digest=sha256:ceb17233ad51a5456227685478fff9f6d7d3d3a2702a175dce92d5b81ca77199

Observation de8a8bae-340f-406b-b2b5-d2ebd0a6a75e · outbound

This paper cites Rethinking On-Policy Self-Distillation for Thinking Models.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Rethinking On-Policy Self-Distillation for Thinking Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.711475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.711475Z digest=sha256:07861a075ea0851d55246bfebabb4b09af57891215afb94a7d2322fde51175cc

Observation c0342944-478a-44d2-bb1f-10c5d83b5640 · outbound

This paper cites What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.901164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.901164Z digest=sha256:9f05d59b7bfba12cd7b70ee58df43306b634d792b92c83d12b5fa3f36768d926

Observation e7417754-ad18-4cea-89c2-5e37ef6c84f9 · outbound

This paper cites Self-Distilled Agentic Reinforcement Learning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Self-Distilled Agentic Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.959982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.959982Z digest=sha256:c17772726dfd5f546bce643b0dedcd38b155d0ff794ffdc0a5fd4f4b4d57bffb

Observation 99fa61dd-c8d5-467b-9b10-929fafded276 · outbound

This paper cites RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:06.084807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:06.084807Z digest=sha256:dfcada7a536de9e9c5f6617b167d5d9e8309f6b2e821bd0e1c5b2fb8c58db8bf

Observation 580e6fef-625d-424c-bd32-c7274ac12b97 · outbound

This paper cites Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.431166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.198738Z digest=sha256:822b35da7132c1e85b94748579fdc522c89b84d1ea431b9de860ba5a7c6db885

Observation 59af0db5-903e-4c9c-9a80-4faba569bfca · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:06.275961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:06.275961Z digest=sha256:ecce951932d8255cd1e9fdc2c7cd07af9611110372f5edac76b2c92ce82e5d77

Observation ce75a16e-460f-4b37-89c4-a00d3130759c · outbound

This paper cites InInternational Conference on Learning Representations, volume 2025, 89490–89520.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation InInternational Conference on Learning Representations, volume 2025, 89490–89520

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:39:09.119131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.491333Z digest=sha256:63c25d6dc8b1cb1c665781e2b2140e63cd70a491e12e2dc8bfa8399f97047e9b

Observation 38de8315-b88a-46a4-82dc-19a2ec9350d9 · outbound

This paper cites PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.247916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.584525Z digest=sha256:36ca20dcbcffbc2780ef1e3f4244fca115ec794ddd0068154e8c5d7840a22f39

Observation a54fa22c-d4a9-4e5c-b091-7fcec56b98e6 · outbound

This paper cites SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.054402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.711304Z digest=sha256:640b0873f23e66e568946081fcb430ebe2696b1f4a622eb721f190ac36e99ed7

Observation 1b4eb0f6-4658-4a91-9f08-3b6b2855efc0 · outbound

This paper cites On the Position Bias of On-Policy Distillation.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation On the Position Bias of On-Policy Distillation

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:07.816552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.783612Z digest=sha256:3b734090c1c2321e885bd49755dc717018635d371d5675ad6b9692dda089ed0a

Observation 7533a37c-84f9-42ba-ace3-070ad07a3177 · outbound

This paper cites Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:07.586658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.851794Z digest=sha256:f65d952fa9791eb40fbd0d071fd62ce224b5be298b38229196d3733d6b5e9ee2

Observation 0bb6da02-090c-412b-b582-0e459ef46997 · outbound

This paper cites Beyond Absolute Imitation: Anchored Residual Guidance for Privileged On-Policy Distillation.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Beyond Absolute Imitation: Anchored Residual Guidance for Privileged On-Policy Distillation

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:07.350446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.947603Z digest=sha256:20e92be0b8795a6e4425a5b08ba64e746d86fcb70852754b80e006707f73f222

Observation 5f0a1dd7-5d11-46a4-a987-f8e4f6439289 · outbound

This paper cites StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:07.025064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:07.025064Z digest=sha256:a04da4c2f3f41ecf2f8b55cda06dc875ea171a709277ffa51365ed11c54e1910

Observation ff4914d2-bbe1-4cef-8ac5-defa24f6e7f9 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:07.147879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:07.147879Z digest=sha256:37cd74cb7a8f7e0a86bff85c3cdaf3d3174156b8c6fd9c611e8330dfd241a07e

Observation 50411b58-ae63-4659-8875-6ab149674276 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:06.391541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:06.391541Z digest=sha256:0c5064fb220473166bea2174fe6359f56f3000cd113ccdb4430bfcc75d1f81d0

Observation 70e865c4-6f3f-409b-8835-4cdeb391d57d · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.816250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.816250Z digest=sha256:2d9b1fbe2f3ef808310438239bfecb99a872734387c205366e14eef6eaf5ed46

Observation a89a410f-ac14-4e8c-a089-12281d578d34 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.647351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.647351Z digest=sha256:7edb91596f28682db5370bc51e513c609cdc10742a90f44ee78f41736838ccde

Observation 51b017bc-e9e1-47fc-b9ac-05e60814ae8a · outbound

This paper cites Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy

Reference 2026

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.874780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:05.412903Z digest=sha256:4bfb501ad69a6fc41f08b541dbd2f131daff9abae30998540b0ad840471a895f

Pith citing papers

No inbound Pith citation observations are available.