Pith. sign in

Paper Citation Record · LEDGER

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation

As of 8 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2608.04788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04788 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:07.147879Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c1ad2c2e-7897-4a58-936b-c7abc4d9cffe · outbound

This paper cites On-Policy Delta Distillation.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation On-Policy Delta Distillation

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.666528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:05.523794Z digest=sha256:d6ff673c68b25ef2c271b5839aac5427a13052463af05c17c2beb7a79e0f7d6c

Observation de8a8bae-340f-406b-b2b5-d2ebd0a6a75e · outbound

This paper cites Rethinking On-Policy Self-Distillation for Thinking Models.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Rethinking On-Policy Self-Distillation for Thinking Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.711475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.711475Z digest=sha256:92a6e8d90a0cd858feefef1da8b8cb3eb562483bd8913ceba38061843ebaa114

Observation c0342944-478a-44d2-bb1f-10c5d83b5640 · outbound

This paper cites What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.901164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.901164Z digest=sha256:e3d8a825244f8ee32d5d088f965bca6b984d48478508acbbf183466b43dd2f6b

Observation e7417754-ad18-4cea-89c2-5e37ef6c84f9 · outbound

This paper cites Self-Distilled Agentic Reinforcement Learning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Self-Distilled Agentic Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.959982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.959982Z digest=sha256:8c002496f80a70e79078c85e402db49d871c3afd01ca0c88fd288d4fef3a5840

Observation 99fa61dd-c8d5-467b-9b10-929fafded276 · outbound

This paper cites RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:06.084807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:06.084807Z digest=sha256:096273f58af5f804151ffbf583341fd21b4afef5d3b6cfb07acc0c3b9e199683

Observation 580e6fef-625d-424c-bd32-c7274ac12b97 · outbound

This paper cites Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.431166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.198738Z digest=sha256:4c9b22ef7c2db1a2a66a8746dbde7aa708a80ce3697c88b0b799795b11ddb68c

Observation 59af0db5-903e-4c9c-9a80-4faba569bfca · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:06.275961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:06.275961Z digest=sha256:feb0a4674e409fb965630733966432aaed487086eb492c417a57a43294c83039

Observation ce75a16e-460f-4b37-89c4-a00d3130759c · outbound

This paper cites InInternational Conference on Learning Representations, volume 2025, 89490–89520.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation InInternational Conference on Learning Representations, volume 2025, 89490–89520

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:39:09.119131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.491333Z digest=sha256:48f7689bff0724b736c0742461acfda9d6c5fabab866a32104426fa15338d52c

Observation 38de8315-b88a-46a4-82dc-19a2ec9350d9 · outbound

This paper cites PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.247916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.584525Z digest=sha256:12f4cdd86f86f79a1378a176f6b7692d1688f406645c3ec31f624f9ff4b0a2be

Observation a54fa22c-d4a9-4e5c-b091-7fcec56b98e6 · outbound

This paper cites SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.054402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.711304Z digest=sha256:8d984f8c83a930084723c71c2c9455a454d22a900d4bfc15b261fb9b8c229d67

Observation 1b4eb0f6-4658-4a91-9f08-3b6b2855efc0 · outbound

This paper cites On the Position Bias of On-Policy Distillation.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation On the Position Bias of On-Policy Distillation

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:07.816552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.783612Z digest=sha256:49fe26b7797b3326b87ff91d7646f19d154f767d0dfcb7a2916d3889c2b0e670

Observation 7533a37c-84f9-42ba-ace3-070ad07a3177 · outbound

This paper cites Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:07.586658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.851794Z digest=sha256:1e5254e2799fd198ffd1001978f1f654783adefca2c98f06a352cb94b3f34ae6

Observation 0bb6da02-090c-412b-b582-0e459ef46997 · outbound

This paper cites Beyond Absolute Imitation: Anchored Residual Guidance for Privileged On-Policy Distillation.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Beyond Absolute Imitation: Anchored Residual Guidance for Privileged On-Policy Distillation

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:07.350446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.947603Z digest=sha256:4b71a4ef94fc09a6eec8b83112beadb62d1d431d85d6f05fc02ce8c9f30a4c72

Observation 5f0a1dd7-5d11-46a4-a987-f8e4f6439289 · outbound

This paper cites StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:07.025064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:07.025064Z digest=sha256:4d50b0553b0f666becf8457dffb2b9df0e22553c29a6c3af6fbe67fc61e4f55a

Observation ff4914d2-bbe1-4cef-8ac5-defa24f6e7f9 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:07.147879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:07.147879Z digest=sha256:7239a992cbf1510c4652ff74f037e428a1d0dd023e5e7c74070f8a1bcfeee347

Observation 50411b58-ae63-4659-8875-6ab149674276 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:06.391541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:06.391541Z digest=sha256:911a730ff748bb28aca418bf9606a354e0219dffa960e7eca88e3e2d3283bf19

Observation 70e865c4-6f3f-409b-8835-4cdeb391d57d · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.816250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.816250Z digest=sha256:fb3e44820d565042eb790e73d524313215fb59d05a05351aeb0976c3e48f6791

Observation a89a410f-ac14-4e8c-a089-12281d578d34 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.647351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.647351Z digest=sha256:a037c94276bda7909c7f4ac9ade5186a88ec1b0ef4e3374b55b0347624252eac

Observation 51b017bc-e9e1-47fc-b9ac-05e60814ae8a · outbound

This paper cites Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy

Reference 2026

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.874780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:05.412903Z digest=sha256:75034fb18629fac5af4269f0d630c19e7961239a5a0cb7ade86ad4fc57967e58

Pith citing papers

No inbound Pith citation observations are available.