Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:07.147879Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2608.04788.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:07.147879Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c1ad2c2e-7897-4a58-936b-c7abc4d9cffe · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation On-Policy Delta Distillation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation de8a8bae-340f-406b-b2b5-d2ebd0a6a75e · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Rethinking On-Policy Self-Distillation for Thinking Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0342944-478a-44d2-bb1f-10c5d83b5640 · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7417754-ad18-4cea-89c2-5e37ef6c84f9 · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Self-Distilled Agentic Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99fa61dd-c8d5-467b-9b10-929fafded276 · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 580e6fef-625d-424c-bd32-c7274ac12b97 · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 59af0db5-903e-4c9c-9a80-4faba569bfca · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce75a16e-460f-4b37-89c4-a00d3130759c · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation InInternational Conference on Learning Representations, volume 2025, 89490–89520
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 38de8315-b88a-46a4-82dc-19a2ec9350d9 · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a54fa22c-d4a9-4e5c-b091-7fcec56b98e6 · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1b4eb0f6-4658-4a91-9f08-3b6b2855efc0 · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation On the Position Bias of On-Policy Distillation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7533a37c-84f9-42ba-ace3-070ad07a3177 · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0bb6da02-090c-412b-b582-0e459ef46997 · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Beyond Absolute Imitation: Anchored Residual Guidance for Privileged On-Policy Distillation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5f0a1dd7-5d11-46a4-a987-f8e4f6439289 · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff4914d2-bbe1-4cef-8ac5-defa24f6e7f9 · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50411b58-ae63-4659-8875-6ab149674276 · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70e865c4-6f3f-409b-8835-4cdeb391d57d · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation OpenVLA: An Open-Source Vision-Language-Action Model
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a89a410f-ac14-4e8c-a089-12281d578d34 · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51b017bc-e9e1-47fc-b9ac-05e60814ae8a · outbound
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy
Reference 2026
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.