Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:15:50.932630Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2608.03223.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:15:50.932630Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8675798f-b798-4d3f-8c9c-fdb46028c21f · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ff4f27e9-3352-42d5-86aa-a439fd126d03 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9adefe45-cc4b-402a-a50d-2d780189b19a · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92a3e15b-408b-41b7-aaf8-4ee8842f5ce9 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13156ec5-b935-4634-b9ce-a745c9ddf324 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b138fd6e-ac23-4e3d-87a5-68837a7b8bbb · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Distilling the Knowledge in a Neural Network
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71826fd6-80ac-44a6-9ed0-dd1d21f9e9ef · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baa0c435-e836-4b37-91fd-d4d9b6e538de · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Reinforcement Learning via Self-Distillation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fab17f2f-af4e-492e-8d1e-40c658fa80bc · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ebed2cd-4208-4aa5-bc04-7de91ad3f08c · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d37c7dd3-f9be-4136-8c78-68c3880e9b1d · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb700ff0-c4a3-41e1-926a-7527f6a98784 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Naturalquestions:abenchmarkforquestionansweringresearch
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4bd96b65-717d-43ab-8852-d01b4c418612 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e3be0fea-2032-4c84-8e10-cf1b97b8d851 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Self-Distilled Policy Gradient
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 508bf451-91ee-42c2-ab64-7fdc6610c639 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Self-Distilled Agentic Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 525c4e31-9b8a-4ab2-8e01-d00e5e1783af · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b79254f-3fc5-400c-8af9-09f6132a4b10 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b605da9c-739d-4e15-9d83-ecdac7487221 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1011db5f-b902-4973-b2a2-3ad1da50d45d · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ecdbda85-cc2f-4523-a3cc-f02fe3097b57 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5dfa644a-6dee-4f8e-a030-92755311a9dc · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Qwen2.5 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd82cd44-8337-4a08-8b01-2c0467438122 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 035ec4cb-5dc6-43fc-bbd5-e08891a2cefe · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6d0de95-b5ef-4323-9231-c65a20275e07 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81c9f18d-99bd-4bcc-adde-d2766986be77 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff8d01c5-2bd5-42d1-9e20-50c0d9be4d9a · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Learning by Distilling Context
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9fb4306-7ed0-4c05-affa-c6e28e54c521 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cceeed0-e3aa-4b6f-bc35-7df47754fbee · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 38069836-2da3-4d8a-a6ed-bf13e7a964df · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 332f1b5f-613f-4323-b26f-5a8389fb2d4c · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b747ed9c-5279-4141-b5e1-09f56823bc86 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e49090d-7305-422a-9c4d-dbb8011c04bb · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8857167-3609-4823-94d3-4c04bdb123df · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 580e16ac-39d3-4085-96fb-0f8cca1a2119 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping TIP: Token Importance in On-Policy Distillation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62ec70ad-3164-4a3f-86b3-0c610941cba0 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Qwen3 Technical Report
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e73c43fc-446a-46e9-bf20-de450cbf756f · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Self-Distilled RLVR
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd9b2f95-9564-4cf8-badd-dd4a3e8e1e6f · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cce0ed66-ec96-4eae-b737-28ed9bc82c38 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping WebShop: Towards scalable real-world web interaction with grounded language agents
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b3da775b-c531-4833-81db-c3208aff0302 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping On-Policy Context Distillation for Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5ec1a51-7670-44cb-ae15-62522eb0765f · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24d858c6-4fac-43ac-9b1d-6f58d24bf818 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Proximal Policy Optimization Algorithms
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73286c70-c049-4d27-ad0f-0f428c56a001 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping Transactions of the Association for Computational Linguistics10 (2022), 539–554
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d9033018-4208-4139-a8c5-f5fe7d1a2488 · outbound
Agentic Reinforcement Learning with Self-Distilled Reward Shaping RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.