Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:32:47.290687Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2608.07371.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:32:47.290687Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 43af2e35-8b6b-4a77-b853-75fb78390afb · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning Advances in Neural Information Processing Systems , year =
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d882b2c6-7e5b-40d3-a6bf-9fbb63ddb08f · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning Robotics: Science and Systems XIV , year =
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 45955c89-a13a-4101-a2e5-fb155d762ac8 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2024 , eprint =
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1df07b0-d27a-45fe-b124-7cee04b408f8 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2024 , eprint =
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40cef15e-f13c-4663-947d-2283b7be1d0a · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 590e8453-8ced-4a3d-9d75-1bbe6cf84e16 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning Group-in-Group Policy Optimization for
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 892b1006-3112-42c9-a7be-5675cb4b0bee · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2026 , eprint =
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 598177fa-89c7-4cda-9663-6546f2e75202 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2026 , eprint =
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2004b62c-5367-4e00-8db3-131d370d0c8d · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning Self-Distilled RLVR
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c22479f-2c14-4c62-bf45-b6947037748e · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a497627-3db7-4365-9863-3381caafb392 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning TIP: Token Importance in On-Policy Distillation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40697789-9769-4ddc-8676-b87b8aadbe91 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1457809d-0d90-4605-a119-62cdfcfb6904 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2026 , eprint =
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49d7e6e7-d96b-499d-aff1-cbe99ab49061 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-Distillation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7e4e1bc5-677a-44e2-8e03-d79747a8f281 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2026 , eprint =
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdfa1f23-c784-4d8c-b20b-647e48fbc0f3 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7841065-f830-445d-b8ff-e7406788012d · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6576559-7c75-4aab-835c-1b2429f6e0a4 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32442963-7db3-4cc0-a4f2-71475cb1dfb0 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e93cda6d-58b7-491a-956b-429b01ad9891 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdbfe7a7-1e4e-4aed-af23-6a706b3a777e · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2603.18683 , archivePrefix =
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bdf08b6-28bb-43c5-af36-bf921743eded · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2026 , eprint =
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d65bb8c2-9c52-485c-97f6-ef1902e6c21d · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77e4ed2d-77f4-48d5-801d-8fdba5c16bac · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5823c01d-e450-4f7d-936a-f8d519616ee6 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e3b83a2f-7ea5-4647-958d-16182753004f · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning SAGE-OPD: Selective Agent-Guided Intervention for Multi-Turn On-Policy Distillation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d74bd473-f4b2-426f-8612-6caf4e99b8ba · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdd59071-fe34-4b74-b6ba-4dc6288b7777 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2021 , eprint =
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4debe24c-edf4-44a9-84e5-b13ee961a9eb · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2022 , url =
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4f4d4130-eaeb-44d1-a9bc-c22894d1521b · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2024 , eprint =
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0f9e17d9-cb83-41f2-884e-ca019fbb0777 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2025 , eprint =
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb21a203-6dc1-4dfc-8b53-05e5b062ea59 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2017 , eprint =
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a70b3bff-5819-4816-b7e2-c691bcaf9183 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae0882d1-496a-4562-8be6-b1a4dcaf22aa · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2510.23603 , archivePrefix =
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d8b50da-c695-4d3d-accb-dcc189b748ef · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning InstructSAM: Segment Any Instance with Any Instructions
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e6cf9bf9-8a6c-4b66-8880-d69a209ba3a9 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9f8a0c-ff79-4004-9b1d-67eab96bef69 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2025 , url =
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 185b5178-ac17-4822-a8ed-223a8c273ce1 · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 85e9fcbd-71ac-4e3f-89fb-feff40384c9a · outbound
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
No inbound Pith citation observations are available.