Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T06:59:15.208890Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2607.24720.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T06:59:15.208890Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 16bf0b82-4f8e-47d9-a0e4-0eefd0935772 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation On-policydistillationoflanguagemodels: Learningfromself-generatedmistakes
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c4888ed-fc61-4432-b14c-145e9db723cc · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Large Language Models for Planning: A Comprehensive and Systematic Survey
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82416f95-36f1-44c0-9a27-f2cb38b140c2 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e45157b-2721-467a-bed1-ad066595600d · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Weak-to-Strong Generalization via Direct On-Policy Distillation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a277c381-ed47-4fa7-b03c-e399c8af5ac9 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e836083-1e7c-45f3-ad44-e227f2ecd23b · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 791b8509-062b-4da4-8a16-583ffbff6c0a · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Co-Evolving Policy Distillation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 993cf3aa-df86-4a6e-bf54-8da691da4a8e · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Minillm: Knowledge distillation of large language models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90d96df3-1b9d-42e8-acd5-0898fb8e0d3e · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Reasoning with language model is planning with world model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25238d1b-0ac7-4b82-92e6-9eec9cdd4dfb · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5434bfa-165b-4ea1-ad5c-e9c0b6fdfefd · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Stable on-policy distillation through adaptive target reformulation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc85a82a-ef9a-4e67-80cd-9eb096c1bdf9 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Rag-rewardbench: Benchmarking reward models in retrieval augmented generation for preference alignment
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfc026a5-6244-44de-8492-e6db46745902 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Scaling reasoning efficiently via relaxed on-policy distillation.arXiv preprint arXiv:2603.11137,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9d4bf5b-0352-47b0-8310-7fe651227173 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Unifying group-relative and self-distillation policy optimization via sample routing.arXiv preprint arXiv:2604.02288, 2026a
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe9de7e-21af-4fdb-b42a-ee37a8a0e882 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8967c0f-f30f-43cd-8ae4-1ac988eed026 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Unlocking the future: Exploring look-ahead planning mechanistic interpretability in large language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2925f4a0-efba-41b0-9bf5-0385742deb9b · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Privileged Information Distillation for Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1f6f190-0a52-4073-af01-d3e5c8f4b24a · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Pope: Learning to reason on hard problems via privileged on-policy exploration.arXiv preprint arXiv:2601.18779,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfe573fb-9d52-4881-b841-f3ef30f99529 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation RL's Razor: Why Online Reinforcement Learning Forgets Less
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5d29081-45c2-4793-8c14-da75e205a861 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Self-Distillation Enables Continual Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc8c05bc-dc42-4c9b-82e9-5a241445fbc8 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation A Survey of On-Policy Distillation for Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 689d5f84-04fd-4bde-9d77-4818006ac36e · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41165b66-84c3-480d-a89e-d0edc0d8059d · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation MiMo-V2-Flash Technical Report
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03cac1b0-4a77-45db-8cf2-a9bc2e5f6294 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Trust Region On-Policy Distillation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c45225d9-c962-4d0b-9068-acaee1725a96 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Deepseek-v4: Towards highly efficient million-token context intelligence.arXiv preprint arXiv:2606.19348,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16f790e2-4841-4574-ae82-48ae17764488 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Self-Distilled RLVR
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc1a627d-9072-4c93-8cd4-93f42b3694c9 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation From𝑓(𝑥) and𝑔(𝑥) to𝑓(𝑔(𝑥)) : Llms learn new skills in rl by composing old ones.arXiv preprint arXiv:2509.25123,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05a40bfd-2784-4590-a07b-851eae744fe3 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34b05c67-3f4e-4f84-bae7-38312f8180cd · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation GLM-5: from Vibe Coding to Agentic Engineering
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cb24b19-1589-4847-8660-84d175fa4b34 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation On the interplay of pre-training, mid-training, and rl on reasoning language models.arXiv preprint arXiv:2512.07783, 2025a
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e34d1c5-958c-46d1-ae55-e7b29e03fa3f · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 454ac1e0-413e-430d-b634-fa8d45812554 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation The row label Category gives the synthetic domain
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe12c3c0-4dd9-4eb4-b3d5-ce5bbf7f5f65 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1d1a489-2701-4d79-b64b-ec9676dc5494 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ada1fc28-ca80-4cb8-8d4e-407cdebbe162 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Internalizingworldmodelsviaself-playfinetuningforagenticrl.arXivpreprintarXiv:2510.15047,
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 989fb6e4-b1c9-447b-bb7c-d86e6fb4b460 · outbound
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.