Pith. sign in

Paper Citation Record · LEDGER

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

As of 14 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2607.24720.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.24720 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T06:59:15.208890Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved35
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 16bf0b82-4f8e-47d9-a0e4-0eefd0935772 · outbound

This paper cites On-policydistillationoflanguagemodels: Learningfromself-generatedmistakes.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation On-policydistillationoflanguagemodels: Learningfromself-generatedmistakes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:14.873544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:14.873544Z digest=sha256:3cf535443b89841692a639ce6ad5fc6122ceefec48d19ac7b32ff1a23e146f68

Observation 6c4888ed-fc61-4432-b14c-145e9db723cc · outbound

This paper cites Large Language Models for Planning: A Comprehensive and Systematic Survey.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Large Language Models for Planning: A Comprehensive and Systematic Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:14.883715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:14.883715Z digest=sha256:0368583c66c0f47bc331fb3f8c96746629a591b55496444c0dd46c5b48374aa1

Observation 82416f95-36f1-44c0-9a27-f2cb38b140c2 · outbound

This paper cites AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:14.907995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:14.907995Z digest=sha256:e3604d9278b54bfb3fc5426236d3453bcd65d76d3731d2c6a651ad3e924f6bc2

Observation 9e45157b-2721-467a-bed1-ad066595600d · outbound

This paper cites Weak-to-Strong Generalization via Direct On-Policy Distillation.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Weak-to-Strong Generalization via Direct On-Policy Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:14.911227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:14.911227Z digest=sha256:661ec0e10cdf2da5cd5332340a2babc05414e862163a84c87f1ab90e74f9290c

Observation a277c381-ed47-4fa7-b03c-e399c8af5ac9 · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:14.914845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:14.914845Z digest=sha256:c22a6afcf62fb3887db6c64ba0eb3f0b780ffaa195440973efd31c3a710d8dbb

Observation 9e836083-1e7c-45f3-ad44-e227f2ecd23b · outbound

This paper cites MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation MAGELLAN: Metacognitive predictions of learning progress guide autotelic LLM agents in large goal spaces

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:14.917904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:14.917904Z digest=sha256:a38cddad7e4b68e5cecc1125fd1d6c9a287abbc56763758dca9df173c5dfaf97

Observation 791b8509-062b-4da4-8a16-583ffbff6c0a · outbound

This paper cites Co-Evolving Policy Distillation.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Co-Evolving Policy Distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:14.920841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:14.920841Z digest=sha256:4dcd3bf6fcd0cb34effcad37a2530cef670d0182c7480219bb7204abad3f7f20

Observation 993cf3aa-df86-4a6e-bf54-8da691da4a8e · outbound

This paper cites Minillm: Knowledge distillation of large language models.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Minillm: Knowledge distillation of large language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:14.923552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:14.923552Z digest=sha256:c89a6d211a11eb61c65c02d4f65671df8278f20f0d005f71c0437f19cfebcbef

Observation 90d96df3-1b9d-42e8-acd5-0898fb8e0d3e · outbound

This paper cites Reasoning with language model is planning with world model.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Reasoning with language model is planning with world model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:14.926673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:14.926673Z digest=sha256:9dcd862991bfee99adad00081cfff3a9392c0f0c6924862353d3869b562e0e50

Observation 25238d1b-0ac7-4b82-92e6-9eec9cdd4dfb · outbound

This paper cites Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.011598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.011598Z digest=sha256:d69de8adbfcca760958e3a465ac73867083f58faa084b1fde4d671c1c609cb93

Observation e5434bfa-165b-4ea1-ad5c-e9c0b6fdfefd · outbound

This paper cites Stable on-policy distillation through adaptive target reformulation.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Stable on-policy distillation through adaptive target reformulation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.044749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.044749Z digest=sha256:5b008e4320280b9fed4ed48a2f2f00c04a9cf23059520b1f0dbe5e143ad3f3d3

Observation cc85a82a-ef9a-4e67-80cd-9eb096c1bdf9 · outbound

This paper cites Rag-rewardbench: Benchmarking reward models in retrieval augmented generation for preference alignment.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Rag-rewardbench: Benchmarking reward models in retrieval augmented generation for preference alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.089974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.089974Z digest=sha256:dc718a6c1219785f39e3340364ec9e018ba2b1e28485d9d90b1c237a07734e32

Observation cfc026a5-6244-44de-8492-e6db46745902 · outbound

This paper cites Scaling reasoning efficiently via relaxed on-policy distillation.arXiv preprint arXiv:2603.11137,.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Scaling reasoning efficiently via relaxed on-policy distillation.arXiv preprint arXiv:2603.11137,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.134752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.134752Z digest=sha256:5897a744af59a870357b84d496320f35b2b0ba36a5af7cd6cfc3516a3088701e

Observation b9d4bf5b-0352-47b0-8310-7fe651227173 · outbound

This paper cites Unifying group-relative and self-distillation policy optimization via sample routing.arXiv preprint arXiv:2604.02288, 2026a.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Unifying group-relative and self-distillation policy optimization via sample routing.arXiv preprint arXiv:2604.02288, 2026a

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.159156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.159156Z digest=sha256:14c26ad7dec53197ad9b6edb1142601654ca6d48305f6e2f542f7c125c241c2b

Observation 5fe9de7e-21af-4fdb-b42a-ee37a8a0e882 · outbound

This paper cites MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.161670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.161670Z digest=sha256:4a41230cc1fc5164741c21ba123a44d7d9323e36986582f7e26885fc810dee22

Observation a8967c0f-f30f-43cd-8ae4-1ac988eed026 · outbound

This paper cites Unlocking the future: Exploring look-ahead planning mechanistic interpretability in large language models.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Unlocking the future: Exploring look-ahead planning mechanistic interpretability in large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.164590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.164590Z digest=sha256:216846914edaf29ed701313b7e73ed159e4c7afcd76b6b3ff5ecfcead7de43d3

Observation 2925f4a0-efba-41b0-9bf5-0385742deb9b · outbound

This paper cites Privileged Information Distillation for Language Models.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Privileged Information Distillation for Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.167530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.167530Z digest=sha256:accbc0b97b715a72335b44a5fd1d1fadef7c985ed10c95cc00fe8c50d99d8be9

Observation f1f6f190-0a52-4073-af01-d3e5c8f4b24a · outbound

This paper cites Pope: Learning to reason on hard problems via privileged on-policy exploration.arXiv preprint arXiv:2601.18779,.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Pope: Learning to reason on hard problems via privileged on-policy exploration.arXiv preprint arXiv:2601.18779,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.170697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.170697Z digest=sha256:af08368289cf2462f094ace47be1c956d59911aac3ee8bf984188bf67d29719b

Observation cfe573fb-9d52-4881-b841-f3ef30f99529 · outbound

This paper cites RL's Razor: Why Online Reinforcement Learning Forgets Less.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation RL's Razor: Why Online Reinforcement Learning Forgets Less

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.173154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.173154Z digest=sha256:ef1697dd5cb994428a8d93a59bc61a2645f1943e6bc1c902af36a977a80f7f0a

Observation f5d29081-45c2-4793-8c14-da75e205a861 · outbound

This paper cites Self-Distillation Enables Continual Learning.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Self-Distillation Enables Continual Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.175710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.175710Z digest=sha256:7154322548170f3f78befe5ea31eea568ad2dba8a9544a9d5baa9af34140ce41

Observation cc8c05bc-dc42-4c9b-82e9-5a241445fbc8 · outbound

This paper cites A Survey of On-Policy Distillation for Large Language Models.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation A Survey of On-Policy Distillation for Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.178315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.178315Z digest=sha256:2caad38435f5f47d826c3b98eb1f13540572a9264ad7819150acabcf92b5c30d

Observation 689d5f84-04fd-4bde-9d77-4818006ac36e · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.181371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.181371Z digest=sha256:810e86a6159b949323b381d03bab62a9b00a809cb02d4b00e9277d58a07d1a90

Observation 41165b66-84c3-480d-a89e-d0edc0d8059d · outbound

This paper cites MiMo-V2-Flash Technical Report.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation MiMo-V2-Flash Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.184145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.184145Z digest=sha256:4e54c3fa4d6ac5e23d21d4dc7248218f022ed27825789c47036541eaa8ca489b

Observation 03cac1b0-4a77-45db-8cf2-a9bc2e5f6294 · outbound

This paper cites Trust Region On-Policy Distillation.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Trust Region On-Policy Distillation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.187338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.187338Z digest=sha256:dc17f325605a752d8ba2b9faf55081902f296eee6bc12a41c1c1047bd2d0493f

Observation c45225d9-c962-4d0b-9068-acaee1725a96 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence.arXiv preprint arXiv:2606.19348,.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Deepseek-v4: Towards highly efficient million-token context intelligence.arXiv preprint arXiv:2606.19348,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.189967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.189967Z digest=sha256:ea04dbc9f3472b4b10812452a5a75d26f784ae527c1c3acf0b079e1996c918c8

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.192803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.192803Z digest=sha256:7ed1d5ecbeda4772c773d2dab9877d2cc2ea3cf338f7306de137a06b7221aca4

Observation cc1a627d-9072-4c93-8cd4-93f42b3694c9 · outbound

This paper cites From𝑓(𝑥) and𝑔(𝑥) to𝑓(𝑔(𝑥)) : Llms learn new skills in rl by composing old ones.arXiv preprint arXiv:2509.25123,.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation From𝑓(𝑥) and𝑔(𝑥) to𝑓(𝑔(𝑥)) : Llms learn new skills in rl by composing old ones.arXiv preprint arXiv:2509.25123,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.195656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.195656Z digest=sha256:259f5f4f0c859b11cfaed70a2e7f956ce4c4b439ad4f2eed1d94967d9faeefaf

Observation 05a40bfd-2784-4590-a07b-851eae744fe3 · outbound

This paper cites OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.198026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.198026Z digest=sha256:c07450d47bc136a21548e952477ed33150d3a5c8513501651bb016945af367d0

Observation 34b05c67-3f4e-4f84-bae7-38312f8180cd · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation GLM-5: from Vibe Coding to Agentic Engineering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.200541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.200541Z digest=sha256:f6aaf43b872443524ec0547e7f1306817aaeffb4c0092e909d7665a6b16873a3

Observation 2cb24b19-1589-4847-8660-84d175fa4b34 · outbound

This paper cites On the interplay of pre-training, mid-training, and rl on reasoning language models.arXiv preprint arXiv:2512.07783, 2025a.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation On the interplay of pre-training, mid-training, and rl on reasoning language models.arXiv preprint arXiv:2512.07783, 2025a

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.203153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.203153Z digest=sha256:02bdcac22154c1f2b512b6a72616fa683082fd4d9739fa88cfede98cb3195777

Observation 4e34d1c5-958c-46d1-ae55-e7b29e03fa3f · outbound

This paper cites EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.205579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.205579Z digest=sha256:98a131bab47561dbaa03b8351807274bfee1ad5f52e96ab844630105b3ea9b2e

Observation 454ac1e0-413e-430d-b634-fa8d45812554 · outbound

This paper cites The row label Category gives the synthetic domain.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation The row label Category gives the synthetic domain

Reference 36

Resolution
malformed identifier
no resolver link, observed 2026-07-31T06:59:15.208890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.208890Z digest=sha256:561d0c32e2f79d2c5314bd1ddd081564d2887bfef1778ad7525335b922c17bff

Observation fe12c3c0-4dd9-4eb4-b3d5-ce5bbf7f5f65 · outbound

This paper cites Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:14.974745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:14.974745Z digest=sha256:ac977c54aec2ee008f1c3eca1d2abecfa5ea1e517d3cd6ef5269e824fafdf137

Observation b1d1a489-2701-4d79-b64b-ec9676dc5494 · outbound

This paper cites Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:14.877055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:14.877055Z digest=sha256:17be419042d25860a6bd8441c78c87f91006343dd6155d4a59b017001cd8d455

Observation ada1fc28-ca80-4cb8-8d4e-407cdebbe162 · outbound

This paper cites Internalizingworldmodelsviaself-playfinetuningforagenticrl.arXivpreprintarXiv:2510.15047,.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Internalizingworldmodelsviaself-playfinetuningforagenticrl.arXivpreprintarXiv:2510.15047,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:14.904753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:14.904753Z digest=sha256:64acf53c4eb2bda7dc6fbe3de1da8978f7df4902030e58f43857cb7061bb3104

Observation 989fb6e4-b1c9-447b-bb7c-d86e6fb4b460 · outbound

This paper cites Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:14.880496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:14.880496Z digest=sha256:7628a153cf59a17dad0ac852ca1ecbbf160d2afbb3e0ba0c3393ae04564fce04

Pith citing papers

No inbound Pith citation observations are available.