Pith. sign in

Paper Citation Record · LEDGER

Multi-Turn On-Policy Distillation with Prefix Replay

As of 11 August 2026, this Paper Citation Record lists 100 of 300 outbound references and 1 inbound Pith citation observation for arXiv:2607.04763.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04763 v3

Coverage vector

measured 100 of 300 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:40:42.872339Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T15:30:01.059028Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 300 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved99
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9f443ffd-8362-49eb-bfae-22b322c58694 · outbound

This paper cites 2025 , howpublished =.

Multi-Turn On-Policy Distillation with Prefix Replay 2025 , howpublished =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:30.224369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:30.224369Z digest=sha256:576848ad5c5b7d2b061f715c5a1f7ddc2e0140cf46752ed0b3df682b92a59892

Observation f1336d3e-9266-4ac5-bbfc-97c85160e992 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

Multi-Turn On-Policy Distillation with Prefix Replay Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:30.324022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:30.324022Z digest=sha256:40ae8c8e9dce7948f41fd231b6c255bb8947799635fead3b1be85d4e5868c164

Observation 7fc72b79-7772-4889-8a16-0f2edb942c82 · outbound

This paper cites Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) , pages=.

Multi-Turn On-Policy Distillation with Prefix Replay Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) , pages=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:30.457462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:30.457462Z digest=sha256:e2d9c27d3a01ef5106cb942679b23fa4e3ab5f070e8d01314bc81fef278c2ce2

Observation baed1806-3062-421c-8bc1-68798732de3c · outbound

This paper cites Proceedings of the 2018 conference on empirical methods in natural language processing , pages=.

Multi-Turn On-Policy Distillation with Prefix Replay Proceedings of the 2018 conference on empirical methods in natural language processing , pages=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:30.574860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:30.574860Z digest=sha256:8dbbb30cd87c7f528cd737b559f3dc030cf8983c8871d80a2fd44d113b658ab6

Observation 3ce64a91-a090-4428-a27f-b80b9bbbac3f · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Multi-Turn On-Policy Distillation with Prefix Replay Transactions of the Association for Computational Linguistics , volume=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:30.695664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:30.695664Z digest=sha256:036068703fb240cc25fd79a7f822eac651afcf27107155b915e38e8fb3571d84

Observation 94141de7-8be5-4d4a-8104-c3f0c383da82 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Multi-Turn On-Policy Distillation with Prefix Replay Measuring Mathematical Problem Solving With the MATH Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:30.860607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:30.860607Z digest=sha256:05fcf5c6ee884e7c36f9ad11c438805bc02a76ecb10a7e7dd46672e1ff3da218

Observation b86297d8-c780-4a18-bb0c-ab0c0c567073 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Multi-Turn On-Policy Distillation with Prefix Replay Advances in Neural Information Processing Systems , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:30.987077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:30.987077Z digest=sha256:abb63da77f760c739771db07b75632544db39f5ef0748750673ed56dee081fe0

Observation 1cc13f30-31aa-462b-8616-15789eef77d7 · outbound

This paper cites Hugging Face repository , volume=.

Multi-Turn On-Policy Distillation with Prefix Replay Hugging Face repository , volume=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:31.063677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:31.063677Z digest=sha256:ac67a8679a82ec653ada9f12d8d3e13b4543a1932890daf781d5ea7e58f30915

Observation ef84c5aa-781c-4547-860d-5637c28636db · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

Multi-Turn On-Policy Distillation with Prefix Replay ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:31.207661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:31.207661Z digest=sha256:c661287cc1529110d429856cc0cb451abff46b77b0403fdf10b4a5b29f206904

Observation 24806139-c91e-4364-aebe-37f62c053b82 · outbound

This paper cites Qwen3 Technical Report.

Multi-Turn On-Policy Distillation with Prefix Replay Qwen3 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:31.270884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:31.270884Z digest=sha256:0f195dd1296d3071461d5cbf38bcc7aab1562cd5d5b1090594957dbe0921c1c0

Observation 3998160e-d330-41db-a29b-35ab63f05034 · outbound

This paper cites Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL) , year=.

Multi-Turn On-Policy Distillation with Prefix Replay Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL) , year=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:31.392647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:31.392647Z digest=sha256:fad88d4c3183dc1870826cbc38c722e52f37318a42a7e4de642d005c32efad10

Observation 676fddcd-e256-4d32-8ec9-d2ae00ecc602 · outbound

This paper cites Proceedings of the 28th International Conference on Computational Linguistics (COLING) , year=.

Multi-Turn On-Policy Distillation with Prefix Replay Proceedings of the 28th International Conference on Computational Linguistics (COLING) , year=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:31.522460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:31.522460Z digest=sha256:6300393bf038ec8380db06bfd087b87c202b99fbd6cd83ede0f6330b55fed215

Observation 6c0f9ccf-b2c0-442c-b62e-d55546d7bef9 · outbound

This paper cites Transactions of the Association for Computational Linguistics (TACL) , volume=.

Multi-Turn On-Policy Distillation with Prefix Replay Transactions of the Association for Computational Linguistics (TACL) , volume=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:31.670002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:31.670002Z digest=sha256:717798d6c5dde7dd0a88cfc3d975a8eecd4bb9fb960610107529f43f7377e4ef

Observation a404641d-57cb-4d5d-9a28-6cfaf69c11d0 · outbound

This paper cites Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL) , year=.

Multi-Turn On-Policy Distillation with Prefix Replay Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL) , year=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:31.816310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:31.816310Z digest=sha256:b53b85377b5e4359e5083d465ad8b02fb59000261ded9c329d2178536df42123

Observation c3391810-ad72-4917-9f6c-efdb4428422a · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2023 , year=.

Multi-Turn On-Policy Distillation with Prefix Replay Findings of the Association for Computational Linguistics: EMNLP 2023 , year=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:31.962421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:31.962421Z digest=sha256:107f19f8e06cda5d5e3260ea196db96c414950ae18297a5d0245599405f0b7bc

Observation 68de4c08-32ba-451a-a155-16fbfde9defe · outbound

This paper cites Langley , title =.

Multi-Turn On-Policy Distillation with Prefix Replay Langley , title =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:32.146843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:32.146843Z digest=sha256:18243ffd52be0dd76ac013ca78e8c05ff595a30c88f2bb3e1245c8661ca6d4fc

Observation dde49871-d451-4f36-8bd2-8851a462c72f · outbound

This paper cites 2025 , eprint=.

Multi-Turn On-Policy Distillation with Prefix Replay 2025 , eprint=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:32.281887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:32.281887Z digest=sha256:3aecb9c0f278298216bb8f994ede6805ba2a538bfd40ef26381c5efb3f760b79

Observation b3ecfb01-0095-44ce-ad21-fd3c930ee143 · outbound

This paper cites an unresolved cited work.

Multi-Turn On-Policy Distillation with Prefix Replay Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:32.394207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:32.394207Z digest=sha256:6c48fe71bf5e7575e23eff1d6ead80b90fd5278e54fadf623c9b7f750ce53c8b

Observation aa145834-047c-4961-902a-9b90d23d035a · outbound

This paper cites an unresolved cited work.

Multi-Turn On-Policy Distillation with Prefix Replay Unresolved cited work

Reference 19

Resolution
parse uncertain
no resolver link, observed 2026-08-02T08:40:32.510605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:32.510605Z digest=sha256:a63d26dbf0f7cbd89ce76b5060ff446129fb835442c6131f768d7f29a9ecd637

Observation a544614c-13b1-4c90-bd04-7d41bff4ffd3 · outbound

This paper cites an unresolved cited work.

Multi-Turn On-Policy Distillation with Prefix Replay Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:32.622142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:32.622142Z digest=sha256:f6de8398074ca04b3a6a78c85c17ebc3e3b1aecf05d82435a18ffd12f420c42e

Observation 0ac50a24-4dd5-4fcb-97a7-15ba90179030 · outbound

This paper cites an unresolved cited work.

Multi-Turn On-Policy Distillation with Prefix Replay Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:32.747711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:32.747711Z digest=sha256:dd19c32bf6e7ab5faf2c0f592b43549e8f0c05a46c3cef557afab431512b795d

Observation 6ce12b81-7b75-45fc-b22a-aba9c048ca31 · outbound

This paper cites an unresolved cited work.

Multi-Turn On-Policy Distillation with Prefix Replay Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:32.889724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:32.889724Z digest=sha256:b04a0841a608c7af280a3b9025841800cfd1ef618b146d3c22d698f723e1a563

Observation 17405cd1-3d3b-47b8-8f4a-1a2f7a51319b · outbound

This paper cites Newell and P.

Multi-Turn On-Policy Distillation with Prefix Replay Newell and P

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:32.978498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:32.978498Z digest=sha256:e68bcb17d7b3b37aa062a10e566b4a9a51a3b1f5489ba288e5de1bee0ea9c6ac

Observation a5688e5f-6f64-4ba4-9d60-1984f0e11167 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Multi-Turn On-Policy Distillation with Prefix Replay The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:33.092361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:33.092361Z digest=sha256:ccc26035019192ee5b41183c90c2f052d3ff577ac9cc90b478e181372befa16e

Observation fe31b435-ad89-4af7-8f3c-35f651116117 · outbound

This paper cites GitHub repository , howpublished =.

Multi-Turn On-Policy Distillation with Prefix Replay GitHub repository , howpublished =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:33.219015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:33.219015Z digest=sha256:665b13aabe3f1243005081d6d9b4a0cc68b41c8600f65162132cd0cc16b41e9e

Observation 1762d531-7f5d-4e71-beba-e3e37b041e23 · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Multi-Turn On-Policy Distillation with Prefix Replay ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:33.379008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:33.379008Z digest=sha256:7d075013136ab0da1066837f8891b5b9af0b821a9eaab42cb88d1cea2d730340

Observation 67f1bd67-2b41-48bc-9aaa-25bbc030fc27 · outbound

This paper cites an unresolved cited work.

Multi-Turn On-Policy Distillation with Prefix Replay Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:33.502772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:33.502772Z digest=sha256:b067a04c66177ebfb8f4f5daf0294edc23f9ae64e5fcd550adf3d766e6a5e149

Observation 752cbdd9-4649-4b5e-82b5-18645ac4b389 · outbound

This paper cites Scaling Learning Algorithms Towards.

Multi-Turn On-Policy Distillation with Prefix Replay Scaling Learning Algorithms Towards

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:33.598903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:33.598903Z digest=sha256:b5aa6b2b6feb8c2bf9a0ad2d8c7f531120db4de6476906d703d7867b4d62e27d

Observation 9a6ea6c5-23bb-4b5c-bc83-bc02b64b8147 · outbound

This paper cites and Osindero, Simon and Teh, Yee Whye , journal =.

Multi-Turn On-Policy Distillation with Prefix Replay and Osindero, Simon and Teh, Yee Whye , journal =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:33.758155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:33.758155Z digest=sha256:af3568b6d92cc0455a78259452852f3d890ee079395f41a4eb1f2efaaff0f532

Observation 2795fa85-3430-48d4-b071-b3f454479979 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Multi-Turn On-Policy Distillation with Prefix Replay Advances in Neural Information Processing Systems , volume=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:33.857352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:33.857352Z digest=sha256:e2b59d9d5affd0779cc9941add01f31d8aebea550e885ab3884cfedc960231e0

Observation b7fa51f1-e716-4864-af9b-10dd0af16192 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Multi-Turn On-Policy Distillation with Prefix Replay Advances in Neural Information Processing Systems , volume=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:33.969923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:33.969923Z digest=sha256:955b871c8431eadb2a8bea8581eb3f75243c5fdbd7df42d664a4b538e7917487

Observation c9493bcf-bb4b-4c58-91b9-018ed6a6aabe · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Multi-Turn On-Policy Distillation with Prefix Replay Advances in Neural Information Processing Systems , volume=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:34.175817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:34.175817Z digest=sha256:8392041bf11802d7c43c32949932d35636cc1e479c0579954b6ead9b152efc22

Observation d2302034-396f-4799-9db9-60a4e84d5d7c · outbound

This paper cites Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning.

Multi-Turn On-Policy Distillation with Prefix Replay Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:34.285961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:34.285961Z digest=sha256:23923f13e5d7083a4e5a037e74ede5043906eff6e467a3e41cf3c64aced6bb39

Observation 24b67dc4-17e7-4270-8a97-7b57822cb507 · outbound

This paper cites Generating Sequences by Learning to Self-Correct.

Multi-Turn On-Policy Distillation with Prefix Replay Generating Sequences by Learning to Self-Correct

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:34.391896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:34.391896Z digest=sha256:3450380043b8332914f3a1f021f69d6d31874dcab545e7368596fc2facca13d8

Observation 49439340-2569-4632-95ad-47fb2e9ae9a0 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Multi-Turn On-Policy Distillation with Prefix Replay Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:34.589724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:34.589724Z digest=sha256:e968b89245ff8e76d62f71c6a68ffda56e2694c63e9d7228f9fd64646a0f6807

Observation c7b4aa52-d7cf-430f-bdad-2f7118defa76 · outbound

This paper cites Recursive Introspection: Teaching Language Model Agents How to Self-Improve.

Multi-Turn On-Policy Distillation with Prefix Replay Recursive Introspection: Teaching Language Model Agents How to Self-Improve

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:34.748668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:34.748668Z digest=sha256:6c42a232ad125ba5430d98e22e29aef4af03298707cdc2ec830a9a268f92210a

Observation 3c81f251-c07b-4059-a2c3-79fcf80c5546 · outbound

This paper cites RRM: Robust Reward Model Training Mitigates Reward Hacking.

Multi-Turn On-Policy Distillation with Prefix Replay RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:34.864629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:34.864629Z digest=sha256:5b05edc3ea55919affce2ae9fc83bcbca59316f6c0dc2487960b21e1cfc6c38f

Observation 9e4255ad-0064-422d-bb92-ed138bd982b2 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Multi-Turn On-Policy Distillation with Prefix Replay Training Language Models to Self-Correct via Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:35.009644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:35.009644Z digest=sha256:c7b0cdf73eeb02c00309a9fab418e7cba97c19d7aa95a8d2f616f25b3d60b25b

Observation b300545e-6bc3-4a35-b8d3-9a1585203b2e · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Multi-Turn On-Policy Distillation with Prefix Replay Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:35.127508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:35.127508Z digest=sha256:95c6d53d924b119d8a4165c130268da3e8a7c1127fb8b5d31a15913dff8efbe8

Observation 09af1700-6f18-4f2f-ad7e-09b2672d4da5 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Multi-Turn On-Policy Distillation with Prefix Replay Large Language Models Cannot Self-Correct Reasoning Yet

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:35.247088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:35.247088Z digest=sha256:1283de158fd75b585f0dbf7bf1ebceb66e064def66f1eb2f111d11cfacd06a04

Observation 9b049245-6f1f-4954-a740-096908245922 · outbound

This paper cites Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization.

Multi-Turn On-Policy Distillation with Prefix Replay Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:35.403262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:35.403262Z digest=sha256:7cd8dc43abde4f8f9478c64f160b76b237b9518ae655d1ecec66e952dd8a60fc

Observation 337bc30f-ea88-46cc-851c-97563e60a527 · outbound

This paper cites 2016 , publisher=.

Multi-Turn On-Policy Distillation with Prefix Replay 2016 , publisher=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:35.548203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:35.548203Z digest=sha256:753dd85c1c1c51cf1d76f2e6f02aca6a155d6832f540794bc6deff06a5c75dc5

Observation e9477601-761a-464d-ba86-7f81ed5390dd · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

Multi-Turn On-Policy Distillation with Prefix Replay Disentangling Length from Quality in Direct Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:35.672115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:35.672115Z digest=sha256:062b1a410e238bf827ea0dff3585461ee6be57f58d75319c32520e4be8fda765

Observation 45b25298-d644-4be9-a4e3-e3cd9ab28394 · outbound

This paper cites Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level.

Multi-Turn On-Policy Distillation with Prefix Replay Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:35.809095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:35.809095Z digest=sha256:3765d75b1116bbe0fc782b38bb1c06b516f60abd4385aa31fd5af503eef8fcc7

Observation 15e17f8c-3d1d-4117-a752-b61ea7913795 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

Multi-Turn On-Policy Distillation with Prefix Replay Teaching Large Language Models to Reason with Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:36.004785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:36.004785Z digest=sha256:36f959abbfd89fca2ffc121775f86cb0c370caf576bb8d8f1dd3d4a1c8488500

Observation d3636e34-c2e4-4ae9-88f7-2d8b3818795f · outbound

This paper cites Hashimoto , title =.

Multi-Turn On-Policy Distillation with Prefix Replay Hashimoto , title =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:36.184768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:36.184768Z digest=sha256:b7b9243bd6015176dc45885707d6158efe2c8def439c916f44c480890e8a9354

Observation 43ad6cec-431b-48e7-a216-abfaa752c32c · outbound

This paper cites 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=.

Multi-Turn On-Policy Distillation with Prefix Replay 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:36.356815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:36.356815Z digest=sha256:19e9b00cd1a0c9012337477c4042b0e6154d18b2ee2a564c96c491a9a88aac15

Observation 5a86c0ce-2077-42e8-9bfc-5fe801ea1dec · outbound

This paper cites Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning.

Multi-Turn On-Policy Distillation with Prefix Replay Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:36.488598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:36.488598Z digest=sha256:062717fb5d55f3083391f5156ed220714084d67d3d88b782b9a3cdfd18636490

Observation 4d4d236d-1d70-41a1-be45-83fe3dc9b27d · outbound

This paper cites Know What You Don't Know: Unanswerable Questions for SQuAD.

Multi-Turn On-Policy Distillation with Prefix Replay Know What You Don't Know: Unanswerable Questions for SQuAD

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:36.627944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:36.627944Z digest=sha256:85163d3bb4731611172492dca6a4d2d035059485ae6334f01296dda994bddf8c

Observation 38ca4a2a-9194-4b5e-b5fd-21168d18833b · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Multi-Turn On-Policy Distillation with Prefix Replay Advances in Neural Information Processing Systems , volume=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:36.784779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:36.784779Z digest=sha256:bd04e0f5743db6fc25d5cae67df3919e8d8a55bbc25c5b718b3b16c81135da73

Observation af13d23a-4ffe-4f47-aded-4a9c6e2c9efd · outbound

This paper cites International Conference on Machine Learning , pages=.

Multi-Turn On-Policy Distillation with Prefix Replay International Conference on Machine Learning , pages=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:36.911975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:36.911975Z digest=sha256:0f8f0f53aac695a90191046e6178264215a879281aedaed25da4db0f8871de6d

Observation ab1836fa-823a-4bf7-b867-5974bfa74b56 · outbound

This paper cites 2024 , publisher =.

Multi-Turn On-Policy Distillation with Prefix Replay 2024 , publisher =

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:37.091342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:37.091342Z digest=sha256:527cb4b516d856a95a5ecc088d0fbd97ad427ea16e3627125882fa1cfaeb26f4

Observation 5f3c4dd8-ab74-45ef-b9a4-2c089d436537 · outbound

This paper cites ACM Transactions on Information Systems (TOIS) , volume=.

Multi-Turn On-Policy Distillation with Prefix Replay ACM Transactions on Information Systems (TOIS) , volume=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:37.204832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:37.204832Z digest=sha256:82810d4a520caa7e155e81b1d6eebba0b99acdb8a74bb72f50fcddca0ae6a74b

Observation 5ea71156-63e4-4bb1-bd8a-be90a8fbe852 · outbound

This paper cites doi:10.57967/hf/0513 , publisher =.

Multi-Turn On-Policy Distillation with Prefix Replay doi:10.57967/hf/0513 , publisher =

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:37.285743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:37.285743Z digest=sha256:eedb5dddfdb30078ff992ce1378039f3265cd30e5af93682f5f0501d3e85d203

Observation 79dd8db8-7b58-4a8b-b936-fc16176c280b · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Multi-Turn On-Policy Distillation with Prefix Replay Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:37.369447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:37.369447Z digest=sha256:3e185f3ca63ccf2515b8c68cd6606622effafc25e3c1a4b9329e1ba0e7fa7d14

Observation cde80cf6-ed03-41cb-a4c6-1dded28bf69c · outbound

This paper cites Generative Reward Models.

Multi-Turn On-Policy Distillation with Prefix Replay Generative Reward Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:37.485022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:37.485022Z digest=sha256:a22ab618155cce915f172873706e95f81591c503d866f21fb85f565bb1deaab7

Observation 42f096b6-45c0-42cb-8ad9-9578d5e75310 · outbound

This paper cites WARM: On the Benefits of Weight Averaged Reward Models.

Multi-Turn On-Policy Distillation with Prefix Replay WARM: On the Benefits of Weight Averaged Reward Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:37.580679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:37.580679Z digest=sha256:a2b9bc76df411a09f5d92ee667dc47211867bd583302ecca992cb5ca390abcb3

Observation 06d6db6a-3d7b-41b0-ab22-7d0822e031e6 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Multi-Turn On-Policy Distillation with Prefix Replay Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:37.766104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:37.766104Z digest=sha256:7533093c138a3ba8f233242da4fd3ecd714708bff48769d297d9fa7f63103d7e

Observation d1771369-4264-4d3c-9503-6f2f51d0efeb · outbound

This paper cites DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving.

Multi-Turn On-Policy Distillation with Prefix Replay DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:37.931798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:37.931798Z digest=sha256:14966f68cfc49db1a2358b82deb3bac5b5271a892b0b9a4eb56fc3dc3c5487ad

Observation d7d425d7-542b-4f87-8f9f-41bd278ad717 · outbound

This paper cites BOND: Aligning LLMs with Best-of-N Distillation.

Multi-Turn On-Policy Distillation with Prefix Replay BOND: Aligning LLMs with Best-of-N Distillation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:38.041034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:38.041034Z digest=sha256:c3c52cdd8790bac2ba43b850f8ca119390d5bb5034b49c10a7bd09ed75219e4c

Observation ced94b72-b359-4ad9-b517-b3e03c37d5db · outbound

This paper cites 2023 , eprint=.

Multi-Turn On-Policy Distillation with Prefix Replay 2023 , eprint=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:38.173978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:38.173978Z digest=sha256:aadf7dd36c27ea70ab474a19bc79ae3ea721fe84129018d66433e008a5c8e74d

Observation 768ec640-82d1-4148-9f71-238efeaac97c · outbound

This paper cites ODIN: Disentangled Reward Mitigates Hacking in RLHF.

Multi-Turn On-Policy Distillation with Prefix Replay ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:38.315887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:38.315887Z digest=sha256:2b14ac478b827d8f88f936ca11af40232b39c7497120197bf21abef555e5f736

Observation ed92659c-72d6-4e18-963b-ab19742f152e · outbound

This paper cites The Llama 3 Herd of Models.

Multi-Turn On-Policy Distillation with Prefix Replay The Llama 3 Herd of Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:38.483807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:38.483807Z digest=sha256:79c35ede00381b7ddee7782cda0f13fc96b2216965e67e1b36d75b99e38fb751

Observation bc1f96a0-d5b4-447c-bfa8-bc4840830c6a · outbound

This paper cites 2023 , publisher=.

Multi-Turn On-Policy Distillation with Prefix Replay 2023 , publisher=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:38.588624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:38.588624Z digest=sha256:dcfd6587a689bcba21563d7bfaaab963c690ce482e9e1ab2a23b030e72b13380

Observation 22d5c9f1-4e25-49d2-8275-3fb8638b6810 · outbound

This paper cites 2024 , journal =.

Multi-Turn On-Policy Distillation with Prefix Replay 2024 , journal =

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:38.643258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:38.643258Z digest=sha256:0a774482c60bec339de778c61d6be041dda58342cbe910750569c52c76b22e8b

Observation 3902999d-d6ef-4a8d-8d30-d40e98904309 · outbound

This paper cites LIMA: Less Is More for Alignment.

Multi-Turn On-Policy Distillation with Prefix Replay LIMA: Less Is More for Alignment

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:38.769023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:38.769023Z digest=sha256:4b2a7dc43f67c28d88bd3ffa8f38c4cb07e8691fa0a7fe3e1c6f9613be63c2fd

Observation 88a6371e-ac61-4295-bbac-8c2669572b83 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

Multi-Turn On-Policy Distillation with Prefix Replay ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:38.941759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:38.941759Z digest=sha256:9e4482f0b5f389fef8b1a2885ac0703268ff002ff089952cca6809c168872893

Observation 350e3b83-e95c-497e-96fa-f4c9f3db9a24 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

Multi-Turn On-Policy Distillation with Prefix Replay Advancing LLM Reasoning Generalists with Preference Trees

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.023948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.023948Z digest=sha256:22fc84c796c3ca8f612fff945f15d12c6f3302a8406723b11cc9dd347e051f60

Observation 7ed1519e-8eb9-481f-995a-7df541fbaee7 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Multi-Turn On-Policy Distillation with Prefix Replay WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.153039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.153039Z digest=sha256:e844e69150610b55c81f94e9b51e6aaa357b0147ac60a93ab2186f074cc1a76c

Observation 511bdc20-4e3c-465d-951f-fbe531153738 · outbound

This paper cites Iterative Reasoning Preference Optimization.

Multi-Turn On-Policy Distillation with Prefix Replay Iterative Reasoning Preference Optimization

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.278322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.278322Z digest=sha256:af3c554e933429bfc6b7e38f520fa580ad8479a58c5af03699bab1cc0f4a35c8

Observation b8458bb9-1226-4e19-842e-bb86044fa609 · outbound

This paper cites Let's Verify Step by Step.

Multi-Turn On-Policy Distillation with Prefix Replay Let's Verify Step by Step

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.434503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.434503Z digest=sha256:8a0fdc9656aae507023822ef856c093b462a6ad33811b931b8e7d5c1fc8ae2dc

Observation 21481048-a15e-45ca-a546-1ff1ed5e919a · outbound

This paper cites MAmmoTH2: Scaling Instructions from the Web.

Multi-Turn On-Policy Distillation with Prefix Replay MAmmoTH2: Scaling Instructions from the Web

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.565899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.565899Z digest=sha256:f1c63389f8c989a114176e9aed0e82dff9c38d865363e5c1706743f4e6f3f35c

Observation 053ca9e2-88b8-4a03-8e0f-f300343c83cb · outbound

This paper cites REBEL: Reinforcement Learning via Regressing Relative Rewards.

Multi-Turn On-Policy Distillation with Prefix Replay REBEL: Reinforcement Learning via Regressing Relative Rewards

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.691665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.691665Z digest=sha256:7ea9fb639173a7f3b1c21e4b939c49d852ee5b790211b427243f324f6993fffe

Observation 611ae85f-0e67-41c6-b984-776c7bf72877 · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

Multi-Turn On-Policy Distillation with Prefix Replay Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.787640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.787640Z digest=sha256:4836b06c3d7c596695c67b5efd6cdea44f8d66cee8f765a7c3624771c33a40ec

Observation 7bec1b78-d10d-4bae-a9da-fd714f79ffbc · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Multi-Turn On-Policy Distillation with Prefix Replay Advances in Neural Information Processing Systems , volume=

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.909147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.909147Z digest=sha256:951f2dd48377e27c15370d7e7517c0781f02f1a99224a948eec20e6907c5996f

Observation a9acd4b2-4ebd-42bf-90a9-62837c84ed50 · outbound

This paper cites Step-level Value Preference Optimization for Mathematical Reasoning.

Multi-Turn On-Policy Distillation with Prefix Replay Step-level Value Preference Optimization for Mathematical Reasoning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.974766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.974766Z digest=sha256:3a14e8e1eb5a95e1b0c5e37f2f2528aa8a7b8f454d74ebd0a85bfcfcbaf924ca

Observation ad0225b5-6ede-4c58-8fe0-b39213941d98 · outbound

This paper cites nature , volume=.

Multi-Turn On-Policy Distillation with Prefix Replay nature , volume=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:40.055037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:40.055037Z digest=sha256:cb6dcfdda13d7cb7259e68b23a5a9698bd084f9b538cc53a859829db8a09e5ce

Observation 1c34ea7d-e742-47ae-855b-9943f89da350 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Multi-Turn On-Policy Distillation with Prefix Replay Playing Atari with Deep Reinforcement Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:40.132166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:40.132166Z digest=sha256:ee6798ea0b969ea21fe1a268048795b482f852d1795b88126c2372f056b817e3

Observation 6f322826-12d6-45ff-ba02-ecde46fa59be · outbound

This paper cites ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent.

Multi-Turn On-Policy Distillation with Prefix Replay ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:40.326974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:40.326974Z digest=sha256:1ba6741ded7fd1852efc1a7d4c387dd1a35b019a795ac7325100173e33da5aab

Observation 1e708e9c-c64e-4a44-9328-82b071028f5d · outbound

This paper cites OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset.

Multi-Turn On-Policy Distillation with Prefix Replay OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:40.446065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:40.446065Z digest=sha256:c8688f4a0be35450d29b3e5cd5ae4867449aa0263659a605c521dd107eb96f3b

Observation 57f4433b-5429-4247-b0f5-941d4c3dde53 · outbound

This paper cites OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data.

Multi-Turn On-Policy Distillation with Prefix Replay OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:40.504438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:40.504438Z digest=sha256:ddf54ac378b66401e3a04d7fe14cb3e669afaaea11659a09df0b9a2102865dee

Observation 1b9b7728-032c-4f5a-808e-391398aebdce · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Multi-Turn On-Policy Distillation with Prefix Replay Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:40.571083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:40.571083Z digest=sha256:028278d3df83008dc481cc0d2a470c7f1248f3c1746cb5f0cf0689408051bf1f

Observation 72982cdf-9a09-4958-a459-8667fbca858e · outbound

This paper cites Connection Science , volume=.

Multi-Turn On-Policy Distillation with Prefix Replay Connection Science , volume=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:40.667991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:40.667991Z digest=sha256:63a03bc453024133cbf08b0a1413c2880a295f3083b2e9a51d87e97f5d0614e8

Observation 8286283c-0397-4dbe-83b8-de7a5c2c8f27 · outbound

This paper cites 2010 , publisher=.

Multi-Turn On-Policy Distillation with Prefix Replay 2010 , publisher=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:40.706168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:40.706168Z digest=sha256:433db384b6cdf914c88c7021180ecebae7ce28c99788039504c025f80fc7f8c0

Observation 973cf8ba-f9c7-4386-a43c-032f5b5074b8 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Multi-Turn On-Policy Distillation with Prefix Replay Training Verifiers to Solve Math Word Problems

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:40.786290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:40.786290Z digest=sha256:d32eba04208e997b50ef0c9de22ff6e1c5fd5a099ca8661cac6265c112cabd32

Observation 4134c048-ba08-401b-99fa-8c8a51f261df · outbound

This paper cites 2023 , eprint=.

Multi-Turn On-Policy Distillation with Prefix Replay 2023 , eprint=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:40.885199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:40.885199Z digest=sha256:c10620a9ee7d8771eb3f72b85b51ccbe2f4641b280d106da620c8304c0067641

Observation febdbadf-647c-400e-bf99-eb058a8397f6 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Multi-Turn On-Policy Distillation with Prefix Replay From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:40.977636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:40.977636Z digest=sha256:e852ac480164f4fc3768f99299a5268f25f141313642120eb353c94e765e2062

Observation 779f7c7f-44bf-413e-bb64-6a5a6ac6338b · outbound

This paper cites Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer.

Multi-Turn On-Policy Distillation with Prefix Replay Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:41.047599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:41.047599Z digest=sha256:424564f03f419d2dfa577a2d8672d547e6e00ef7966f55054ac04cf248f763f3

Observation 910aa229-e412-462f-b930-696349ee7d06 · outbound

This paper cites LiPO: Listwise Preference Optimization through Learning-to-Rank.

Multi-Turn On-Policy Distillation with Prefix Replay LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:41.165786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:41.165786Z digest=sha256:c1c1caaaa5d546a5c1aa142ce96c6078638c4909e5131e7bcc6aea15d850fa94

Observation ab88c21d-49a8-4d9f-a7dc-eb3da9c9903d · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Multi-Turn On-Policy Distillation with Prefix Replay Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:41.352028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:41.352028Z digest=sha256:9947adb74baacb449000e079165fe29ab3d39e92cff1d56676fc04f7846d99f7

Observation 942f357b-7f09-40e8-a06f-eb05d914a8f9 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Multi-Turn On-Policy Distillation with Prefix Replay Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:41.531173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:41.531173Z digest=sha256:d78bedcaf8d0b92952386c5550058674fd706f957257ae0e239e7b5b7a440d8a

Observation fcfab427-7345-4fdc-88e9-a53c74b1df81 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Multi-Turn On-Policy Distillation with Prefix Replay Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:41.705346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:41.705346Z digest=sha256:2a75e0825db3cf362a1c9b52c6c12f50624a6022ad2d28af7c4256e1ab86ad01

Observation 594a8fd3-7461-4573-91c8-b2fdff369f9e · outbound

This paper cites Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing.

Multi-Turn On-Policy Distillation with Prefix Replay Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:41.861888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:41.861888Z digest=sha256:00b6aa3ad3925c576b74d542e75e152dbc5826c2ca3875063ad81689098ec540

Observation a0191aac-f91f-4460-b458-110a9358eb47 · outbound

This paper cites European conference on machine learning , pages=.

Multi-Turn On-Policy Distillation with Prefix Replay European conference on machine learning , pages=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:42.030702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:42.030702Z digest=sha256:e693e5ea52eac1560376d957c5d7eae96b62fa306f6a3b6424f745e4cb07a1ff

Observation c55cb8ae-dc47-45c7-9e27-fe2c1bb069fc · outbound

This paper cites Self-Consistency Preference Optimization.

Multi-Turn On-Policy Distillation with Prefix Replay Self-Consistency Preference Optimization

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:42.168233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:42.168233Z digest=sha256:32bfb0161f953f9c788f35439af9d07a904ad4d460a2110ad84bf23ac91ae0f9

Observation 52949c34-af80-4431-a8f7-2a010ee7a163 · outbound

This paper cites 2019 , publisher=.

Multi-Turn On-Policy Distillation with Prefix Replay 2019 , publisher=

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:42.312787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:42.312787Z digest=sha256:3a6d03337fb0947c884c8344d28c5fb35f35068d311dd9e7fd146d148ce176b0

Observation 85614d0d-d12e-484e-976e-8ffddea0d4f4 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Multi-Turn On-Policy Distillation with Prefix Replay Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:42.479453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:42.479453Z digest=sha256:df0dce87abb8661794b12c1a7640f38500dc3f3e0e418995ded29d6a690517d7

Observation 6b851126-94ad-4101-a683-bfb82f6e21df · outbound

This paper cites Machine learning , volume=.

Multi-Turn On-Policy Distillation with Prefix Replay Machine learning , volume=

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:42.599605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:42.599605Z digest=sha256:5cb0e4ddfc58df78a71896ac7d3c5763364c91ae1034ad868a6bed77630ffb76

Observation e66fc51d-fff0-4293-9736-cf39e2ae36c2 · outbound

This paper cites Advances in neural information processing systems , volume=.

Multi-Turn On-Policy Distillation with Prefix Replay Advances in neural information processing systems , volume=

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:42.741707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:42.741707Z digest=sha256:852a6116c066ba27c2fd26b4ea4588afae25f39695a329fdc23a6b2f6f891181

Observation 06ae92ae-f1ae-40d8-b836-5f97188dcad0 · outbound

This paper cites Orca-Math: Unlocking the potential of SLMs in Grade School Math.

Multi-Turn On-Policy Distillation with Prefix Replay Orca-Math: Unlocking the potential of SLMs in Grade School Math

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:42.872339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:42.872339Z digest=sha256:96a4c47faa1e831f8239a8769197b0f56265f9f65355d9a8248ca4aaa7e29d43

Pith citing papers

Observation 09d3c1f3-c066-4967-a71c-eaaffacc1212 · inbound

EvoReason: Self-Evolving Reasoning Primitive-Guided On-Policy Distillation for Latent Reasoning in Generative Recommendation cites this paper.

EvoReason: Self-Evolving Reasoning Primitive-Guided On-Policy Distillation for Latent Reasoning in Generative Recommendation Multi-Turn On-Policy Distillation with Prefix Replay

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T15:30:01.059028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:30:01.059028Z digest=sha256:13282f9b4d2f0100882e1eff5679e3b08ca0c5e32d72a32bb67522b654d27e0e