Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T08:40:42.872339Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 100 of 300 outbound references and 1 inbound Pith citation observation for arXiv:2607.04763.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T08:40:42.872339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T15:30:01.059028Z
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 300 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9f443ffd-8362-49eb-bfae-22b322c58694 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay 2025 , howpublished =
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1336d3e-9266-4ac5-bbfc-97c85160e992 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Text Embeddings by Weakly-Supervised Contrastive Pre-training
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fc72b79-7772-4889-8a16-0f2edb942c82 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) , pages=
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baed1806-3062-421c-8bc1-68798732de3c · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Proceedings of the 2018 conference on empirical methods in natural language processing , pages=
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ce64a91-a090-4428-a27f-b80b9bbbac3f · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Transactions of the Association for Computational Linguistics , volume=
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94141de7-8be5-4d4a-8104-c3f0c383da82 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Measuring Mathematical Problem Solving With the MATH Dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b86297d8-c780-4a18-bb0c-ab0c0c567073 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Advances in Neural Information Processing Systems , volume=
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cc13f30-31aa-462b-8616-15789eef77d7 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Hugging Face repository , volume=
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef84c5aa-781c-4547-860d-5637c28636db · outbound
Multi-Turn On-Policy Distillation with Prefix Replay ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24806139-c91e-4364-aebe-37f62c053b82 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Qwen3 Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3998160e-d330-41db-a29b-35ab63f05034 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL) , year=
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 676fddcd-e256-4d32-8ec9-d2ae00ecc602 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Proceedings of the 28th International Conference on Computational Linguistics (COLING) , year=
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c0f9ccf-b2c0-442c-b62e-d55546d7bef9 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Transactions of the Association for Computational Linguistics (TACL) , volume=
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a404641d-57cb-4d5d-9a28-6cfaf69c11d0 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL) , year=
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3391810-ad72-4917-9f6c-efdb4428422a · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Findings of the Association for Computational Linguistics: EMNLP 2023 , year=
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68de4c08-32ba-451a-a155-16fbfde9defe · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Langley , title =
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dde49871-d451-4f36-8bd2-8851a462c72f · outbound
Multi-Turn On-Policy Distillation with Prefix Replay 2025 , eprint=
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3ecfb01-0095-44ce-ad21-fd3c930ee143 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa145834-047c-4961-902a-9b90d23d035a · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a544614c-13b1-4c90-bd04-7d41bff4ffd3 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ac50a24-4dd5-4fcb-97a7-15ba90179030 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ce12b81-7b75-45fc-b22a-aba9c048ca31 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17405cd1-3d3b-47b8-8f4a-1a2f7a51319b · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Newell and P
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5688e5f-6f64-4ba4-9d60-1984f0e11167 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay The Lessons of Developing Process Reward Models in Mathematical Reasoning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe31b435-ad89-4af7-8f3c-35f651116117 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay GitHub repository , howpublished =
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1762d531-7f5d-4e71-beba-e3e37b041e23 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay ProcessBench: Identifying Process Errors in Mathematical Reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67f1bd67-2b41-48bc-9aaa-25bbc030fc27 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 752cbdd9-4649-4b5e-82b5-18645ac4b389 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Scaling Learning Algorithms Towards
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a6ea6c5-23bb-4b5c-bc83-bc02b64b8147 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay and Osindero, Simon and Teh, Yee Whye , journal =
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2795fa85-3430-48d4-b071-b3f454479979 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Advances in Neural Information Processing Systems , volume=
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7fa51f1-e716-4864-af9b-10dd0af16192 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Advances in Neural Information Processing Systems , volume=
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9493bcf-bb4b-4c58-91b9-018ed6a6aabe · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Advances in Neural Information Processing Systems , volume=
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2302034-396f-4799-9db9-60a4e84d5d7c · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24b67dc4-17e7-4270-8a97-7b57822cb507 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Generating Sequences by Learning to Self-Correct
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49439340-2569-4632-95ad-47fb2e9ae9a0 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7b4aa52-d7cf-430f-bdad-2f7118defa76 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Recursive Introspection: Teaching Language Model Agents How to Self-Improve
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c81f251-c07b-4059-a2c3-79fcf80c5546 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay RRM: Robust Reward Model Training Mitigates Reward Hacking
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e4255ad-0064-422d-bb92-ed138bd982b2 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Training Language Models to Self-Correct via Reinforcement Learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b300545e-6bc3-4a35-b8d3-9a1585203b2e · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09af1700-6f18-4f2f-ad7e-09b2672d4da5 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Large Language Models Cannot Self-Correct Reasoning Yet
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b049245-6f1f-4954-a740-096908245922 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 337bc30f-ea88-46cc-851c-97563e60a527 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay 2016 , publisher=
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9477601-761a-464d-ba86-7f81ed5390dd · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Disentangling Length from Quality in Direct Preference Optimization
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45b25298-d644-4be9-a4e3-e3cd9ab28394 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15e17f8c-3d1d-4117-a752-b61ea7913795 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Teaching Large Language Models to Reason with Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3636e34-c2e4-4ae9-88f7-2d8b3818795f · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Hashimoto , title =
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43ad6cec-431b-48e7-a216-abfaa752c32c · outbound
Multi-Turn On-Policy Distillation with Prefix Replay 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year=
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a86c0ce-2077-42e8-9bfc-5fe801ea1dec · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d4d236d-1d70-41a1-be45-83fe3dc9b27d · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Know What You Don't Know: Unanswerable Questions for SQuAD
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38ca4a2a-9194-4b5e-b5fd-21168d18833b · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Advances in Neural Information Processing Systems , volume=
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af13d23a-4ffe-4f47-aded-4a9c6e2c9efd · outbound
Multi-Turn On-Policy Distillation with Prefix Replay International Conference on Machine Learning , pages=
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab1836fa-823a-4bf7-b867-5974bfa74b56 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay 2024 , publisher =
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f3c4dd8-ab74-45ef-b9a4-2c089d436537 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay ACM Transactions on Information Systems (TOIS) , volume=
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ea71156-63e4-4bb1-bd8a-be90a8fbe852 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay doi:10.57967/hf/0513 , publisher =
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79dd8db8-7b58-4a8b-b936-fc16176c280b · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cde80cf6-ed03-41cb-a4c6-1dded28bf69c · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Generative Reward Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42f096b6-45c0-42cb-8ad9-9578d5e75310 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay WARM: On the Benefits of Weight Averaged Reward Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06d6db6a-3d7b-41b0-ab22-7d0822e031e6 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1771369-4264-4d3c-9503-6f2f51d0efeb · outbound
Multi-Turn On-Policy Distillation with Prefix Replay DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7d425d7-542b-4f87-8f9f-41bd278ad717 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay BOND: Aligning LLMs with Best-of-N Distillation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ced94b72-b359-4ad9-b517-b3e03c37d5db · outbound
Multi-Turn On-Policy Distillation with Prefix Replay 2023 , eprint=
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 768ec640-82d1-4148-9f71-238efeaac97c · outbound
Multi-Turn On-Policy Distillation with Prefix Replay ODIN: Disentangled Reward Mitigates Hacking in RLHF
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed92659c-72d6-4e18-963b-ab19742f152e · outbound
Multi-Turn On-Policy Distillation with Prefix Replay The Llama 3 Herd of Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc1f96a0-d5b4-447c-bfa8-bc4840830c6a · outbound
Multi-Turn On-Policy Distillation with Prefix Replay 2023 , publisher=
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22d5c9f1-4e25-49d2-8275-3fb8638b6810 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay 2024 , journal =
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3902999d-d6ef-4a8d-8d30-d40e98904309 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay LIMA: Less Is More for Alignment
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88a6371e-ac61-4295-bbac-8c2669572b83 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 350e3b83-e95c-497e-96fa-f4c9f3db9a24 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Advancing LLM Reasoning Generalists with Preference Trees
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ed1519e-8eb9-481f-995a-7df541fbaee7 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 511bdc20-4e3c-465d-951f-fbe531153738 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Iterative Reasoning Preference Optimization
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8458bb9-1226-4e19-842e-bb86044fa609 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Let's Verify Step by Step
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21481048-a15e-45ca-a546-1ff1ed5e919a · outbound
Multi-Turn On-Policy Distillation with Prefix Replay MAmmoTH2: Scaling Instructions from the Web
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 053ca9e2-88b8-4a03-8e0f-f300343c83cb · outbound
Multi-Turn On-Policy Distillation with Prefix Replay REBEL: Reinforcement Learning via Regressing Relative Rewards
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 611ae85f-0e67-41c6-b984-776c7bf72877 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bec1b78-d10d-4bae-a9da-fd714f79ffbc · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Advances in Neural Information Processing Systems , volume=
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9acd4b2-4ebd-42bf-90a9-62837c84ed50 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Step-level Value Preference Optimization for Mathematical Reasoning
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad0225b5-6ede-4c58-8fe0-b39213941d98 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay nature , volume=
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c34ea7d-e742-47ae-855b-9943f89da350 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Playing Atari with Deep Reinforcement Learning
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f322826-12d6-45ff-ba02-ecde46fa59be · outbound
Multi-Turn On-Policy Distillation with Prefix Replay ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e708e9c-c64e-4a44-9328-82b071028f5d · outbound
Multi-Turn On-Policy Distillation with Prefix Replay OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57f4433b-5429-4247-b0f5-941d4c3dde53 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b9b7728-032c-4f5a-808e-391398aebdce · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72982cdf-9a09-4958-a459-8667fbca858e · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Connection Science , volume=
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8286283c-0397-4dbe-83b8-de7a5c2c8f27 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay 2010 , publisher=
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 973cf8ba-f9c7-4386-a43c-032f5b5074b8 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Training Verifiers to Solve Math Word Problems
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4134c048-ba08-401b-99fa-8c8a51f261df · outbound
Multi-Turn On-Policy Distillation with Prefix Replay 2023 , eprint=
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation febdbadf-647c-400e-bf99-eb058a8397f6 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 779f7c7f-44bf-413e-bb64-6a5a6ac6338b · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 910aa229-e412-462f-b930-696349ee7d06 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay LiPO: Listwise Preference Optimization through Learning-to-Rank
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab88c21d-49a8-4d9f-a7dc-eb3da9c9903d · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 942f357b-7f09-40e8-a06f-eb05d914a8f9 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcfab427-7345-4fdc-88e9-a53c74b1df81 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 594a8fd3-7461-4573-91c8-b2fdff369f9e · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0191aac-f91f-4460-b458-110a9358eb47 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay European conference on machine learning , pages=
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c55cb8ae-dc47-45c7-9e27-fe2c1bb069fc · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Self-Consistency Preference Optimization
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52949c34-af80-4431-a8f7-2a010ee7a163 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay 2019 , publisher=
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85614d0d-d12e-484e-976e-8ffddea0d4f4 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b851126-94ad-4101-a683-bfb82f6e21df · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Machine learning , volume=
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e66fc51d-fff0-4293-9736-cf39e2ae36c2 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Advances in neural information processing systems , volume=
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06ae92ae-f1ae-40d8-b836-5f97188dcad0 · outbound
Multi-Turn On-Policy Distillation with Prefix Replay Orca-Math: Unlocking the potential of SLMs in Grade School Math
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09d3c1f3-c066-4967-a71c-eaaffacc1212 · inbound
EvoReason: Self-Evolving Reasoning Primitive-Guided On-Policy Distillation for Latent Reasoning in Generative Recommendation Multi-Turn On-Policy Distillation with Prefix Replay
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.