Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T20:06:18.832821Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 0 inbound Pith citation observations for arXiv:2607.28026.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T20:06:18.832821Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
99 of 99 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a57c262b-3c68-486a-9bde-0585c68292be · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52e5a969-9f83-4658-a4da-67d500890c02 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1d53203-cac8-45b2-86a7-edbd7fbc4c42 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 079aaf0e-e246-4db6-bf08-ea3eb1e5d7f2 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 032d3af7-9043-4df6-94e3-156e805cafde · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdaf69e1-27d4-406c-9c1e-d4b9f591388e · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Agentic Reinforced Policy Optimization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8f6096c-d3e5-408f-bc1e-5e230eb3c674 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Rubric-based On-policy Distillation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 401ce279-e4e8-4778-9d94-1d0ef45a49be · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2ab60e4-2204-4626-a93f-86058ab74c34 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2495cca-22a0-49d6-8930-800e7c4a11a7 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation D.; Sugawara, S.; and Aizawa, A
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8002fa03-45d5-4f2f-939f-4add6acc0de7 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62a27e60-91f8-4996-b4a6-8d16fdcd8b86 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Reinforcement Learning via Self-Distillation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae311834-9899-4aa2-a9bc-0655e4d0be03 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cc5831b-2012-4f58-8e88-b6bfa626a414 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8e9cb95-6517-4128-a56a-3e4e6a7bb866 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation CoRT: Code-integrated Reasoning within Thinking
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed3bda88-b286-4d89-a83b-fb5cbb5a78a3 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation WebSailor: Navigating Super-human Reasoning for Web Agent
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 678c0b74-c1b6-41ef-88e5-a4d5f5c207a5 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93e703b5-485e-4aea-812a-07769b483e2d · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eb363cf-e8b6-45dc-8dd4-30cba790929d · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab20a02f-1cb9-419a-a9b7-c57271aa8c02 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e7f2b59-2a78-424b-be5d-41d9ff2012af · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Self-Distilled Agentic Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e82e2748-b1b2-4c7d-8af1-c7965ec7ffd1 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb1f136a-c96c-44d0-988e-ba3c11983bbc · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9de57c66-ddfb-431a-a362-d76b9ef0e045 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ada30bc4-d367-4783-84d5-4505a4ff3a49 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0963b69d-513b-4722-9f53-b6e98d6d77b2 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Humanity's Last Exam
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9fa9b8f-0a71-4404-bfa8-50257e2d7f2b · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation A.; and Lewis, M
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4092235b-ccae-4176-85bf-b185cc3f057f · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation ToolRL: Reward is All Tool Learning Needs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a1e834-d20b-47b2-8a3c-e1d6e14a4129 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a661eac-8563-4f2a-a7ef-5e93bd6d0096 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Qwen2.5 Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91871e85-3f73-4f91-a227-2b77ec7785f3 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Trust Region Policy Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62bbf2f3-e045-4d4e-8366-0fc3a5a9b147 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72b3b020-c5c2-47dd-a3f4-1e6b8b5d2a7e · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dab872a8-80f4-47d3-bcc0-6e660dfdf5c8 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation A Survey of On-Policy Distillation for Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 811a75d3-c651-446d-a832-ee15c4909a8d · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation X.; Liu, Z.; Fang, L.; Wang, Z.; and Wen, J.-R
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 018541f5-f4e5-452d-952a-a769dc764789 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e131f767-25f6-42a7-932d-9ec8ed7670e5 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Qwen3 Technical Report
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8624e6d0-bb2b-480e-895f-50db18ebc1c5 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5e99fc8-d4e3-458f-8d85-93fc80964da7 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b21d4aa-800b-41f8-9efe-4f2e9bf4e5a6 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Acting Less is Reasoning More! Teaching Model to Act Efficiently
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8602b83e-74a1-4df6-adf4-69a7ced1361d · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3e667ac-d082-4c4b-b015-1688e38ef6e1 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c68a7ed-1f71-49fa-b8e9-d74a21fbd5aa · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation WebWalker: Benchmarking LLMs in Web Traversal
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 846cd24f-ac38-48f1-8360-f23f935f241b · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Self-Distilled RLVR
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 550d5251-4169-4b22-9ba6-d06c224b8420 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25ca512c-9c07-46af-ba54-6d7c2d508b3b · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88a9b1c6-d37d-4c86-8a4f-7534b224ec5a · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44cd3696-1fcd-4a30-86eb-60f4ca868a4a · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59d94601-e7fe-4ece-8812-da39d05c3d6a · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d86c6bd7-1c4e-4495-bdfa-6b80e3d6fe9b · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2019 , eprint=
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8608f900-1a77-4b0d-b0db-4c1925a5e8e6 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2017 , eprint=
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d395184-cc6c-49d2-9bd7-04652f2b188c · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44a8c549-a6a1-499d-a19f-0e7be510b8ba · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Neural computation , volume=
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3332816d-4ffc-4dfd-b74a-ee3d6859d39f · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fe8b07f-b698-4f43-b9be-3220b8175815 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac0b075e-f6f5-4a60-98ee-80c00cfee5d2 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e09070b-5bb1-44e6-ac6e-36e3ff0e8031 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b371e495-e319-4735-bc9a-64b54857c711 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 543e41ae-722a-4286-bf4f-673b48dd2459 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12c26a91-dfee-44a8-8300-8892d7bd539b · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1743c439-9a8d-4aa0-a6b2-86ae8b402a13 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Advances in Neural Information Processing Systems , volume=
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0f59ccf-2626-4a01-b237-66fb8e5e2152 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cb86448-61a3-4dab-9654-2f6fa6622c70 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e6b0060-32ca-450a-ad26-ed13e820aa18 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation International conference on machine learning , pages=
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e012e7c-5b89-4f09-8281-fcd6c9086c2d · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16202c2f-256f-4467-af26-5fe53b79dd4d · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a41fa58-5dbc-451f-944c-d9fd867cd7e4 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bd9eeee-7a96-43ad-b2b9-ea856e3848a8 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 054422f6-b294-4362-91c1-c4782176615d · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce9cee68-264c-40f9-956e-2a305d5ed643 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , publisher=
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5d52a3a-126f-4b35-89aa-50842525c4a2 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2024 , eprint=
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5061c51f-b42d-4e67-a605-111aef5eb10d · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2024 , eprint=
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d18457c-9127-4333-9e0b-d59a3fcef5b3 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 849d2471-d40d-4995-a925-01da668dae6c · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b890c80-d7df-4a71-a5e3-2eb1bbbda0eb · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ebbab12-40db-4e2d-9e42-fbcbbfaf0d50 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3f64341-ca73-4245-ab5d-8bf92be0bd53 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 131a81a1-2ec6-4d91-9692-897bedef8767 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97aa4d34-18bf-484b-96ae-cd0a026c30df · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28dd0634-a439-4a5b-ad1e-3d9850fe217e · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eeb9cd58-0ba9-44e5-97a3-6401d897ea80 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation The Twelfth International Conference on Learning Representations,
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ee5a9a5-131e-4bba-93ad-230fc8fb2c91 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Measuring Mathematical Problem Solving With the
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 263eb880-823c-416d-8b72-850373b39ef7 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation WebWalker: Benchmarking LLMs in Web Traversal
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4807d86d-2c0d-42c5-babe-8f7a14350a00 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation H otpot QA : A Dataset for Diverse, Explainable Multi-hop Question Answering
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b670483-abdb-4cea-939a-bb0cf1118122 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Constructing
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b17e206-de3f-45bd-a77a-f48fb3f1a918 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Transactions of the Association for Computational Linguistics , volume=
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9293f171-e8e9-4a4b-ba3f-b906d1218c37 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Smith and Mike Lewis , editor =
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 890ed648-0450-47f9-8569-e41e3502dfbe · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation The Twelfth International Conference on Learning Representations,
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9ad37d2-76ef-456d-b4bb-305020f4f8d5 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Humanity's Last Exam
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9dd0a5b-95db-457e-a828-5b1ddcc9d7e0 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa5ab322-c608-4204-a62f-46fe6816f7df · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6ccbbc4-bf13-4618-b828-bd37700535d2 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ada9f8a-0811-4840-b194-490a576432e4 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2024 , eprint=
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0641966a-50f4-4f8c-af2d-8acff291ea1f · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation The Llama 3 Herd of Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87f63080-8ba1-4606-bc5d-7bcb9447781c · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Qwen3 Technical Report
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ddbcd7c-9a1b-45ab-b37f-798aa9b91dc1 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a017e6ba-202c-4485-93bb-6e5508be01ba · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation CoRT: Code-integrated Reasoning within Thinking
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3517e788-ab76-46b9-9ce7-f6e87a9da2d9 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) , address=
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a783ed2-84d5-4b31-9c67-0cbccbdaa270 · outbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.