Pith. sign in

Paper Citation Record · LEDGER

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation

As of 18 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 0 inbound Pith citation observations for arXiv:2607.28026.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28026 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T20:06:18.832821Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

99 of 99 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved99
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a57c262b-3c68-486a-9bde-0585c68292be · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:09.660217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:09.660217Z digest=sha256:2003958148671f83556196db232b41e34ff3fa8c291178d1ecc556da77867b2e

Observation 52e5a969-9f83-4658-a4da-67d500890c02 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:09.797967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:09.797967Z digest=sha256:dd385135077a18e9fd8804e58cefadae67938cb91c43f4fa2d667676a38b2040

Observation c1d53203-cac8-45b2-86a7-edbd7fbc4c42 · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:09.948016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:09.948016Z digest=sha256:a426e1cfd0a23062ff9272afffbf7ad2c536f9789914615d7c3e64c57a0558e1

Observation 079aaf0e-e246-4db6-bf08-ea3eb1e5d7f2 · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:10.099044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:10.099044Z digest=sha256:54534a459ffe3efee2d40fffee14f0d8cd8457699c14e8efda61ced19c7bd4c0

Observation 032d3af7-9043-4df6-94e3-156e805cafde · outbound

This paper cites Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:10.315097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:10.315097Z digest=sha256:e794f95c1b8814aa07155d9f9861e8f1031440e1b579e6eef90cc0cf6a059a89

Observation cdaf69e1-27d4-406c-9c1e-d4b9f591388e · outbound

This paper cites Agentic Reinforced Policy Optimization.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Agentic Reinforced Policy Optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:10.418692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:10.418692Z digest=sha256:fcf9c5776f76471a44fcfd81c57393d8db0416a6458e0b1c71524055bdd8af98

Observation e8f6096c-d3e5-408f-bc1e-5e230eb3c674 · outbound

This paper cites Rubric-based On-policy Distillation.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Rubric-based On-policy Distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:10.692173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:10.692173Z digest=sha256:3b591d2ef3fa4736b8eb887354edb8391038f11070a3606322bfc511da297daf

Observation 401ce279-e4e8-4778-9d94-1d0ef45a49be · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:10.812527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:10.812527Z digest=sha256:58596435a4f2ccae2b04de32d2d5935773227b86d39027eb068cebfa37fa160a

Observation a2ab60e4-2204-4626-a93f-86058ab74c34 · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:10.902305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:10.902305Z digest=sha256:cb252fd112af470533f653c082d09e753a91bceacf86224636b8341d56a8a84d

Observation a2495cca-22a0-49d6-8930-800e7c4a11a7 · outbound

This paper cites D.; Sugawara, S.; and Aizawa, A.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation D.; Sugawara, S.; and Aizawa, A

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:11.010362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:11.010362Z digest=sha256:a0b82894acada9e717ab5b612d5af7ddbab1be938b04befbd95b5731d038fb66

Observation 8002fa03-45d5-4f2f-939f-4add6acc0de7 · outbound

This paper cites Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:11.131377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:11.131377Z digest=sha256:acb836f169f9f506e698bf817569df337090fc163d5e2e9a4c617a61e205b7b7

Observation 62a27e60-91f8-4996-b4a6-8d16fdcd8b86 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Reinforcement Learning via Self-Distillation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:11.255501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:11.255501Z digest=sha256:60abf243f812de5d4f30153f5ad3c046ba6c944c3b830f9afe83cf49373e9c7d

Observation ae311834-9899-4aa2-a9bc-0655e4d0be03 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:11.327263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:11.327263Z digest=sha256:f64289ef53e3d24af5bf905fbeacb0dab69a89a94f1f69f85697795965c7913e

Observation 6cc5831b-2012-4f58-8e88-b6bfa626a414 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:11.452137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:11.452137Z digest=sha256:88fcd1ea8687568ca3a7c34fc80d92797c19a42d4fc84fc834b6b29a7b06384e

Observation d8e9cb95-6517-4128-a56a-3e4e6a7bb866 · outbound

This paper cites CoRT: Code-integrated Reasoning within Thinking.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation CoRT: Code-integrated Reasoning within Thinking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:11.549348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:11.549348Z digest=sha256:2c0bd1ebb2d91829007d2ffc31b095c1a8bcb7af24d646718dfa97e514f9e655

Observation ed3bda88-b286-4d89-a83b-fb5cbb5a78a3 · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:11.655729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:11.655729Z digest=sha256:58fac6bef452c22f598dfa7ee69a6b58d4cf73340c746db02d642d2b449159c4

Observation 678c0b74-c1b6-41ef-88e5-a4d5f5c207a5 · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:11.775197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:11.775197Z digest=sha256:52fae84ec73d786ffcdb27446fecdf6a291a409b4a889378abca055cd59a0053

Observation 93e703b5-485e-4aea-812a-07769b483e2d · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:11.846074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:11.846074Z digest=sha256:849c4f7133d24fe9dcc6dbe0fe60e28d7bf4f928733bdbd1a0ab385c44d521ff

Observation 9eb363cf-e8b6-45dc-8dd4-30cba790929d · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:11.916136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:11.916136Z digest=sha256:dc59ffec4e1ef3c530ee556816d26397991b77f468376c22e3bf099b319f36c3

Observation ab20a02f-1cb9-419a-a9b7-c57271aa8c02 · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:11.982246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:11.982246Z digest=sha256:1c6852951985812e625ea6f336e6bc486a772fbe547d91a580c08021c7f3c806

Observation 9e7f2b59-2a78-424b-be5d-41d9ff2012af · outbound

This paper cites Self-Distilled Agentic Reinforcement Learning.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Self-Distilled Agentic Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:12.095360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:12.095360Z digest=sha256:a1114b3fe7c2eb18852d7eddb358d9f4da7e23b668c5a59f86b9a5f185ebbdec

Observation e82e2748-b1b2-4c7d-8af1-c7965ec7ffd1 · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:12.224386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:12.224386Z digest=sha256:13e8a386d460f5e31d7cc862dfffc2fcdeeab869cb31989d3aaee22b3d0ec536

Observation cb1f136a-c96c-44d0-988e-ba3c11983bbc · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:12.347223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:12.347223Z digest=sha256:be7d99de9d1f9b483fe4e5456b2992f9b8a9311beab23da7a7c304f9d1d464f1

Observation 9de57c66-ddfb-431a-a362-d76b9ef0e045 · outbound

This paper cites RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:12.421263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:12.421263Z digest=sha256:538f3239212e9ef086a42d6d50ae4c5322475f31dadf7f4d4e57e642a0ffd567

Observation ada30bc4-d367-4783-84d5-4505a4ff3a49 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:12.487930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:12.487930Z digest=sha256:ce2b5c8b9d86acf8f54f1c5a24f4c2e9caa48104811b3d8babe2a2b33302e105

Observation 0963b69d-513b-4722-9f53-b6e98d6d77b2 · outbound

This paper cites Humanity's Last Exam.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Humanity's Last Exam

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:12.558874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:12.558874Z digest=sha256:5ee7adfb9091b643d4c6e53e1cba97807b0acc1321f8ae4910369ffb677ab8a1

Observation c9fa9b8f-0a71-4404-bfa8-50257e2d7f2b · outbound

This paper cites A.; and Lewis, M.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation A.; and Lewis, M

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:12.679843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:12.679843Z digest=sha256:8a57c47c42195df7b8bfe89984f9fd0b4932b398e69448ee5aa4c0a9bfd82650

Observation 4092235b-ccae-4176-85bf-b185cc3f057f · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation ToolRL: Reward is All Tool Learning Needs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:12.771078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:12.771078Z digest=sha256:f9a26b93b391d6f14d37aa9e53c89b429740abb80b64068f159e654678f272fb

Observation 34a1e834-d20b-47b2-8a3c-e1d6e14a4129 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:12.826891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:12.826891Z digest=sha256:411d0dcd799f821d73b906ce4774104520d9eca8053317517a35cdcb0388ece9

Observation 9a661eac-8563-4f2a-a7ef-5e93bd6d0096 · outbound

This paper cites Qwen2.5 Technical Report.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Qwen2.5 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:12.878590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:12.878590Z digest=sha256:70f173aa0b14a5718a1f0725e3b04403bd26596452f3fcbdd3e6cb25065093b5

Observation 91871e85-3f73-4f91-a227-2b77ec7785f3 · outbound

This paper cites Trust Region Policy Optimization.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Trust Region Policy Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:12.984770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:12.984770Z digest=sha256:1e97b939637fb60942dfe96b3db2fc9cfbfa2a3dd3bd77da3dcea530a006514a

Observation 62bbf2f3-e045-4d4e-8366-0fc3a5a9b147 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:13.171364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:13.171364Z digest=sha256:48640a156e47bc3d1edd2480fbbb2b5fb763ca8012684b9fc5d28315ad5c674a

Observation 72b3b020-c5c2-47dd-a3f4-1e6b8b5d2a7e · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:13.270674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:13.270674Z digest=sha256:047339b3660449a471e9bf905bfc15ac1f73b48f7c8391d3cef7e47be7cf36e2

Observation dab872a8-80f4-47d3-bcc0-6e660dfdf5c8 · outbound

This paper cites A Survey of On-Policy Distillation for Large Language Models.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation A Survey of On-Policy Distillation for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:13.333132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:13.333132Z digest=sha256:fd111748bf3e3c89634706eaf5c0dc30e0f8c540cfd8df92488ff3f6c2974dd1

Observation 811a75d3-c651-446d-a832-ee15c4909a8d · outbound

This paper cites X.; Liu, Z.; Fang, L.; Wang, Z.; and Wen, J.-R.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation X.; Liu, Z.; Fang, L.; Wang, Z.; and Wen, J.-R

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:13.436571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:13.436571Z digest=sha256:30d564d938b91e8d4fdea3c329519f883dcd9b3faa73de3853ee90360e3fc298

Observation 018541f5-f4e5-452d-952a-a769dc764789 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:13.484761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:13.484761Z digest=sha256:cd53bf3c2314070bfb2fe9f2c071b9b09f052b6deeb354250a27f0b87785beaa

Observation e131f767-25f6-42a7-932d-9ec8ed7670e5 · outbound

This paper cites Qwen3 Technical Report.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Qwen3 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:13.667751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:13.667751Z digest=sha256:ddf4075ebe3fa8db1c6c0646bca48863c6244a3bb5cbcb38dab12e2da70bc755

Observation 8624e6d0-bb2b-480e-895f-50db18ebc1c5 · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:13.724129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:13.724129Z digest=sha256:df8fdd4b499952ac62c7ce10a11a0d4eee6d9a3960cbe0f9f2c8441296e4bbac

Observation b5e99fc8-d4e3-458f-8d85-93fc80964da7 · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:13.786194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:13.786194Z digest=sha256:0664ebdf25330a16d2d16b0b3e1112bec4b522e6c2e6beeb1e7c5a81c483097f

Observation 9b21d4aa-800b-41f8-9efe-4f2e9bf4e5a6 · outbound

This paper cites Acting Less is Reasoning More! Teaching Model to Act Efficiently.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Acting Less is Reasoning More! Teaching Model to Act Efficiently

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:13.834925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:13.834925Z digest=sha256:7c902df1cbb7339464c4f7e94c345461c97bd65542c918d31eceabb7a573ad88

Observation 8602b83e-74a1-4df6-adf4-69a7ced1361d · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:13.895257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:13.895257Z digest=sha256:9e455d0e69ce4e83959ede0d425682b0cf4c9568cdaadc121e5aef0cf225113e

Observation b3e667ac-d082-4c4b-b015-1688e38ef6e1 · outbound

This paper cites TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:13.989520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:13.989520Z digest=sha256:e6b9179bb3b395f508936a8c46f72d506d9d1dd9ca620ed23a9e6194f553f909

Observation 5c68a7ed-1f71-49fa-b8e9-d74a21fbd5aa · outbound

This paper cites WebWalker: Benchmarking LLMs in Web Traversal.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation WebWalker: Benchmarking LLMs in Web Traversal

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:14.128623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:14.128623Z digest=sha256:3b0754f4f1abc114143e9ffb0dab9f9d5885ececf9876a07acf34c80bc3b3a3b

Observation 846cd24f-ac38-48f1-8360-f23f935f241b · outbound

This paper cites Self-Distilled RLVR.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Self-Distilled RLVR

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:14.199650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:14.199650Z digest=sha256:7b1a3f98aac757a200f3d28777f8ec366bcff042d3fe47a146e72e10d0479ff1

Observation 550d5251-4169-4b22-9ba6-d06c224b8420 · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:14.300743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:14.300743Z digest=sha256:02f562b0eee9d7c3a884ee561e6bed65efb335faffa7fd5b89e360155f46d263

Observation 25ca512c-9c07-46af-ba54-6d7c2d508b3b · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:14.419599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:14.419599Z digest=sha256:ffe0ac0e62ef1bccbcddd1c9212bc318afd8260875f40fc2adf4f83f5da6b961

Observation 88a9b1c6-d37d-4c86-8a4f-7534b224ec5a · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:14.537242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:14.537242Z digest=sha256:405c604e83125a612162bdb837d2b656c60d47efdd1fc44c6f5bbe4a667d98df

Observation 44cd3696-1fcd-4a30-86eb-60f4ca868a4a · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:14.625513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:14.625513Z digest=sha256:af5dfb5f4893d8b474d099d1e08618b1b219c1d870a9d7ac99d3fefc0433d1ff

Observation 59d94601-e7fe-4ece-8812-da39d05c3d6a · outbound

This paper cites 2026 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:14.711800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:14.711800Z digest=sha256:f68ee8e6f3e830f1ebafa987e50c65db0b62b81622689464a07e1aefbcd27bb7

Observation d86c6bd7-1c4e-4495-bdfa-6b80e3d6fe9b · outbound

This paper cites 2019 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2019 , eprint=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:14.766543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:14.766543Z digest=sha256:60030fc7eb71cf1bff4e91f9e557623eb8c822eaf41c0ac0dca032d7f50d56c4

Observation 8608f900-1a77-4b0d-b0db-4c1925a5e8e6 · outbound

This paper cites 2017 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2017 , eprint=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:14.839334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:14.839334Z digest=sha256:a16a6696cda6c2ed4653808b0676e8973218febb2fbc9110497c7ac8a7e34924

Observation 9d395184-cc6c-49d2-9bd7-04652f2b188c · outbound

This paper cites 2026 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:14.952774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:14.952774Z digest=sha256:79a344a5a65f98f92424d17e41b88e81bda6dc6c98925577c4d8aa64325ee6d1

Observation 44a8c549-a6a1-499d-a19f-0e7be510b8ba · outbound

This paper cites Neural computation , volume=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Neural computation , volume=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:15.046759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:15.046759Z digest=sha256:71af6ecee59fa0f33eed59511835dafddfa4cb236c226156092ed876c026a90a

Observation 3332816d-4ffc-4dfd-b74a-ee3d6859d39f · outbound

This paper cites 2026 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:15.104480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:15.104480Z digest=sha256:fb2db99715daea965784c2ec036b60c84c4538239d12889471e616b24cb82ab4

Observation 6fe8b07f-b698-4f43-b9be-3220b8175815 · outbound

This paper cites 2026 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:15.192925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:15.192925Z digest=sha256:7fea4273274515754960bd0325b608edab70e69fcf300eb18ccf7dc364f5e333

Observation ac0b075e-f6f5-4a60-98ee-80c00cfee5d2 · outbound

This paper cites 2026 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:15.298526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:15.298526Z digest=sha256:c4890d4b56c0bc0a107b64651e2eba16babcc96cb47ff1da2d2de8f3ace8f3d2

Observation 9e09070b-5bb1-44e6-ac6e-36e3ff0e8031 · outbound

This paper cites 2026 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:15.424298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:15.424298Z digest=sha256:65299f4850c560170c45cf9e029dca241046ebb3912b4cc7ff4723e89095e090

Observation b371e495-e319-4735-bc9a-64b54857c711 · outbound

This paper cites 2025 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:15.516068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:15.516068Z digest=sha256:e7621ade9357d843dd1c5770d66cd3a90a21d5241bda596199318fbe6a50c02b

Observation 543e41ae-722a-4286-bf4f-673b48dd2459 · outbound

This paper cites 2025 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:15.562071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:15.562071Z digest=sha256:e327b1761caacfe651568da6a1836938ecad6d5f1664f9c2920fb735a3ed089c

Observation 12c26a91-dfee-44a8-8300-8892d7bd539b · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:15.623282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:15.623282Z digest=sha256:da5cc820b0fd16a070c1a0651967a8c738ebbf70455b7c27fcf5bd20930eb849

Observation 1743c439-9a8d-4aa0-a6b2-86ae8b402a13 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Advances in Neural Information Processing Systems , volume=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:15.683841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:15.683841Z digest=sha256:b4eb0d400319877fdefca321cf0032475ab4816d4e6b3a2adda3006077efb1f6

Observation e0f59ccf-2626-4a01-b237-66fb8e5e2152 · outbound

This paper cites 2026 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:15.749648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:15.749648Z digest=sha256:ba95c1173dead9fd85b74455411dc9b722d8c4c956ef0ff85561dc22f89c3035

Observation 6cb86448-61a3-4dab-9654-2f6fa6622c70 · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:15.862134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:15.862134Z digest=sha256:18b1ec9e976a6a1425853e6e76defc9b5a72550cfbace1fbe04fa10bf5aa40ee

Observation 1e6b0060-32ca-450a-ad26-ed13e820aa18 · outbound

This paper cites International conference on machine learning , pages=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation International conference on machine learning , pages=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:15.924752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:15.924752Z digest=sha256:4d1ddf289b6306974f4c95c4a444c332444531ef5fd518fcf219d92cfe2644af

Observation 4e012e7c-5b89-4f09-8281-fcd6c9086c2d · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:15.975629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:15.975629Z digest=sha256:be747e5f680d4ba208b63a0e8f77a4daa15c2a787f883c57d0748cf4c97b46a2

Observation 16202c2f-256f-4467-af26-5fe53b79dd4d · outbound

This paper cites 2026 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:16.060463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:16.060463Z digest=sha256:d7d1a7b6e46b290a72de479a4629fdfd05c71c7bd321d5d26dc3e278c5b4ab41

Observation 0a41fa58-5dbc-451f-944c-d9fd867cd7e4 · outbound

This paper cites 2025 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:16.128351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:16.128351Z digest=sha256:b1e4e6867d1fe9956ba9112a5e8c11eff4a9e5c52efe17d27619472d05c84fa6

Observation 3bd9eeee-7a96-43ad-b2b9-ea856e3848a8 · outbound

This paper cites 2026 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:16.194956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:16.194956Z digest=sha256:52c8398874636109b049b26341257a45a54fc888204f5bae96e30440a60dc260

Observation 054422f6-b294-4362-91c1-c4782176615d · outbound

This paper cites 2025 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:16.276264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:16.276264Z digest=sha256:529b40c8d346b2feaa67af5190dacd89ab14d43cda4810c288d02404f815305f

Observation ce9cee68-264c-40f9-956e-2a305d5ed643 · outbound

This paper cites 2025 , publisher=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , publisher=

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:16.314303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:16.314303Z digest=sha256:0d3f504ad81644ef6207eaf8499c96b0cbbc642fe06e50821e1f1d9b28af2344

Observation f5d52a3a-126f-4b35-89aa-50842525c4a2 · outbound

This paper cites 2024 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2024 , eprint=

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:16.383494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:16.383494Z digest=sha256:e532646e9aca5d7e1a87acf89a90873cbb0ff8cee806ed2b99e19f7af5b6599e

Observation 5061c51f-b42d-4e67-a605-111aef5eb10d · outbound

This paper cites 2024 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2024 , eprint=

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:16.433285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:16.433285Z digest=sha256:1654e554868e0514c9d35530bebf6123bb10384fef1bd1a4fcb2b1834ab555fe

Observation 5d18457c-9127-4333-9e0b-d59a3fcef5b3 · outbound

This paper cites 2026 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:16.493620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:16.493620Z digest=sha256:0dbcf345886e2d69597aeabc31a2ac0fea56d54d5c8fd79107be67fb0018a6fd

Observation 849d2471-d40d-4995-a925-01da668dae6c · outbound

This paper cites 2026 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:16.544150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:16.544150Z digest=sha256:f643937bf4181ea1a4a646cd992f4e972772bd39785668322475d95070e0d2cb

Observation 8b890c80-d7df-4a71-a5e3-2eb1bbbda0eb · outbound

This paper cites 2025 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:16.612575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:16.612575Z digest=sha256:83a76c48dc57dca5945c9c837ee67581fe1c016692d05a5d1983692d9ad793ae

Observation 5ebbab12-40db-4e2d-9e42-fbcbbfaf0d50 · outbound

This paper cites 2025 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:16.664956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:16.664956Z digest=sha256:43457ac6d92e0cb8be62fa44b76cc65db7d61e1a6f0d1686775a1c627ae15e8e

Observation d3f64341-ca73-4245-ab5d-8bf92be0bd53 · outbound

This paper cites 2025 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:16.733403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:16.733403Z digest=sha256:8a956ca1ae0211d439657578db115784c2087adcfbbfb0596be546b9094d2c1f

Observation 131a81a1-2ec6-4d91-9692-897bedef8767 · outbound

This paper cites 2025 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:16.811324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:16.811324Z digest=sha256:80185305fc070279a8cbf02112e408dba6eb54f24d896f9d9c6c497f696847e7

Observation 97aa4d34-18bf-484b-96ae-cd0a026c30df · outbound

This paper cites 2026 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2026 , eprint=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:16.922134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:16.922134Z digest=sha256:f4a2b5963a9172af6991a8347ab778c7ea1534574edbc7fa0d7017ddf95cffc8

Observation 28dd0634-a439-4a5b-ad1e-3d9850fe217e · outbound

This paper cites 2025 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:16.999661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:16.999661Z digest=sha256:83f731b34c50bdc79b2aa0eff058d2a1b0d2a7888bc258172af396ef7ad5b2da

Observation eeb9cd58-0ba9-44e5-97a3-6401d897ea80 · outbound

This paper cites The Twelfth International Conference on Learning Representations,.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation The Twelfth International Conference on Learning Representations,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:17.113456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:17.113456Z digest=sha256:125dc4ab2b55495895fb7a63424f9d0c9aa381a8a65f193193ad110014bc139d

Observation 8ee5a9a5-131e-4bba-93ad-230fc8fb2c91 · outbound

This paper cites Measuring Mathematical Problem Solving With the.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Measuring Mathematical Problem Solving With the

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:17.186190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:17.186190Z digest=sha256:8f7ca4b581dfa9ec3aaad867b42a83b7c290ea4685305d0d3cafa693c412cc90

Observation 263eb880-823c-416d-8b72-850373b39ef7 · outbound

This paper cites WebWalker: Benchmarking LLMs in Web Traversal.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation WebWalker: Benchmarking LLMs in Web Traversal

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:17.267411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:17.267411Z digest=sha256:06de5a5091d4d9f3349e8c2af14f0054312ddc5bb8e9b4325a81836aa2d5768e

Observation 4807d86d-2c0d-42c5-babe-8f7a14350a00 · outbound

This paper cites H otpot QA : A Dataset for Diverse, Explainable Multi-hop Question Answering.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation H otpot QA : A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:17.361743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:17.361743Z digest=sha256:8d091096375b6f5e11297b3810dfea731864e9907d5fc8fa3d8eafb6da5a20cd

Observation 2b670483-abdb-4cea-939a-bb0cf1118122 · outbound

This paper cites Constructing.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Constructing

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:17.415396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:17.415396Z digest=sha256:7f9b1958b8d727504ff936c7c7e39b70ec7dd5275ef702bfb4d5960966a5fbe3

Observation 8b17e206-de3f-45bd-a77a-f48fb3f1a918 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Transactions of the Association for Computational Linguistics , volume=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:17.514321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:17.514321Z digest=sha256:914245587f15ad350cad38edccf6e8bea02da89e28570495e8d3b66832002d1d

Observation 9293f171-e8e9-4a4b-ba3f-b906d1218c37 · outbound

This paper cites Smith and Mike Lewis , editor =.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Smith and Mike Lewis , editor =

Reference 90

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:17.617436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:17.617436Z digest=sha256:748be2bae52464420aa6c400464452a5531c8a0306dc8433d6fbc9b919778b21

Observation 890ed648-0450-47f9-8569-e41e3502dfbe · outbound

This paper cites The Twelfth International Conference on Learning Representations,.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation The Twelfth International Conference on Learning Representations,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:17.721536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:17.721536Z digest=sha256:b53060239cfe42feb80cb7eac0ff610021dbe8b2d59d29f57400580b028fec9c

Observation b9ad37d2-76ef-456d-b4bb-305020f4f8d5 · outbound

This paper cites Humanity's Last Exam.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Humanity's Last Exam

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:17.833208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:17.833208Z digest=sha256:36cfd3ed1e6410753ad0f317ca1da1da52fe26f5c7902ced8c640a405b74cc70

Observation f9dd0a5b-95db-457e-a828-5b1ddcc9d7e0 · outbound

This paper cites xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:17.957159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:17.957159Z digest=sha256:9ee0080d18d831739f30588b56c5bfef85fed876f3735805de2d9c656523f049

Observation fa5ab322-c608-4204-a62f-46fe6816f7df · outbound

This paper cites 2025 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2025 , eprint=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:18.039899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:18.039899Z digest=sha256:46e5a6fb6680a71088fb6944fe1514d158f080bc72b7dcc12d7d8bb17676d89d

Observation b6ccbbc4-bf13-4618-b828-bd37700535d2 · outbound

This paper cites Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:18.144305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:18.144305Z digest=sha256:a02bb7118ad0507f71e47787126c7949d4bf67a4b3ce4b7da1c0ed49083ea92d

Observation 9ada9f8a-0811-4840-b194-490a576432e4 · outbound

This paper cites 2024 , eprint=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 2024 , eprint=

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:18.249687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:18.249687Z digest=sha256:e5740d326e74f8faf06a0b11f75d8b4b0c2551192f3b7de3b3ceeb7b74e22e2e

Observation 0641966a-50f4-4f8c-af2d-8acff291ea1f · outbound

This paper cites The Llama 3 Herd of Models.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation The Llama 3 Herd of Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:18.352604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:18.352604Z digest=sha256:dc07b478d2a9245fb27f84fa44aa12b3d5aa97ed1143a5d86b189d8d515bcfe1

Observation 87f63080-8ba1-4606-bc5d-7bcb9447781c · outbound

This paper cites Qwen3 Technical Report.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Qwen3 Technical Report

Reference 98

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:18.447646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:18.447646Z digest=sha256:1584d5d6ecca050a36f8948334b05d104ad3a58ea99d0765aa59d74e1ee40dd9

Observation 3ddbcd7c-9a1b-45ab-b37f-798aa9b91dc1 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:18.552844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:18.552844Z digest=sha256:b947fd0fbdace12a39f5b447e2c71de63d3a493a6d0d60c035ab1c199108ddd1

Observation a017e6ba-202c-4485-93bb-6e5508be01ba · outbound

This paper cites CoRT: Code-integrated Reasoning within Thinking.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation CoRT: Code-integrated Reasoning within Thinking

Reference 100

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:18.658766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:18.658766Z digest=sha256:d5446042e6781355345e3f41e2d1a3231d308ac26d302786a7bbe2c3b317f32a

Observation 3517e788-ab76-46b9-9ce7-f6e87a9da2d9 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) , address=.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) , address=

Reference 101

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:18.759274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:18.759274Z digest=sha256:2cab18c1ac3e4e591f623b40ba4c68d38b9102650ebbd2ca4e51672061be7c5c

Observation 9a783ed2-84d5-4b31-9c67-0cbccbdaa270 · outbound

This paper cites an unresolved cited work.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Unresolved cited work

Reference 102

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:18.832821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:18.832821Z digest=sha256:9e5d1e42f9f8bc9e7e5dd7e56083f07b0b17da7e4bd72b6d05318590dee396ea

Pith citing papers

No inbound Pith citation observations are available.