Pith. sign in

Paper Citation Record · LEDGER

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization

As of 19 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2605.19436.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19436 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T07:25:34.637803Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:17:44.132452Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T04:17:44.759376Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact18
  • verified fuzzy10
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd16c89e-b10c-4460-8175-6f7dd9798967 · outbound

This paper cites Qwen3-VL Technical Report.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.692342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:6b8642a98fb84172dd35e641fa3299ff61707a88234a3f217553f732f9cfc086

Observation 9f7f6f2a-cbd1-4c62-8605-8652dff3c5b9 · outbound

This paper cites Enhancing reinforcement learning with dense rewards from language model critic.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Enhancing reinforcement learning with dense rewards from language model critic

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.789041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:be686359cf6d4bb7629e15adade947555432c3dd8a2c88c9464e37c35e897fd4

Observation de406f7d-30ce-4ab4-9bfa-dee53628e503 · outbound

This paper cites Hdpo: Hybrid distillation policy optimization via privileged self-distillation.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Hdpo: Hybrid distillation policy optimization via privileged self-distillation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:28:06.732172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:855faf4d41e46c649dc666e74e6187539cd3b8a0933ec8e857beef5fe51fce71

Observation 49e1f607-f397-46ee-be6e-07acc6a6a873 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.715382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:1afbf13d0dc5b6487245fe60453e7c9ded248b4a4fb877694e411ced30ba6562

Observation 3bba8b69-66f5-4828-818a-8142788905ab · outbound

This paper cites Segment policy optimization: Ef- fective segment-level credit assignment in rl for large language models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Segment policy optimization: Ef- fective segment-level credit assignment in rl for large language models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:28:06.696139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:7acc7ef9d7384c40105f3bbe972ba24ed28b1f4f38838c1ad72ef138f8555a21

Observation bfb0cdab-fa69-4677-b91b-e458441e5908 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Reinforcement Learning via Self-Distillation

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.740802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:4515bcfab5cbd60c22e074ab64a461e6d2c2e0a26663813e399cad8b0e762129

Observation d41cce22-37b6-41db-8958-43dea5070073 · outbound

This paper cites Vineppo: Refining credit assignment in rl training of llms.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Vineppo: Refining credit assignment in rl training of llms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.806585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:843eae7bd1ff61f92e7cfc57ee9a4329f69b95464df35dcc12d3717452275bd5

Observation f752aaaa-ec60-4f3c-8bca-5ec713ef0259 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Efficient memory management for large language model serving with pagedattention

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.803751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:71ee0390cc7af9df7c346d10cdbe50e324d6e6fa0dcc221518419de261b1e362

Observation fb9b350e-582e-426b-b3b5-7c252a10b441 · outbound

This paper cites Let’s verify step by step.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Let’s verify step by step

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.801519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:8a8a3dc0484eaa02cb8c6e7107f127ae2d3478e142bdf37d1dd80ad58f4b8c97

Observation 9437ccef-8cb4-4beb-b315-9edb91f04ed3 · outbound

This paper cites Decoupled Weight Decay Regularization.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Decoupled Weight Decay Regularization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.743892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:abc287c0efeb1ca267b7c0aaccc52dcf857357cec3534272ccd8a534932335fb

Observation ef5f2040-11c3-4095-ad24-8bf13606edf0 · outbound

This paper cites Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.808483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:9b13840377b9db3fd68c40bb91de9ee89d4a74de258ba27b095c9bb67e81e9af

Observation e0793516-e4ec-483a-b45b-b9f87305e1cf · outbound

This paper cites Privileged Information Distillation for Language Models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Privileged Information Distillation for Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:f3bc2ab266ffe34ad5596863f286670b263e5c539c891b60b56e3914622d483a

Observation 6b6b04f9-7334-44dc-8466-e12506ac3660 · outbound

This paper cites an unresolved cited work.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-20T07:28:07.810293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:04340df5795f15072cf7411f4b4fa3902e8b3d344c58f12585ceadd00c9c4f2d

Observation 813132db-5d51-4a2d-8414-50ecf1838102 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Direct preference optimization: Your language model is secretly a reward model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.797670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:c17bdbb5f1d78087c80daecdb4a3d5be0365ef8a2fe7958e94b2a5f501600d98

Observation 3653660e-4cd6-40b9-8842-ee81e9705f96 · outbound

This paper cites Proximal Policy Optimization Algorithms.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Proximal Policy Optimization Algorithms

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.712482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:b51c9d64e57bccfb7af3100adcf2ac7d530fe8cef41186b625c3c1b259d2cbf0

Observation c424760a-73d0-4362-8a38-f3972fd3a327 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:42:19.263401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:6122291e0dafa6a385939f84b5aba72f2df99029135094813691f890fcc426c3

Observation 74ae5e96-c6ca-4e43-9f32-8705243c0aad · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.699451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:b2686e525b0bb362232c00d058ee6d5997e9bd370272d3316dee4b75704d4dda

Observation 49ba4685-b6b3-42f3-8f79-ec2a17ab8e83 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Measuring multimodal mathematical reasoning with math-vision dataset

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.792830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:9367435bcb9f5a78b50b506fbaabbb4518b5faa516467e0a89a78912fb9284ed

Observation 1dc53f8c-65dc-4127-818e-bb63134e56ae · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.738133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:db0233d7a640e26e417cfdb5d12cfda24d73a837038d56b9d1234c29be7ca39c

Observation c6c886dc-109c-49d2-ab17-458f03f002c4 · outbound

This paper cites Qwen3 Technical Report.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Qwen3 Technical Report

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.735030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:a1e97c3ec7839323a331b76e363a7fdb6839623748fe32d0d1d6440e2617e5de

Observation f4444fb1-8423-45b8-915b-697e7349ccef · outbound

This paper cites Self-Distilled RLVR.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Self-Distilled RLVR

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.709083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:6f58233e8634c50973337be3eb477881de522e9a55a834a09969555aa0185e77

Observation c2f87e32-de1c-4df6-8d41-f21c45c57c44 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.718341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:059685c06b66a167f2af85b7b35a384fea540f0899ee1acf19f5fac55fa13e8a

Observation 1b06546e-8fd5-4bf0-889e-d5238403854e · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.795477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:d9265a5a867d120b0382e633d6f44be64ba5d487e8abd9a8cb2c881d02f3f096

Observation cf422d56-d3d4-4fd6-814f-79948df646ba · outbound

This paper cites From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.747246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:41f491590b8e1ddcaa2b8a661e5a1a64d64b1c13f11b9de9adb960c1dcb6f8bc

Observation fb4dd984-5cab-4c60-b6d5-7b8ee616348b · outbound

This paper cites Lmms-eval: Reality check on the evaluation of large multimodal models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Lmms-eval: Reality check on the evaluation of large multimodal models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.799533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:fb9e19e9feae5226d25fd330fc6a768dd555b5769dd596e514f131ff0683fe02

Observation 19b192e4-371e-4a1b-ae75-3129d962a124 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.706111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:0b5467543f3a2a4dd9db5a4073549af83579cfae8201d0265c2811f241cd3f47

Observation 268fae63-85c0-4afc-8e47-1e13dcb335b2 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.702664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:d7bd509e82217c0a7b5218606e66e3932c9a0274ee0eed377ee028bfe6782e18

Observation 8c601782-78d5-49c8-bcf3-7c09989aa339 · outbound

This paper cites EasyR1: An efficient, scalable, multi-modality RL training framework.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization EasyR1: An efficient, scalable, multi-modality RL training framework

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.790952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:bebe28fbde5957ca44eaee93ba30066fa89a0aeacec4941a8d9aaacd83716f63

Observation 94f731f4-8d62-4817-aa77-2f3d071c6fc5 · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:28:06.725569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:d238b43720d551501a89f6fcdc7efded619b7c13a204608274666b836ec0e15e

Pith citing papers

Observation e331cbeb-d053-4a43-b835-8a72ecaea650 · inbound

H$^2$SD: Hybrid Hindsight Self-Distillation cites this paper.

H$^2$SD: Hybrid Hindsight Self-Distillation CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T13:52:48.791895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:52:48.791895Z digest=sha256:ff790f888f2cbaaf86ac5e0661e9638b4f04409170e51713045103d12a8a35ed

Observation cc1008c9-cd83-4da0-8597-d3369fc394c9 · inbound

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots cites this paper.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T04:17:44.764322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T04:17:44.132452Z digest=sha256:68719d3c6f620f63f5eb2812aa7bbb4bd765cb58dcad75713935e79155ad1a2d