Pith. sign in

Paper Citation Record · LEDGER

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2605.19436.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19436 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T07:25:34.637803Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:52:48.791895Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact18
  • verified fuzzy10
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd16c89e-b10c-4460-8175-6f7dd9798967 · outbound

This paper cites Qwen3-VL Technical Report.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.692342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:d705f482d2619927566ea89db516523c89a444ef6c44905fb3a1f60178da6327

Observation 9f7f6f2a-cbd1-4c62-8605-8652dff3c5b9 · outbound

This paper cites Enhancing reinforcement learning with dense rewards from language model critic.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Enhancing reinforcement learning with dense rewards from language model critic

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.789041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:4da68eb91265bf8197004fffc010fd901aa04712482dab72a0b3333c346a7397

Observation de406f7d-30ce-4ab4-9bfa-dee53628e503 · outbound

This paper cites Hdpo: Hybrid distillation policy optimization via privileged self-distillation.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Hdpo: Hybrid distillation policy optimization via privileged self-distillation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:28:06.732172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:c67e556025ae680705408c2bc032df20320c81ed9bd99d7f7b903ca2b1eb862e

Observation 49e1f607-f397-46ee-be6e-07acc6a6a873 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.715382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:cd791c5d30d2bb952e4d016a0ee4fb54cb99aafe92ed8c161e32eafeb35a71d9

Observation 3bba8b69-66f5-4828-818a-8142788905ab · outbound

This paper cites Segment policy optimization: Ef- fective segment-level credit assignment in rl for large language models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Segment policy optimization: Ef- fective segment-level credit assignment in rl for large language models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:28:06.696139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:5db9bdd78a52179d3528253b7780ed6e4a0a7ba4fc4c9ca45fe0180414dfd317

Observation bfb0cdab-fa69-4677-b91b-e458441e5908 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Reinforcement Learning via Self-Distillation

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.740802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:eeda5f5e89547d7451e05690f0425bbd9ea99f0b43251f967fc33d3648b10fbc

Observation d41cce22-37b6-41db-8958-43dea5070073 · outbound

This paper cites Vineppo: Refining credit assignment in rl training of llms.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Vineppo: Refining credit assignment in rl training of llms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.806585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:98ea4068c5419bcce02901df165a8785a348e2426babdd0bb57c98941f7471fb

Observation f752aaaa-ec60-4f3c-8bca-5ec713ef0259 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Efficient memory management for large language model serving with pagedattention

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.803751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:12246cc264c377588ec9a6fc95e05e8eae65239e0c0541fe0cd0318d6567c36c

Observation fb9b350e-582e-426b-b3b5-7c252a10b441 · outbound

This paper cites Let’s verify step by step.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Let’s verify step by step

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.801519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:8a096a31f4c5328276b4e1742ca5e821dc4e49e759b12779c652eff4a638bdec

Observation 9437ccef-8cb4-4beb-b315-9edb91f04ed3 · outbound

This paper cites Decoupled Weight Decay Regularization.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Decoupled Weight Decay Regularization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.743892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:52c0151d1ea00ebcf2cc611bab7979296db4e6209baf863d03c67273d0808869

Observation ef5f2040-11c3-4095-ad24-8bf13606edf0 · outbound

This paper cites Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.808483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:f0d51467c2f87c67a6e2b5987e4092805020780bd803f410a5aa6aa36ccb657b

Observation e0793516-e4ec-483a-b45b-b9f87305e1cf · outbound

This paper cites Privileged Information Distillation for Language Models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Privileged Information Distillation for Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:dfe1fabfc58112ec49ba238939934f954700e820b00451f96ef65ff277274a9c

Observation 6b6b04f9-7334-44dc-8466-e12506ac3660 · outbound

This paper cites an unresolved cited work.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-20T07:28:07.810293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:576b1b6573b2dcdfaf36b282ab0208346960a26d9182515b750f193364076e27

Observation 813132db-5d51-4a2d-8414-50ecf1838102 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Direct preference optimization: Your language model is secretly a reward model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.797670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:7f1562ea5659a40746a4eaf24e3d28d220567ca49a0289bd5f5409126053fb58

Observation 3653660e-4cd6-40b9-8842-ee81e9705f96 · outbound

This paper cites Proximal Policy Optimization Algorithms.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Proximal Policy Optimization Algorithms

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.712482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:baf1f3c3df997cf3a080abe6ea66a83186c307216b2443ffac0e757723dbaaac

Observation c424760a-73d0-4362-8a38-f3972fd3a327 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:42:19.263401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:b7817f18a06382ffa9ef316ee1c55082d3a257060593be65e61a3c774bfe6992

Observation 74ae5e96-c6ca-4e43-9f32-8705243c0aad · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.699451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:a2eee3f82ce2b64ebb53ddc362e0db665ba8b2af58e91d012a08fdc97b7232e7

Observation 49ba4685-b6b3-42f3-8f79-ec2a17ab8e83 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Measuring multimodal mathematical reasoning with math-vision dataset

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.792830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:453b5d3d60c24a1ba7d3847fdc7626de0c18b47c55ded5b1395d59c8c23250bc

Observation 1dc53f8c-65dc-4127-818e-bb63134e56ae · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.738133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:d1b2cd9ff2bf9c00dd1538d869c2385fcdc685c4d01063169320a04d7398bb23

Observation c6c886dc-109c-49d2-ab17-458f03f002c4 · outbound

This paper cites Qwen3 Technical Report.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Qwen3 Technical Report

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.735030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:1cb8ca577bc85a2a6fd7b378cbfec3c4a382ec96a1e16232c4d20f6400e0efe0

Observation f4444fb1-8423-45b8-915b-697e7349ccef · outbound

This paper cites Self-Distilled RLVR.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Self-Distilled RLVR

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.709083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:c105526036613f1430abfac1099620b25c820fa198f4999375e43bdba7833a1c

Observation c2f87e32-de1c-4df6-8d41-f21c45c57c44 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.718341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:57bf83d47d1f0e1da2d27fb4ec3648ffd804b6af05408fa06a5a14f105bc956c

Observation 1b06546e-8fd5-4bf0-889e-d5238403854e · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.795477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:1b011d1ac1fa1ea7d7e8d404f86f90651b05bf56281a98316dfd3d0530437e63

Observation cf422d56-d3d4-4fd6-814f-79948df646ba · outbound

This paper cites From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.747246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:77a536f27e080f4bfcc7fa8af102575bd0ab20edcfa5de9cd28876ecf283e895

Observation fb4dd984-5cab-4c60-b6d5-7b8ee616348b · outbound

This paper cites Lmms-eval: Reality check on the evaluation of large multimodal models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Lmms-eval: Reality check on the evaluation of large multimodal models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.799533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:a6c091359f1b31ba60734ec8525e19fcf8f93283ce4584cb4c57c98a8b9ed3f7

Observation 19b192e4-371e-4a1b-ae75-3129d962a124 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.706111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:c2e1a7850183175387bc5d30f5b91c1250b2539960ee339bb0e77a590ccef8c4

Observation 268fae63-85c0-4afc-8e47-1e13dcb335b2 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.702664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:d9706f8651398c0b96fb1523ec042c70b5fd92b83346214fec6a70298f9234e6

Observation 8c601782-78d5-49c8-bcf3-7c09989aa339 · outbound

This paper cites EasyR1: An efficient, scalable, multi-modality RL training framework.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization EasyR1: An efficient, scalable, multi-modality RL training framework

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.790952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:e64c85be86ab59343cac0aec52cffbf9e651a4d2df28e50a4aa0eb86b9a415ef

Observation 94f731f4-8d62-4817-aa77-2f3d071c6fc5 · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:28:06.725569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:1548679a8545c5dde3e1bd9f668969e88c3a04cfc3f4b7feb41dd98a907abee2

Pith citing papers

Observation e331cbeb-d053-4a43-b835-8a72ecaea650 · inbound

H$^2$SD: Hybrid Hindsight Self-Distillation cites this paper.

H$^2$SD: Hybrid Hindsight Self-Distillation CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T13:52:48.791895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:52:48.791895Z digest=sha256:29e5bad4db098f44bcc201f34cf87a7bf4d59bb284f1cf1be7996a869b4c87e5