Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T07:25:34.637803Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2605.19436.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T07:25:34.637803Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T13:52:48.791895Z
A source-named dated measurement, never combined with another source.
Source: cited_works
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cd16c89e-b10c-4460-8175-6f7dd9798967 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9f7f6f2a-cbd1-4c62-8605-8652dff3c5b9 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Enhancing reinforcement learning with dense rewards from language model critic
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation de406f7d-30ce-4ab4-9bfa-dee53628e503 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Hdpo: Hybrid distillation policy optimization via privileged self-distillation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 49e1f607-f397-46ee-be6e-07acc6a6a873 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3bba8b69-66f5-4828-818a-8142788905ab · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Segment policy optimization: Ef- fective segment-level credit assignment in rl for large language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bfb0cdab-fa69-4677-b91b-e458441e5908 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Reinforcement Learning via Self-Distillation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d41cce22-37b6-41db-8958-43dea5070073 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Vineppo: Refining credit assignment in rl training of llms
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f752aaaa-ec60-4f3c-8bca-5ec713ef0259 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Efficient memory management for large language model serving with pagedattention
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fb9b350e-582e-426b-b3b5-7c252a10b441 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Let’s verify step by step
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9437ccef-8cb4-4beb-b315-9edb91f04ed3 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Decoupled Weight Decay Regularization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ef5f2040-11c3-4095-ad24-8bf13606edf0 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e0793516-e4ec-483a-b45b-b9f87305e1cf · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Privileged Information Distillation for Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6b6b04f9-7334-44dc-8466-e12506ac3660 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 813132db-5d51-4a2d-8414-50ecf1838102 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Direct preference optimization: Your language model is secretly a reward model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3653660e-4cd6-40b9-8842-ee81e9705f96 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Proximal Policy Optimization Algorithms
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c424760a-73d0-4362-8a38-f3972fd3a327 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 74ae5e96-c6ca-4e43-9f32-8705243c0aad · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 49ba4685-b6b3-42f3-8f79-ec2a17ab8e83 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Measuring multimodal mathematical reasoning with math-vision dataset
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1dc53f8c-65dc-4127-818e-bb63134e56ae · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c6c886dc-109c-49d2-ab17-458f03f002c4 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Qwen3 Technical Report
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f4444fb1-8423-45b8-915b-697e7349ccef · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Self-Distilled RLVR
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c2f87e32-de1c-4df6-8d41-f21c45c57c44 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1b06546e-8fd5-4bf0-889e-d5238403854e · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cf422d56-d3d4-4fd6-814f-79948df646ba · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fb4dd984-5cab-4c60-b6d5-7b8ee616348b · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Lmms-eval: Reality check on the evaluation of large multimodal models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 19b192e4-371e-4a1b-ae75-3129d962a124 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 268fae63-85c0-4afc-8e47-1e13dcb335b2 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8c601782-78d5-49c8-bcf3-7c09989aa339 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization EasyR1: An efficient, scalable, multi-modality RL training framework
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 94f731f4-8d62-4817-aa77-2f3d071c6fc5 · outbound
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e331cbeb-d053-4a43-b835-8a72ecaea650 · inbound
H$^2$SD: Hybrid Hindsight Self-Distillation CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.