Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T01:39:11.464437Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2607.14614.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T01:39:11.464437Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0689e7eb-20f2-4117-8bdf-3be15eeebe3d · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization R., Geist, M., and Bachem, O
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e5efa7-533b-43f0-b14f-cbe4f735064e · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 703dcbbb-8474-4336-9b52-e30328945863 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61b93e40-3d7f-4cdf-9e2a-120b63a30cdc · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Reasoning with Exploration: An Entropy Perspective
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9dea355-746b-4272-b43b-a642bb4cdfdb · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization K., Chen, G., Xu, W., Luu, A
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b230e182-72c6-4661-8c47-925519a34555 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20fe05de-a943-431b-a89d-ac5d7396223a · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Deepseek-r1: incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081): 633–638, September 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0e898bf-91bf-4164-b613-9c1939a72c09 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Measuring mathematical problem solving with the MATH dataset
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65be8710-17e3-4c27-89cc-916385f00bef · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Reinforcement Learning via Self-Distillation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fed4f4cb-834b-4735-97b1-c9820b7609d1 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization F., and Joty, S
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdd53a9a-c87d-4837-9844-6cd8ab6e4b68 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fceea892-3f7f-4d84-bd97-d024ff84b8f4 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization V ., Jeon, M., Vu, K., Lai, V ., and Yang, E
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d47ed9d-be70-458f-ac3e-f1b4e783acb8 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Revisiting LLM Reasoning via Information Bottleneck
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2ef72b9-6fad-4cf5-bc2f-4e97442b67c9 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization RED: Unleashing token-level rewards from holistic feedback via reward redistribution
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80a54d28-261d-4a24-9277-f7be0d804108 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Can we further elicit reasoning in LLMs? critic-guided planning with retrieval-augmentation for solving challenging tasks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71afbd51-17a8-4d77-a281-67b2a3efd58b · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Let’s verify step by step
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29d4ae00-9413-4c24-a698-8104a05dbb6b · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Ravr: Reference-answer-guided variational reasoning for large language models.arXiv preprint arXiv:2510.25206, 2025
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aadada7-4f06-40c2-ae7f-1c52211b643a · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization and Lab, T
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25553960-5153-441d-a7f7-d972e5747e04 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization American mathematics competitions (AMC), 2023
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 378c8d48-05fa-42ab-89c3-d7fbe08828c7 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization American invitational mathematics examination (AIME),
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da0d60d1-89d8-4da0-8fe0-adffee18193d · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization L., Stickland, A
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cf673c1-2b74-44d7-a479-ee9a0a320d1f · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Proximal Policy Optimization Algorithms
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40e7aa23-9973-41a1-bbe8-68ba3ebc29eb · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Accessed: 2025-12-23
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61dfcddc-5072-4349-ab74-4c0092e78eca · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization On entropy control in llm-rl algorithms.arXiv preprint arXiv:2509.03493, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f33537c-e256-496a-9950-69b95093438c · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Self-Distillation Enables Continual Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e94f6c29-ce3c-4863-a626-f99a505f94d6 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1bec4d7-fd43-4282-aa3e-8760191cd133 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Kimi K2: Open Agentic Intelligence
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87738a82-72e2-41dc-bae3-09dd7cd67609 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce0c3b6b-14d8-4b8e-9bbf-26743f2329ba · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Espo: Entropy importance sampling policy optimization.arXiv preprint arXiv:2512.00499, 2025
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f56af7d-adc5-4061-ac52-35a00dfab8da · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff2c5ad0-5fd8-464e-8202-baa1d11ad2fd · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73026038-d1e2-4d6c-926f-ef4f97237b75 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization and Karkhanis, D
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c18f4ab1-a996-42ca-be58-bfbab0daff46 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Beyond the 80/20 rule: High-entropy minority tokens drive effective reinforcement learning for LLM reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83fa065d-7d82-4a64-9b34-c917ca434a1c · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization MMLU-pro: A more robust and challenging multi-task language understanding benchmark
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f2f0074-1d9b-49d1-84ec-92c6fd15cd36 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddb3deaa-f872-47f7-b5f4-6ecc6f2c91f6 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52a76309-0236-4747-90da-a17746c21878 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Quantile advantage estimation for entropy-safe reasoning.arXiv preprint arXiv:2509.22611, 2025
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43079864-152a-4fda-b2cd-4706f00e4322 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization OpenClaw-RL: Train Any Agent Simply by Talking
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf97f670-e25a-4dfc-bc18-fbb7f0b853f9 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Reasons to reject? aligning language models with judgments
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 804dc322-12c0-41bd-b001-027a18c64b29 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5017778e-003b-4130-a602-849d58ae4ddc · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization A., Osten- dorf, M., and Hajishirzi, H
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30e4be79-6d98-45a4-b9e9-669e3bc40d89 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Self-distillation bridges distribution gap in language model fine-tuning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0a49c8e-a0dc-4b59-b561-71e72d61c698 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Tree of thoughts: Deliberate problem solving with large language models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c38c0ee-0f90-401d-af90-170dc0ea9620 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Qwen3 Technical Report
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbdfea30-7b90-42f2-9959-1aa0c07b2398 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c46e2e8f-02b8-4482-9d30-57bcf13ee76d · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17252ee0-4e5a-4a37-bc01-6e96b3b77dce · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization DAPO: An open-source LLM reinforcement learning system at scale
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de45ab0f-6525-4a0b-8a4a-4ac946e6c732 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization R., Kailkhura, B., Lai, F., Zhao, J., and Chen, B
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8331d4b8-dc11-4eda-b24c-c560ab2b1b34 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization First Return, Entropy-Eliciting Explore
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee0017c6-1e68-410b-b189-a4cd75cdf2f4 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Group Sequence Policy Optimization
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d65af65-e512-42f9-b615-bd93fd42e247 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4291b401-e973-4b24-aaa0-f65e5ad55dce · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3aae2d3-5a8a-4f6a-97a7-4faf062ba1d2 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef22e5eb-c616-46c7-90a0-0193d8386361 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6baacee-6734-4eb0-a29b-3a115665d2f3 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 064e667f-ead7-4b58-9de7-827eee6bce35 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization URL https://aclanthology.org/2024
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e3e6b8c-43f5-4ad2-a9d5-4ff845c7fd23 · outbound
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.