Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:26:11.192134Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 16 inbound Pith citation observations for arXiv:2507.21848.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:26:11.192134Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T18:43:17.536451Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T09:47:59.935150Z
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d77ac373-6e99-49ab-a72e-7a0cf74997cd · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Efficient Reasoning Through Suppression of Self-Affirmation Reflections in Large Reasoning Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0b5dd49-9c09-4373-a9ad-b18ec0b14f2e · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Reasoning with Exploration: An Entropy Perspective
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f9b77b4-4d73-438f-ae59-802b873f62e6 · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Process Reinforcement through Implicit Rewards
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae609b57-5ef6-44cf-8f75-b48d85a05bab · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity arXiv preprint arXiv:2504.05185
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43d43ab2-3aa5-49da-8317-3a3b660ecc44 · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity One-shot Entropy Minimization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccdf164c-bf13-48e9-943b-2c4477f3f8e8 · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f60f95d-8c20-4a20-b7b4-c6434b7f16bb · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 636b9b26-0998-42ca-942d-608d8a23df06 · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity OpenAI o1 System Card
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43b0583d-889d-4377-9a2d-eec2da785756 · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity s1: Simple test-time scaling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8f6cb66-e737-4fd9-ae59-baf9f5a0791c · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d869d484-ae55-4538-8f5e-cefedf7cf522 · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10859b75-3b04-4d6a-8f61-5a3f70f0b845 · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5081309-1f68-478d-bee0-953e78a9b6a2 · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67b691d7-64df-43f5-9953-dbe20f528a52 · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity arXiv preprint arXiv:2506.01713
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35a5df91-f167-48b6-8b40-cd7dc059882d · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fe85726-aa59-461a-a040-c4053a797ed2 · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79f6c4c2-a856-43d0-a4a3-0617e3b1b1ab · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c632aa7-4cc9-45bc-b549-9787c239eba3 · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 215984af-1c66-40ea-892b-a9c8e23202bf · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Proximal Policy Optimization Algorithms
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a781e4e-ebb1-4424-b8b4-7cee598c08fd · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Measuring Mathematical Problem Solving With the MATH Dataset
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83f14dab-ca0d-4b49-8f7d-1a69efff6c58 · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Solving Quantitative Reasoning Problems with Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db9b1184-dcf1-439e-b20d-5f4c0624b0df · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69b01b73-13c2-4459-890e-334ee896733b · outbound
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf2cc2bc-117b-4cd6-847a-ba838efe361b · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68eb5094-2f5b-4c44-a765-9f568ccab9f4 · inbound
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 728510a6-8d25-44f4-9359-2b8725dd0455 · inbound
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6766ed2b-4d7d-4775-9f0d-4641884c93d1 · inbound
VIDEOP2R: Video Understanding from Perception to Reasoning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 558a877a-ae86-41e9-b42b-c1f3052f6287 · inbound
Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9622d80b-8959-4571-84e3-629ea3086da1 · inbound
StaRPO: Stability-Augmented Reinforcement Policy Optimization EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c9c248a-4723-453d-942f-2d66754ef541 · inbound
Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 58666f63-cb06-4f1a-9b0a-b2b0612dda90 · inbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 89914b81-2235-4847-bb65-3a9f4ecd70db · inbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b88c5685-4fd8-4dd7-9e91-768ecff1be67 · inbound
Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f711196d-2267-4ebf-9db1-99794f9289d3 · inbound
Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d573ebba-a89f-4ef3-9e76-5f893d5e3882 · inbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d0ebd72-b3e4-4607-a14a-536159c26a1a · inbound
Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e947bb10-d5b3-4419-b4d6-3f8825599d84 · inbound
Consistency as Inductive Bias: Learning Cross-View Invariance for Robust Multimodal Reasoning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71968d98-9133-4017-8a89-4e45d84158cd · inbound
Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a8f6b9b-0cad-4138-b6d3-a4bc95a6b36a · inbound
When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.