Pith. sign in

Paper Citation Record · LEDGER

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

As of 22 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 18 inbound Pith citation observations for arXiv:2507.21848.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21848 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:26:11.192134Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:39:40.838440Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:47:59.935150Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d77ac373-6e99-49ab-a72e-7a0cf74997cd · outbound

This paper cites Efficient Reasoning Through Suppression of Self-Affirmation Reflections in Large Reasoning Models.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Efficient Reasoning Through Suppression of Self-Affirmation Reflections in Large Reasoning Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.159006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.159006Z digest=sha256:9bf77c1d9c35190050542f159c4a35e4641845d4ab7de468ef13745c54b91090

Observation d0b5dd49-9c09-4373-a9ad-b18ec0b14f2e · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Reasoning with Exploration: An Entropy Perspective

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.126635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.126635Z digest=sha256:31d9b93a816b22ad8ff51875faeca0879aec4c0b234cda839a36415c78a6c5b7

Observation 2f9b77b4-4d73-438f-ae59-802b873f62e6 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Process Reinforcement through Implicit Rewards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.130126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.130126Z digest=sha256:0ad8c71cc9eec7f23c84f31a0a48c1a30add875af96715d064426a0e0b8eab77

Observation ae609b57-5ef6-44cf-8f75-b48d85a05bab · outbound

This paper cites arXiv preprint arXiv:2504.05185.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity arXiv preprint arXiv:2504.05185

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.133984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.133984Z digest=sha256:0729606a12e936d01ebef9c4a25a95834845b66e896f177dd4c86fdcd85f3c36

Observation 43d43ab2-3aa5-49da-8317-3a3b660ecc44 · outbound

This paper cites One-shot Entropy Minimization.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity One-shot Entropy Minimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.136900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.136900Z digest=sha256:2b65112149c0feefb0d1596bde5b2989b0a5144f743366ea0006ad0bb9df4311

Observation ccdf164c-bf13-48e9-943b-2c4477f3f8e8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.140006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.140006Z digest=sha256:f75687e2f0ef5681c5551f153eae5d6ff07023a634a9e112c4cac2d74f8b1184

Observation 1f60f95d-8c20-4a20-b7b4-c6434b7f16bb · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.149829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.149829Z digest=sha256:73feec954602224d2f47cd0b4b089bfaa788fcaed5f0ee2fbe19c8d9cc1fecad

Observation 636b9b26-0998-42ca-942d-608d8a23df06 · outbound

This paper cites OpenAI o1 System Card.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.152824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.152824Z digest=sha256:0f405ff7215071ffa6758bbd77013abb0ced749d323ebf18ae34adc9720bc060

Observation 43b0583d-889d-4377-9a2d-eec2da785756 · outbound

This paper cites s1: Simple test-time scaling.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity s1: Simple test-time scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.161943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.161943Z digest=sha256:542678abdc8cd4f9dd9d5dd84323baf163b2b5fbb0bc92144b4ff757e675a61f

Observation e8f6cb66-e737-4fd9-ae59-baf9f5a0791c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.167724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.167724Z digest=sha256:c650f5a6a21034f57a6b7659ebaffbdc754568458d2a9cabc1c9ef463a4acd6a

Observation d869d484-ae55-4538-8f5e-cefedf7cf522 · outbound

This paper cites Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.170585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.170585Z digest=sha256:daad3c6c64af470164f397f64e26b66852438165987f628a5ab4cd5a70f3f1ee

Observation 10859b75-3b04-4d6a-8f61-5a3f70f0b845 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.173550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.173550Z digest=sha256:52eb8db5d6e2ce30f8e89ffef01b4ab1ad299a2a1fe0fee8add521ba2c1c2555

Observation a5081309-1f68-478d-bee0-953e78a9b6a2 · outbound

This paper cites Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.177203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.177203Z digest=sha256:46171229a508cab91a8335af6e5324565ecd41e528a7ef221ba9bedeec771cbc

Observation 67b691d7-64df-43f5-9953-dbe20f528a52 · outbound

This paper cites arXiv preprint arXiv:2506.01713.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity arXiv preprint arXiv:2506.01713

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.180394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.180394Z digest=sha256:753dceb965eef53eb219f50e572b30347ff001733a77b049f5e847bc70e0c6ef

Observation 35a5df91-f167-48b6-8b40-cd7dc059882d · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.183062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.183062Z digest=sha256:9849e9e8326702803306808a68a6c52d2c7744122d0b61e11ceb443e9ab47116

Observation 7fe85726-aa59-461a-a040-c4053a797ed2 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.186011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.186011Z digest=sha256:ce4906fc9cdb0a355538d1f7205100b483da970d751401af4ae089491fc3a2e9

Observation 79f6c4c2-a856-43d0-a4a3-0617e3b1b1ab · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.188931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.188931Z digest=sha256:e2a772c7b23313c3675035b800958fc9c3775edaafb12326b519a2fe9412fd93

Observation 1c632aa7-4cc9-45bc-b549-9787c239eba3 · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.192134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.192134Z digest=sha256:95c31a470ff70573ad47c025c9206d10463857c861798714b40537db83b4fc23

Observation 215984af-1c66-40ea-892b-a9c8e23202bf · outbound

This paper cites Proximal Policy Optimization Algorithms.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.164984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.164984Z digest=sha256:23dc24b025de92984c20afc892e88d9b9de5ab74495d2ed86d0abfb6c3db49c9

Observation 6a781e4e-ebb1-4424-b8b4-7cee598c08fd · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.146760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.146760Z digest=sha256:3a3a76db71aee07eaae7bc2fd39a8bbf5dd523a54fceba7f15dc8fbef31ee82d

Observation 83f14dab-ca0d-4b49-8f7d-1a69efff6c58 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Solving Quantitative Reasoning Problems with Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.155945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.155945Z digest=sha256:e98a8632012566c78aa8607894f68d07e8d26e07620a73656925f95d3a82fd77

Observation db9b1184-dcf1-439e-b20d-5f4c0624b0df · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.143506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.143506Z digest=sha256:f522bb60101610009dad286c083da731ddfcf8546d173144f1d04ea706eaba62

Observation 69b01b73-13c2-4459-890e-334ee896733b · outbound

This paper cites SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.122616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.122616Z digest=sha256:c6063a8374749574749509960eaf06e84030cfa713a6baf130cf760c35f2c872

Pith citing papers

Observation cf2cc2bc-117b-4cd6-847a-ba838efe361b · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.764302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:44bf0dd541e734359c41f01a563c0ed46538e5cadae9ab45d063566dc357b8d4

Observation 68eb5094-2f5b-4c44-a765-9f568ccab9f4 · inbound

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents cites this paper.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.634858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.634858Z digest=sha256:77f007841d03c7b10da7c151fb5dbf9a7f6a7bf81f62d1d6f2f213753230a0b5

Observation 728510a6-8d25-44f4-9359-2b8725dd0455 · inbound

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training cites this paper.

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:30:35.435478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T20:30:17.581842Z digest=sha256:c4c0af9448af0983695bb6b5ddf47b6dc5bdf3b09926e3fadb76daec7bbca29e

Observation 6766ed2b-4d7d-4775-9f0d-4641884c93d1 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.674800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:b800f49c9da8b81355a42ecc8c7b678286d0deeef3329407695b9d84c735f302

Observation 558a877a-ae86-41e9-b42b-c1f3052f6287 · inbound

Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning cites this paper.

Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T20:44:47.413838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:44:47.413838Z digest=sha256:d7fed7bc8afafaf042040ee9ac0c1ad488f0e9593b3fd8cc9502e223e17c3c64

Observation 9622d80b-8959-4571-84e3-629ea3086da1 · inbound

StaRPO: Stability-Augmented Reinforcement Policy Optimization cites this paper.

StaRPO: Stability-Augmented Reinforcement Policy Optimization EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:20:59.951337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:11:49.056805Z digest=sha256:8138a396cbab560f8b1dff2966561c799a5d55165db4d2ac1503e3f3c833ae8e

Observation 4c9c248a-4723-453d-942f-2d66754ef541 · inbound

Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance cites this paper.

Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:13.056711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T08:05:36.206946Z digest=sha256:60fd572009671075dd1f20e9e4e46bf7229c2cbd9669766bd746f4b3c3e03560

Observation 58666f63-cb06-4f1a-9b0a-b2b0612dda90 · inbound

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning cites this paper.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.309046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T14:52:33.141693Z digest=sha256:7840fd188b8e661d74bb4d1bff05cc6d4c2dcbfa95e372f2845aeb8b71c6ec20

Observation 89914b81-2235-4847-bb65-3a9f4ecd70db · inbound

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning cites this paper.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.406585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:1e2449a71f063aaf24043964a53d8adaadd15f4946fbf2a7ac01b4face6deba3

Observation b88c5685-4fd8-4dd7-9e91-768ecff1be67 · inbound

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation cites this paper.

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:09:41.261365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-21T06:08:55.571828Z digest=sha256:16e22e1270bb1e9bdaf8630cdcbb958b9de0d6d16e7e1926a7e1f46991a7aa52

Observation f711196d-2267-4ebf-9db1-99794f9289d3 · inbound

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation cites this paper.

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:15:47.310221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T17:15:53.286925Z digest=sha256:f84af3f59afbc64ae2b129c3527eb06383c0b9161c81b55846869a9b775fdf7e

Observation d573ebba-a89f-4ef3-9e76-5f893d5e3882 · inbound

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO cites this paper.

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:49.989463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T23:54:32.621093Z digest=sha256:4f31be3ea36742674536e5416b03b1c32b0fc2e5f2973a402d208561612ac5b8

Observation 1d0ebd72-b3e4-4607-a14a-536159c26a1a · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:47:59.936872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:bf0009abc20802faec0e18fa35ae2fd632de0f42a7541630b8e88dd531750684

Observation e947bb10-d5b3-4419-b4d6-3f8825599d84 · inbound

Consistency as Inductive Bias: Learning Cross-View Invariance for Robust Multimodal Reasoning cites this paper.

Consistency as Inductive Bias: Learning Cross-View Invariance for Robust Multimodal Reasoning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:14:19.253487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T06:10:41.786328Z digest=sha256:0281f871e5a9c11ba88542d4410245ab9f59717dbd0729532a8a69c2e951e3fb

Observation 71968d98-9133-4017-8a89-4e45d84158cd · inbound

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index cites this paper.

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:42.413060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-01T05:27:22.527220Z digest=sha256:2c7184eb65572fef19ba71c17e7bb8baa5dddfa6b7cb442d20a8749def83bf87

Observation 0a8f6b9b-0cad-4138-b6d3-a4bc95a6b36a · inbound

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO cites this paper.

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T18:43:17.536451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:43:17.536451Z digest=sha256:f9245e2e4c8b213809d267ccea02fa1b0f7d8139c485ec95af0678681345a59c

Observation bbb13f40-929a-4252-ab59-dd58226c2ec5 · inbound

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO cites this paper.

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:47.185967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:52:47.185967Z digest=sha256:0322bef85945b73dfb02872c75a26a6068cd3c9412b7c8b4562e860d4153d713

Observation 3b81bcac-b01f-4ce5-917d-27972528646d · inbound

Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models cites this paper.

Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-12T00:39:40.838440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:39:40.838440Z digest=sha256:62299d80951798325432997b8f2a01c8c865b393fffc3401454fdbbbdb20bbac