Pith. sign in

Paper Citation Record · LEDGER

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

As of 7 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 16 inbound Pith citation observations for arXiv:2507.21848.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21848 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:26:11.192134Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T18:43:17.536451Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:47:59.935150Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d77ac373-6e99-49ab-a72e-7a0cf74997cd · outbound

This paper cites Efficient Reasoning Through Suppression of Self-Affirmation Reflections in Large Reasoning Models.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Efficient Reasoning Through Suppression of Self-Affirmation Reflections in Large Reasoning Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.159006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.159006Z digest=sha256:443a7146319794f8b20150792b5d19b6c2bd555f08575763f4aab15ae369172d

Observation d0b5dd49-9c09-4373-a9ad-b18ec0b14f2e · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Reasoning with Exploration: An Entropy Perspective

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.126635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.126635Z digest=sha256:c1e76da9afc5ce1ce37f234592115e0ff3026ffaf41ecef493dc869d384b417b

Observation 2f9b77b4-4d73-438f-ae59-802b873f62e6 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Process Reinforcement through Implicit Rewards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.130126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.130126Z digest=sha256:63a10165d1a64a503cf80da0763ca6b0c9ef80964d1fd400568ade94d4e7e50a

Observation ae609b57-5ef6-44cf-8f75-b48d85a05bab · outbound

This paper cites arXiv preprint arXiv:2504.05185.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity arXiv preprint arXiv:2504.05185

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.133984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.133984Z digest=sha256:7d969821812d405556bd0531048fa56f58fcbce0f50e9fb8e16017957cfc3fe1

Observation 43d43ab2-3aa5-49da-8317-3a3b660ecc44 · outbound

This paper cites One-shot Entropy Minimization.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity One-shot Entropy Minimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.136900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.136900Z digest=sha256:b2ce27acda49ccae0472abda0710dab94b34b6963b5bfd650ce9449492b3198b

Observation ccdf164c-bf13-48e9-943b-2c4477f3f8e8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.140006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.140006Z digest=sha256:a876ee7e1a8375d809abb125ba260fa3b43e21686dcf0886b1145a3fa125c856

Observation 1f60f95d-8c20-4a20-b7b4-c6434b7f16bb · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.149829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.149829Z digest=sha256:dd80619cd17db85911d338b21798e3e2430529753b9c4417b91ad242e53baa38

Observation 636b9b26-0998-42ca-942d-608d8a23df06 · outbound

This paper cites OpenAI o1 System Card.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.152824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.152824Z digest=sha256:1ef0bd1dd216ffc416fd79ff122c6c6645536ac00b4501d1078a9e91a405ec7f

Observation 43b0583d-889d-4377-9a2d-eec2da785756 · outbound

This paper cites s1: Simple test-time scaling.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity s1: Simple test-time scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.161943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.161943Z digest=sha256:017ed91de73ce8ec1be6877e53c3e6e800c2ace5f7d9c4a55bf4091c068e73fb

Observation e8f6cb66-e737-4fd9-ae59-baf9f5a0791c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.167724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.167724Z digest=sha256:82bdbabb374b9820bc5ba5d7e72c49a4bba862970513250e3c6eb3c7c2b874e0

Observation d869d484-ae55-4538-8f5e-cefedf7cf522 · outbound

This paper cites Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.170585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.170585Z digest=sha256:16634492bf8a62494d9261948d7cef2ff941962e8ac266bc14c5c1a314825069

Observation 10859b75-3b04-4d6a-8f61-5a3f70f0b845 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.173550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.173550Z digest=sha256:76217b5d6c1f44a3a1b9e0155be7fb2e0d8431edb7640f0d230f4346950b9422

Observation a5081309-1f68-478d-bee0-953e78a9b6a2 · outbound

This paper cites Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.177203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.177203Z digest=sha256:8cec70d61600b2f30dc08b6ff4c284585e35fc4456eb90ccdfbd2aa472839967

Observation 67b691d7-64df-43f5-9953-dbe20f528a52 · outbound

This paper cites arXiv preprint arXiv:2506.01713.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity arXiv preprint arXiv:2506.01713

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.180394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.180394Z digest=sha256:2dd63345e1529dc65863a849784233d273d96dbeca233e364b263bc174836bca

Observation 35a5df91-f167-48b6-8b40-cd7dc059882d · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.183062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.183062Z digest=sha256:2da4cda4956c1abf7bac09fa228940d5f423abc14deb5494130373eabe2ab502

Observation 7fe85726-aa59-461a-a040-c4053a797ed2 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.186011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.186011Z digest=sha256:5360e7ec14fa408a79623c24e790a9941565c10608d7f9ad25f09df9558116a7

Observation 79f6c4c2-a856-43d0-a4a3-0617e3b1b1ab · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.188931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.188931Z digest=sha256:9a8cc00518db56cf337e234454412e4322987b3891fa8d2878695f7a962750db

Observation 1c632aa7-4cc9-45bc-b549-9787c239eba3 · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.192134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.192134Z digest=sha256:3012421c7e357f7839eebef0658a7a0c0885706baa19704f0b27e64a987ed9f2

Observation 215984af-1c66-40ea-892b-a9c8e23202bf · outbound

This paper cites Proximal Policy Optimization Algorithms.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.164984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.164984Z digest=sha256:7e718dcaa0a19d0e84f9809e25c3293c405eb5ecda9bf13dd0c8a15b8b43136c

Observation 6a781e4e-ebb1-4424-b8b4-7cee598c08fd · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.146760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.146760Z digest=sha256:80f11376fc9f0a2f78646b46db67c24e5cbcc20a6ca5b709c8667c9c1966d6e3

Observation 83f14dab-ca0d-4b49-8f7d-1a69efff6c58 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity Solving Quantitative Reasoning Problems with Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.155945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.155945Z digest=sha256:f4456bcfd540500f28adbb89955b711bdca2a1e44954fb2ce0e6999f7f12b748

Observation db9b1184-dcf1-439e-b20d-5f4c0624b0df · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.143506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.143506Z digest=sha256:b310cbfb789f5af07c2048e8952d90c6a6c0f8dd07680b8f40ff6085a778c0be

Observation 69b01b73-13c2-4459-890e-334ee896733b · outbound

This paper cites SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization.

EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:11.122616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:11.122616Z digest=sha256:9d941cce452e464be51c6a3698a687458e50d101f2e55c9687910fd00ebb4fd9

Pith citing papers

Observation cf2cc2bc-117b-4cd6-847a-ba838efe361b · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.764302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:d2eb5a10ecd50e13be41665fb2743fbfb741b55fef8abda733fba348d430c62a

Observation 68eb5094-2f5b-4c44-a765-9f568ccab9f4 · inbound

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents cites this paper.

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T19:29:00.634858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:29:00.634858Z digest=sha256:b0a5331ef4ff03bf6865bd30d480809a49d9275a32a0bad61c65d88152149d6c

Observation 728510a6-8d25-44f4-9359-2b8725dd0455 · inbound

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training cites this paper.

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:30:35.435478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T20:30:17.581842Z digest=sha256:7e76c7848e8ce3ac9b34b867a2083434668d97b9ee4fca47b969b9e66e500144

Observation 6766ed2b-4d7d-4775-9f0d-4641884c93d1 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.674800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:885d4d60e8890bcb8052c08239159b0baf7b2cc17dea07e258f55b8c748e0153

Observation 558a877a-ae86-41e9-b42b-c1f3052f6287 · inbound

Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning cites this paper.

Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T20:44:47.413838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:44:47.413838Z digest=sha256:f0a98c04923787a6520d90b26ff693d10ab19aeee7d3c78cf6b55d6582655d64

Observation 9622d80b-8959-4571-84e3-629ea3086da1 · inbound

StaRPO: Stability-Augmented Reinforcement Policy Optimization cites this paper.

StaRPO: Stability-Augmented Reinforcement Policy Optimization EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:20:59.951337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:11:49.056805Z digest=sha256:13e4120b119e477875b70f40d6437dd8e0a3e5834a10334c63b5db1d9c0ae607

Observation 4c9c248a-4723-453d-942f-2d66754ef541 · inbound

Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance cites this paper.

Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:13.056711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T08:05:36.206946Z digest=sha256:9d77864d9aa60772492f1bb6fe6e07994369d68e636f2512b6cb4bd3c0f3d66c

Observation 58666f63-cb06-4f1a-9b0a-b2b0612dda90 · inbound

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning cites this paper.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.309046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T14:52:33.141693Z digest=sha256:4aae34faed21bf0acea8f727c6a15af6895208282827d40190d11b79bf03011c

Observation 89914b81-2235-4847-bb65-3a9f4ecd70db · inbound

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning cites this paper.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.406585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:fa27a92f779b26b41b6a69cab8c978179a20bfcfcdea76744781e996254c625f

Observation b88c5685-4fd8-4dd7-9e91-768ecff1be67 · inbound

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation cites this paper.

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:09:41.261365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T06:08:55.571828Z digest=sha256:257ed39a24848ddd37fdb765b2db9c1845bce6198e176be11796ccd48e8790af

Observation f711196d-2267-4ebf-9db1-99794f9289d3 · inbound

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation cites this paper.

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:15:47.310221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T17:15:53.286925Z digest=sha256:7db7cb4e7415e2f2e7f57fee5df0732dfa3651d8e00d2f19e4918b61b6721a9b

Observation d573ebba-a89f-4ef3-9e76-5f893d5e3882 · inbound

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO cites this paper.

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:49.989463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T23:54:32.621093Z digest=sha256:5028c9faaffbf66f1d53acb9ffd50ee5fbc4ec4a74ad4361b5426c410a5fd60f

Observation 1d0ebd72-b3e4-4607-a14a-536159c26a1a · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:47:59.936872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:8beb4408631f6f72a7577dae9431b25f1821492a131f2cebb6479b669c930193

Observation e947bb10-d5b3-4419-b4d6-3f8825599d84 · inbound

Consistency as Inductive Bias: Learning Cross-View Invariance for Robust Multimodal Reasoning cites this paper.

Consistency as Inductive Bias: Learning Cross-View Invariance for Robust Multimodal Reasoning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:14:19.253487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T06:10:41.786328Z digest=sha256:c1fcc6aa08345fb4a390b0021f3433cc30497be75f0ab0769959a0b86b75e915

Observation 71968d98-9133-4017-8a89-4e45d84158cd · inbound

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index cites this paper.

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:42.413060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:27:22.527220Z digest=sha256:bee8085ef835dbb6a00509a3e3e29595df4632c5d69ef65d0ec8ed20dc2f027a

Observation 0a8f6b9b-0cad-4138-b6d3-a4bc95a6b36a · inbound

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO cites this paper.

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T18:43:17.536451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:43:17.536451Z digest=sha256:c3820e067f0cc877b2becccb5599b9d4c10f3a9de1a09a33f0dd572a6a959d78