Pith. sign in

Paper Citation Record · LEDGER

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning

As of 4 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2605.17333.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.17333 v2

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T19:03:10.268534Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:50:58.824041Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T02:06:27.263175Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact25
  • verified fuzzy9
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b1ba4e5-b7ca-4f4b-aed8-b5f84a16258e · outbound

This paper cites Program Synthesis with Large Language Models.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Program Synthesis with Large Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.420630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:d7a1c1b7df03ae64f64b1b9726ff90f9f8b13050dbf48cde1023f28fc0316a18

Observation 9a366783-c26f-4655-a718-8b1a770589a0 · outbound

This paper cites Matharena: Evaluating llms on uncontaminated math competitions, february 2025.https://matharena.ai, 8.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Matharena: Evaluating llms on uncontaminated math competitions, february 2025.https://matharena.ai, 8

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:34:27.166592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:9992e9bb686c32712635fdeffc8d3201a6d75be2d261031495d6a407b2449deb

Observation 9a99b717-4914-450d-b475-106ec02b200c · outbound

This paper cites Post-training as reweighting: A stochastic view of reasoning trajectories in language models.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Post-training as reweighting: A stochastic view of reasoning trajectories in language models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.423181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:206c952bcb36d9058547ef6de5ea6a8a458b263e00704fe4b772f51098ce2ef1

Observation 1e6cacf5-f354-4274-8065-8e7706cdfbb0 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Evaluating Large Language Models Trained on Code

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.410488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:481717ddb1e7bf15d9bf462b49ffe9dafbea3dbc0c43168fc6212863bd565b92

Observation 1ae4db99-e03f-4ad9-b81e-c18b5c139651 · outbound

This paper cites arXiv preprint arXiv:2505.09655 , year=.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning arXiv preprint arXiv:2505.09655 , year=

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.408150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:90277c8744c4f4b8e249fed4ef9787d25cc7cee3bb0ea46550f7f3bf555154ad

Observation a087e67a-aee1-4f92-ba25-bba483404220 · outbound

This paper cites Reasoning with exploration: An entropy perspective.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Reasoning with exploration: An entropy perspective

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:34:27.168504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:8373a862b65ddf540e2914ba4e1b960576867f84982710cbf5f4c78b5ca411a3

Observation bab5517e-8183-4245-920f-7286a4c45e73 · outbound

This paper cites American invitational mathematics examination-aime 2024.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning American invitational mathematics examination-aime 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:34:27.170368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:5a82e2632fd13124919ae3df4fa746c3089444bafbe419b2cb26049af9b9da44

Observation 2ffba4e4-c132-4868-95e1-e07535b29155 · outbound

This paper cites Harder is better: Boost- ing mathematical reasoning via difficulty-aware grpo and multi-aspect question reformulation.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Harder is better: Boost- ing mathematical reasoning via difficulty-aware grpo and multi-aspect question reformulation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.425821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:701b79c98564a775e04cd022b3c4680ba3db15df18d739db0eb3b48313928f81

Observation 515a2898-7620-4fc6-917e-df207c8fe663 · outbound

This paper cites Interleaved latent visual reasoning with selective perceptual modeling.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Interleaved latent visual reasoning with selective perceptual modeling

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.420310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:fc5efcd594cd64092f7644e874c1df4d263d65605fcfb319f891657fc9db5dad

Observation cd543f84-2aed-4200-a98c-9b4d694d84d1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.433031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:5251073f30432c178217ba5abcbd1b261aba835b42df1ef27ca140fb128ea0e3

Observation be741e74-6ada-4578-b50d-cafb9fb6b427 · outbound

This paper cites Reasoning through exploration: A reinforcement learning framework for robust function calling.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Reasoning through exploration: A reinforcement learning framework for robust function calling

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.372591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:674938ca8a391c45da900209570e1c114cc71c0654565610641657bc4884d7f8

Observation 7640a5c3-5593-45b7-9692-81e770d0db2d · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:34:27.162531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:05ad1c23eca21dd7c3b10bf73dc9118e3185cfda309dd26a4babdfc9db613f89

Observation ca3964a8-5ba9-41df-9c34-a73d013d2318 · outbound

This paper cites OpenAI o1 System Card.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning OpenAI o1 System Card

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.388344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:0ca38282a31531552d65caf13651af32211105f43f3c583b792358bb256eddf9

Observation bd7e061d-91b4-4aae-bd34-9b09ef6de84e · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.413123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:3969dcc4c95d480f894f9c3a991346be295ed9741fe6ca1611b3cbc3a13485eb

Observation 7c1f0d5b-7377-427c-ab7f-b81e6b50be0e · outbound

This paper cites Setpo: Set-level policy optimization for diversity-preserving llm reasoning.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Setpo: Set-level policy optimization for diversity-preserving llm reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.435940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:0aca52dcbd1b418a1901509ecc20a69b7e347babd4354f897c40298306f29574

Observation 72e8effc-8268-4ac5-a301-fd26280d3594 · outbound

This paper cites DeepSeek-V3 Technical Report.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning DeepSeek-V3 Technical Report

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.425260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:9d6398ea870e9f451848029be3b90e25c6f6ddb01626d11b39ce8ab2311f5520

Observation 828c286c-4b7f-42bc-bc3b-e780b5ada7e6 · outbound

This paper cites Code-r1: Reproducing r1 for code with reliable rewards.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Code-r1: Reproducing r1 for code with reliable rewards

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:34:27.159023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:c9b892f89e5b2100ea60265c719ab75837502a516607016fc9491ff1a5a46c2b

Observation 36188bb1-2d4a-41c3-a9a1-e14b8d9f3a4f · outbound

This paper cites an unresolved cited work.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-07-08T01:34:27.160665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:890832aa420a8f24bd43aa8f9da8e48571a7e48cb48ea4ba3e4e1beba5b7e6b3

Observation 8d5fe6b1-9741-47ca-ae9e-077feedbf07f · outbound

This paper cites Diversity-aware training for test-time scaling.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Diversity-aware training for test-time scaling

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:34:27.164596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:6a4e3a85d1b60ce8f1ff5e6e80c4a7ae201d2bae211fce6eaaab2499248121a3

Observation 5a8455af-7cf7-42a6-b624-df92f5cbc661 · outbound

This paper cites American mathematics competitions - amc.https://maa.org/.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning American mathematics competitions - amc.https://maa.org/

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:34:27.155163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:601ef34c504cf3dbcf389dbbf2a6d39d1537aee0935a2f1c6e607e4439c51e89

Observation 73a0841b-21a0-4296-8266-2b0c081e654e · outbound

This paper cites American invitational mathematics examination-AIME 2025.https://maa.org/.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning American invitational mathematics examination-AIME 2025.https://maa.org/

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:34:27.157135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:790172ef6120598ef7b0a53f8027dd05a0c7bb2ea4f818c8728b0af8abee4057

Observation 0b0e6343-3c06-41c6-9278-e0671f7d8c48 · outbound

This paper cites American invitational mathematics examination-AIME 2026.https://maa.org/.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning American invitational mathematics examination-AIME 2026.https://maa.org/

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:34:27.153189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:456593ff8672d7fd85c8256228c3fcf1bb8eed55bcd9fc2d329a1207315f0841

Observation a2c47669-c47c-4152-b605-5bb3e996df66 · outbound

This paper cites Ngrpo: Negative-enhanced group relative policy optimization.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Ngrpo: Negative-enhanced group relative policy optimization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.402722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:5714530f8ef18ec324b753a6642e31fe09e1be545d24e00b71d7e56c6be42863

Observation d0b9eb4a-8f38-4e2e-b05c-42ed6bf2bafe · outbound

This paper cites CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.422916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:6e1a2918a0f9bcfefc87bd213827e89f3b542d0484880b6dd32d25e2f7cd5ebe

Observation a8602bf4-fdd3-43d0-81b9-7ac7c85e116b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.430705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:c7a853732ae773d2cfcf459f9ba935884adcad9ee78166ba957f2c4d05e04f1d

Observation 7b6013b6-fc6d-4461-87c8-9f57e21f19c5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.394706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:b202cb31d9b89358f4716a70cfbae1d99c12306541c84f751f0725acde3aec81

Observation adae003e-4066-4f99-a89f-c9ac8719bf8d · outbound

This paper cites Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.378626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:5db1298a602dbf872447edd8927b64b1889699cb5796050c1bbd8f17d1f9f552

Observation 21f9fa6e-f89d-41a0-bc2c-22798f808b0d · outbound

This paper cites MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.401145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:63a3d5cc50f70883417b856cb63f8bfa5a3af06833573405b80ae52be0f302b2

Observation 1b62b345-12eb-4937-b131-b52703c3bcec · outbound

This paper cites Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.415466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:377b3565b2dc193fd61e9f8f9a0740b298e52e29714fdb105587930d6dfac834

Observation 90f96817-8764-417f-915d-28ad14c00973 · outbound

This paper cites Grouter: Decoupling Routing from Representation for Accelerated MoE Training.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Grouter: Decoupling Routing from Representation for Accelerated MoE Training

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.369568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:82a9ff4f2d94614c1b65eb1e7248a20d8c6cbe86b4aee2b86f1dff3996aa4c5e

Observation 43f199f6-69b9-4a14-ad0a-0185143ee4d6 · outbound

This paper cites Qwen3 Technical Report.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Qwen3 Technical Report

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.408984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:5e7efd687250536133ab49c7d7668bdf6d08a30b4cc0ec62872f559d3a947e62

Observation 19724671-9237-4f97-9bd1-03974b406884 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.428210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:6a92d49f4b5418cb021f6ae98169c51fd9a5f6f3e7928a4ad855790cd835cfb5

Observation 522341bc-f461-4681-81fb-fff501ead888 · outbound

This paper cites RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.411604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:d823db81bd91b4a813fa8e1446781bf634595da80b68d606f32a833c26a8b4df

Observation 89914b81-2235-4847-bb65-3a9f4ecd70db · outbound

This paper cites EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.406585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:2b059b5402a760c8ec4624efe3e5f51927ddbce4ce389da8548852994fd05316

Observation 7fb99a93-ad11-4443-8c89-252b3c0098a7 · outbound

This paper cites Reinforced Efficient Reasoning via Semantically Diverse Exploration.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Reinforced Efficient Reasoning via Semantically Diverse Exploration

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.390998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:b87786b10ab39f448aec10e1436e0a6431fbf1a1974fab26c4f3ea3e55fe6d15

Pith citing papers

Observation f3364cc3-5a39-48aa-8a57-b0a171f6c7c0 · inbound

Grouter: Decoupling Routing from Representation for Accelerated MoE Training cites this paper.

Grouter: Decoupling Routing from Representation for Accelerated MoE Training Leveraging Error Diversity in Group Rollouts for Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T21:50:58.824041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:50:58.824041Z digest=sha256:48a23c994b4a6818b3f926dbcc38a6799447d564a26d1ac4c4fb551a8339bf80

Observation ba8ccdcc-64d8-4c07-b1a1-472c9b830409 · inbound

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning cites this paper.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Leveraging Error Diversity in Group Rollouts for Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.264764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:4f6334b3fda8e3f7f431b56123f3a6f0c646b2a5eea4b538389256f35331285b