Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T19:03:10.268534Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2605.17333.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T19:03:10.268534Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T21:50:58.824041Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T02:06:27.263175Z
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4b1ba4e5-b7ca-4f4b-aed8-b5f84a16258e · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Program Synthesis with Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9a366783-c26f-4655-a718-8b1a770589a0 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Matharena: Evaluating llms on uncontaminated math competitions, february 2025.https://matharena.ai, 8
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9a99b717-4914-450d-b475-106ec02b200c · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Post-training as reweighting: A stochastic view of reasoning trajectories in language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1e6cacf5-f354-4274-8065-8e7706cdfbb0 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1ae4db99-e03f-4ad9-b81e-c18b5c139651 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning arXiv preprint arXiv:2505.09655 , year=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a087e67a-aee1-4f92-ba25-bba483404220 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Reasoning with exploration: An entropy perspective
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bab5517e-8183-4245-920f-7286a4c45e73 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning American invitational mathematics examination-aime 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2ffba4e4-c132-4868-95e1-e07535b29155 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Harder is better: Boost- ing mathematical reasoning via difficulty-aware grpo and multi-aspect question reformulation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 515a2898-7620-4fc6-917e-df207c8fe663 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Interleaved latent visual reasoning with selective perceptual modeling
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cd543f84-2aed-4200-a98c-9b4d694d84d1 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation be741e74-6ada-4578-b50d-cafb9fb6b427 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Reasoning through exploration: A reinforcement learning framework for robust function calling
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7640a5c3-5593-45b7-9692-81e770d0db2d · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ca3964a8-5ba9-41df-9c34-a73d013d2318 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning OpenAI o1 System Card
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bd7e061d-91b4-4aae-bd34-9b09ef6de84e · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7c1f0d5b-7377-427c-ab7f-b81e6b50be0e · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Setpo: Set-level policy optimization for diversity-preserving llm reasoning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 72e8effc-8268-4ac5-a301-fd26280d3594 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning DeepSeek-V3 Technical Report
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 828c286c-4b7f-42bc-bc3b-e780b5ada7e6 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Code-r1: Reproducing r1 for code with reliable rewards
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 36188bb1-2d4a-41c3-a9a1-e14b8d9f3a4f · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8d5fe6b1-9741-47ca-ae9e-077feedbf07f · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Diversity-aware training for test-time scaling
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5a8455af-7cf7-42a6-b624-df92f5cbc661 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning American mathematics competitions - amc.https://maa.org/
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 73a0841b-21a0-4296-8266-2b0c081e654e · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning American invitational mathematics examination-AIME 2025.https://maa.org/
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0b0e6343-3c06-41c6-9278-e0671f7d8c48 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning American invitational mathematics examination-AIME 2026.https://maa.org/
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a2c47669-c47c-4152-b605-5bb3e996df66 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Ngrpo: Negative-enhanced group relative policy optimization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d0b9eb4a-8f38-4e2e-b05c-42ed6bf2bafe · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a8602bf4-fdd3-43d0-81b9-7ac7c85e116b · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7b6013b6-fc6d-4461-87c8-9f57e21f19c5 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation adae003e-4066-4f99-a89f-c9ac8719bf8d · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 21f9fa6e-f89d-41a0-bc2c-22798f808b0d · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1b62b345-12eb-4937-b131-b52703c3bcec · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 90f96817-8764-417f-915d-28ad14c00973 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Grouter: Decoupling Routing from Representation for Accelerated MoE Training
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 43f199f6-69b9-4a14-ad0a-0185143ee4d6 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Qwen3 Technical Report
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 19724671-9237-4f97-9bd1-03974b406884 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 522341bc-f461-4681-81fb-fff501ead888 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 89914b81-2235-4847-bb65-3a9f4ecd70db · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7fb99a93-ad11-4443-8c89-252b3c0098a7 · outbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning Reinforced Efficient Reasoning via Semantically Diverse Exploration
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f3364cc3-5a39-48aa-8a57-b0a171f6c7c0 · inbound
Grouter: Decoupling Routing from Representation for Accelerated MoE Training Leveraging Error Diversity in Group Rollouts for Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba8ccdcc-64d8-4c07-b1a1-472c9b830409 · inbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Leveraging Error Diversity in Group Rollouts for Reinforcement Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.