Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T14:53:35.133157Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2605.04077.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T14:53:35.133157Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-13T02:39:02.891861Z
A source-named dated measurement, never combined with another source.
Source: cited_works
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 211aca48-d816-4b43-b352-a918446314ba · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4bd1f855-baca-4e31-bcde-2229ff47f0b5 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Polaris: A post-training recipe for scaling reinforcement learning on advanced reasoning models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 71b97418-39dd-4d49-9561-98ae38de3798 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO doi: 10.1038/s41586-025-09422-z
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d95e123c-cefd-437d-979e-8175a6a61e12 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c3b6d048-6406-4d07-b0de-7542497de3ab · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5a63f6d2-3f06-4cc1-88fc-82de48b8aa5d · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ab27407c-c671-4305-986c-f5f44e43a829 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Solving Quantitative Reasoning Problems with Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0a107278-ab2f-4440-b8ad-b167c0ea7ae2 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Let's Verify Step by Step
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8ba900d0-bf9d-43fb-add6-c3263eed1c68 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO When speed kills stability: Demystifying RL collapse from the training-inference mismatch, sep 2025
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b1e7bdfa-110b-4308-8d7e-8e5e2340dc99 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Understanding R1-Zero-Like Training: A Critical Perspective
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 341b9cd6-8374-41d2-8e61-dd5f19fbd4c8 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Part i: Tricks or traps? a deep dive into rl for llm reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cb09adba-3b2b-4458-a298-8f46e1910865 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b3e036a8-8ebb-43b7-bced-178e70046601 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7044776b-ea3f-4f8e-bf8e-deecf929a894 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO OpenAI o1 System Card
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 39166dec-8cb5-46c4-9d29-5926cd8c6210 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Proximal Policy Optimization Algorithms
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cbfb241c-9b98-4775-be60-bda02c02e715 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c784cbe9-8f83-4373-a686-8bb6783d64bc · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Opencompass: A universal evaluation platform for foundation models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b92f68a3-4c5a-43b4-a0bc-3aa3aa6c15f0 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 25cb378c-f7eb-4da7-8bf1-4bccb1395ec8 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 91f37ab9-165c-4042-9941-65b18def4bbd · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4cb9fe47-170a-4206-a6e0-9ab33945155a · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Qwen3 Technical Report
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0a9b4965-476c-4916-ad3b-697a13ac4864 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Your efficient rl framework secretly brings you off-policy rl training, aug 2025
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bce42b06-c2ac-4fea-90e5-3f3fb739f9da · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c7ae73c4-ac59-494e-99c4-9fad838e55d0 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Scaling of search and learning: A roadmap to reproduce o1 from reinforcement learning perspective
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 417f197f-201b-43ff-a2f1-911c9467be71 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Geometric-mean policy optimization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c9207f86-320b-46a1-858b-26b9e0ab66fc · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Geometric-mean policy optimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 903fd567-d6bc-45b4-b150-5dcadcd2cccc · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Group Sequence Policy Optimization
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f27d840a-d571-415c-93ae-f940d582561b · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Rloop: An self-improving framework for reinforcement learning with iterative policy initialization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9788758c-caa1-4f6d-84f0-c343fb9e6a04 · outbound
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO 14 Appendix A: Why Use Sequence-Count Weights in BA? Here we briefly justify the choice of weights k/G and (G−k)/G in BA
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 419567d6-c8b2-4ee9-907b-633a9b827c4d · inbound
Multimodal Reward Hacking in Reinforcement Learning Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.