Pith. sign in

Paper Citation Record · LEDGER

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO

As of 5 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2605.04077.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.04077 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T14:53:35.133157Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T02:39:02.891861Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact19
  • verified fuzzy8
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 211aca48-d816-4b43-b352-a918446314ba · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:26:04.692861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:1303d1ea9fac012c86b7d5aca5edff86317f2e5e8b107f031951b3f5f4d98dbb

Observation 4bd1f855-baca-4e31-bcde-2229ff47f0b5 · outbound

This paper cites Polaris: A post-training recipe for scaling reinforcement learning on advanced reasoning models.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Polaris: A post-training recipe for scaling reinforcement learning on advanced reasoning models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:21:20.349344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:2fa029a2dbd557593054152b0098b5b07ae50539d3a29e5522518bf3733c75a0

Observation 71b97418-39dd-4d49-9561-98ae38de3798 · outbound

This paper cites doi: 10.1038/s41586-025-09422-z.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO doi: 10.1038/s41586-025-09422-z

Reference 3

Resolution
verified exact
doi, observed 2026-05-10T14:55:31.456353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:8f7a7114e5be05bec0a925207b5aa5b6e494cd58387bb30f7ad365388f21d170

Observation d95e123c-cefd-437d-979e-8175a6a61e12 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:21:20.345349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:6c6da132fc6d2a328c315ed89093433d43ab036793c83c9cc4c2a6e95afdc206

Observation c3b6d048-6406-4d07-b0de-7542497de3ab · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T11:26:04.734541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:6a6a53358f90db48a9753b9a01fb543d5fb2596a4eb092a349b312d8588d4384

Observation 5a63f6d2-3f06-4cc1-88fc-82de48b8aa5d · outbound

This paper cites The Art of Scaling Reinforcement Learning Compute for LLMs.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.085942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:b7baca03a0b9a7298f33aefa1d4dcc09ac0e4f592e1028b0e0a9fc05f929f333

Observation ab27407c-c671-4305-986c-f5f44e43a829 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Solving Quantitative Reasoning Problems with Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T22:43:59.458883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:80ce85f67a91f70074bffe1cba8df625a1f46c9fd8bf99eed8836834accdd590

Observation 0a107278-ab2f-4440-b8ad-b167c0ea7ae2 · outbound

This paper cites Let's Verify Step by Step.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Let's Verify Step by Step

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:26:04.724691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:de1754f58406180e119928ca58b99491919e5c0c82e9855092a04fd1e5b1f8f9

Observation 8ba900d0-bf9d-43fb-add6-c3263eed1c68 · outbound

This paper cites When speed kills stability: Demystifying RL collapse from the training-inference mismatch, sep 2025.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO When speed kills stability: Demystifying RL collapse from the training-inference mismatch, sep 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:21:20.324768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:99308c38a24aefc32cd77330e30774bad006e6449040efb93e86040fb1f85df5

Observation b1e7bdfa-110b-4308-8d7e-8e5e2340dc99 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Understanding R1-Zero-Like Training: A Critical Perspective

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:26:04.738791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:fd4147df4245080306c5f9dc3c87aec27e1bee0c5690c112368d7b022cf34a71

Observation 341b9cd6-8374-41d2-8e61-dd5f19fbd4c8 · outbound

This paper cites Part i: Tricks or traps? a deep dive into rl for llm reasoning.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Part i: Tricks or traps? a deep dive into rl for llm reasoning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:04.667651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:518b81184b88c5a0fc82e384b6b40162d8d59db9e4c64726a060028df48d2219

Observation cb09adba-3b2b-4458-a298-8f46e1910865 · outbound

This paper cites Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:04.671207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:4735196bd453f81be1191a4e440e11bee727875400c14553acbb8f8d9afe3244

Observation b3e036a8-8ebb-43b7-bced-178e70046601 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:7fbf4f2d1348d6f38dc48bf0bce02d2ba8fd69edf1f5881341d6271d9bdbb89d

Observation 7044776b-ea3f-4f8e-bf8e-deecf929a894 · outbound

This paper cites OpenAI o1 System Card.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO OpenAI o1 System Card

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:26:04.647479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:59a97a341b56ada1dd539181a5bdf52cee3523c5327e68a33c6b5a2a0ac1e5b9

Observation 39166dec-8cb5-46c4-9d29-5926cd8c6210 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Proximal Policy Optimization Algorithms

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:26:04.683005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:32d55df1d4fd1b9caf08165151d6da1464a6b21e31498ef9a252850464a68610

Observation cbfb241c-9b98-4775-be60-bda02c02e715 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:26:04.689525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:ffc24dcb6297550f2ab4ae54859ab5ded1ca8a22d9e8042304f902ac488bacc4

Observation c784cbe9-8f83-4373-a686-8bb6783d64bc · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Opencompass: A universal evaluation platform for foundation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:21:20.355195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:0f8f603bc61893a6d8e9dd5590ca1c11068a39fc6ec74167f567b982efb7a8b2

Observation b92f68a3-4c5a-43b4-a0bc-3aa3aa6c15f0 · outbound

This paper cites CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:04.658739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:5a796c86b49dc95bcf9ac067331dc96cb049513d52d0dbb929cd635ff2df9f8b

Observation 25cb378c-f7eb-4da7-8bf1-4bccb1395ec8 · outbound

This paper cites When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:00:25.356150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:32026ad0f09923e5d08e3c855b6dd516ce592e9adae9e21de2787b4ae9d9c5b7

Observation 91f37ab9-165c-4042-9941-65b18def4bbd · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:26:04.730699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:3aad1f8689f8133104a85602e349795d7e0a730b73e61702714261dc41be77b1

Observation 4cb9fe47-170a-4206-a6e0-9ab33945155a · outbound

This paper cites Qwen3 Technical Report.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Qwen3 Technical Report

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:26:04.634949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:40562b4b7987b2a17465965602cb103bce6fcc88f9747f2e8365a29ff3d711bf

Observation 0a9b4965-476c-4916-ad3b-697a13ac4864 · outbound

This paper cites Your efficient rl framework secretly brings you off-policy rl training, aug 2025.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Your efficient rl framework secretly brings you off-policy rl training, aug 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:21:20.337051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:e68c4d6202b62b34ad51e0b9f891bba6dab5c595d0fd0bff6d847ddbbccb6668

Observation bce42b06-c2ac-4fea-90e5-3f3fb739f9da · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:26:04.639218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:dcf2248c4a79172b932d0d0161e4aeecbba08fc8d81061d7e892a56d381b79e7

Observation c7ae73c4-ac59-494e-99c4-9fad838e55d0 · outbound

This paper cites Scaling of search and learning: A roadmap to reproduce o1 from reinforcement learning perspective.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Scaling of search and learning: A roadmap to reproduce o1 from reinforcement learning perspective

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:21:20.328972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:8c7a16926f142b3cf717a4d342f00ad94e6633df5fc3370b093d0f06dd1c0dce

Observation 417f197f-201b-43ff-a2f1-911c9467be71 · outbound

This paper cites Geometric-mean policy optimization.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Geometric-mean policy optimization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:21:20.332965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:41a679e6353026a2b48e11e07489d9da88f87f95731faedb88b4d6756d96ee1e

Observation c9207f86-320b-46a1-858b-26b9e0ab66fc · outbound

This paper cites Geometric-mean policy optimization.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Geometric-mean policy optimization

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:04.616905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:8c4119db716013a6a4663a38d962968a299e26d094e9d6cc472506ad6931d7cc

Observation 903fd567-d6bc-45b4-b150-5dcadcd2cccc · outbound

This paper cites Group Sequence Policy Optimization.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Group Sequence Policy Optimization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:26:04.630684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:0440b3663209180282cb085654e0f12c7a264353848e04a94c1856f897e6690d

Observation f27d840a-d571-415c-93ae-f940d582561b · outbound

This paper cites Rloop: An self-improving framework for reinforcement learning with iterative policy initialization.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO Rloop: An self-improving framework for reinforcement learning with iterative policy initialization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T11:21:20.341335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:ab02f9c9f8309293592561752d2251acf5bb077bf84d687477acf5835edeb73c

Observation 9788758c-caa1-4f6d-84f0-c343fb9e6a04 · outbound

This paper cites 14 Appendix A: Why Use Sequence-Count Weights in BA? Here we briefly justify the choice of weights k/G and (G−k)/G in BA.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO 14 Appendix A: Why Use Sequence-Count Weights in BA? Here we briefly justify the choice of weights k/G and (G−k)/G in BA

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:04.624712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:d4f20acb958c944848d10e6d3d4d7bc3b00c5bb91eec43fa74c30ca746166f43

Pith citing papers

Observation 419567d6-c8b2-4ee9-907b-633a9b827c4d · inbound

Multimodal Reward Hacking in Reinforcement Learning cites this paper.

Multimodal Reward Hacking in Reinforcement Learning Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:5572dfc9135e8aff8fe7bbb53371923437cdb9dda6bef22db9c77dee862923aa