Pith. sign in

Paper Citation Record · LEDGER

BNPO: Beta Normalization Policy Optimization

As of 22 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 11 inbound Pith citation observations for arXiv:2506.02864.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02864 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:20:55.897193Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:49:07.612952Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.713683Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0294f35-78d3-4939-b50c-55a588b1b7b6 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

BNPO: Beta Normalization Policy Optimization Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.811097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.811097Z digest=sha256:e7119784a7d80bddfb41e161b3962dea184befdd8393fee8e100097f5085124e

Observation 5bb22eda-c1b8-48e3-bead-b82e5b41d5f3 · outbound

This paper cites Aime problems and solutions, 2025 a.

BNPO: Beta Normalization Policy Optimization Aime problems and solutions, 2025 a

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:56.129057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:20:55.816379Z digest=sha256:5d42d3c200a957b65976b97404598272a3a76966a939b140792ec4692d1f6a86

Observation 4cf3e0aa-173e-432d-9620-05a096a9706d · outbound

This paper cites Amc problems and solutions, 2025 b.

BNPO: Beta Normalization Policy Optimization Amc problems and solutions, 2025 b

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:56.115926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:20:55.821455Z digest=sha256:2c62ee675714f403b56dcc4256844adb5efccddb454270e7b1a30872c8fa5122

Observation 8aff3bcc-1730-4ccb-b064-900e97e15618 · outbound

This paper cites Reinforcement learning: An introduction.

BNPO: Beta Normalization Policy Optimization Reinforcement learning: An introduction

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:56.102198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:20:55.826091Z digest=sha256:98de776cdd21c191bf7717410b8c678293c46e8388078554d2661a8e9c87689d

Observation a3d766c9-e77f-4c17-9a74-f67746ab7989 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

BNPO: Beta Normalization Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.831118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.831118Z digest=sha256:8a5866d74e10b5e94f079dd7b64e591e5fb69828f28bdef2d33e9f4c361273e2

Observation 4e04a79b-5b7a-4e7b-97f0-cf4fb1bea277 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

BNPO: Beta Normalization Policy Optimization Measuring Mathematical Problem Solving With the MATH Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.836739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.836739Z digest=sha256:e573980d0d7cc8d33324a4bf5af3f76324be5debbaeddda5390b8b6b2b89416b

Observation a927a665-0fdc-42de-985a-abdf099b8613 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

BNPO: Beta Normalization Policy Optimization REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.841916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.841916Z digest=sha256:f0cd4b0911b4b0484780370263850f66a87096ca3ddd46f2f456d1f62aad2a89

Observation 5de13090-1d38-48ba-8fcc-448e26f3e692 · outbound

This paper cites Buy 4 REINFORCE samples, get a baseline for free!, 2019.

BNPO: Beta Normalization Policy Optimization Buy 4 REINFORCE samples, get a baseline for free!, 2019

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:56.090147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:20:55.846031Z digest=sha256:6688adac86511f7a7f784c9f5b9b51da213dcb9b933694e6420658d211590d0b

Observation 4f0ba18f-0091-48a6-ad62-8ea5cfe1b05d · outbound

This paper cites Remax: a simple, effective, and efficient reinforcement learning method for aligning large language models.

BNPO: Beta Normalization Policy Optimization Remax: a simple, effective, and efficient reinforcement learning method for aligning large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:56.077493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T11:20:55.850087Z digest=sha256:1e91a62effece624b5939c7972f44f242885850e523977417073ebeee502c94f

Observation ddf762d4-f6a2-4d57-9851-b2fbc931dd95 · outbound

This paper cites Let's verify step by step.

BNPO: Beta Normalization Policy Optimization Let's verify step by step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.854090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.854090Z digest=sha256:e0bf48f74a594d7a469485fcd3b085caa4982ad968085e7240a5779ec16e4788

Observation b169dba6-9989-489d-be39-f451faea0e9d · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

BNPO: Beta Normalization Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.857749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.857749Z digest=sha256:5a222a0efeed8df3092a95d943253848350b73265e5553ccb00055c076fbef1e

Observation 1744213d-6b24-47be-9fbf-aa287f875d3d · outbound

This paper cites Training language models to follow instructions with human feedback.

BNPO: Beta Normalization Policy Optimization Training language models to follow instructions with human feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.862927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.862927Z digest=sha256:ad868c5ca678248a42cc6ffd2237d3281760fbe08af26da8127d4459235d4399

Observation 00bdef7b-6df4-4c50-8fe0-14e33526e5ff · outbound

This paper cites Proximal Policy Optimization Algorithms.

BNPO: Beta Normalization Policy Optimization Proximal Policy Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.866604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.866604Z digest=sha256:95bc83389a8845ee3547eb5e147f2ed11d8dbf9adab30af674d6a241735dc4c0

Observation 1a3cc03e-4159-4fc3-bc86-9771e7e771c6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

BNPO: Beta Normalization Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.871831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.871831Z digest=sha256:c398fc439e5ae8576e6eeb9a911516d099b474d98690f4143e29bb65453b16a1

Observation 849dfdd8-d8fc-4af0-924d-a6f58e89a1c0 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

BNPO: Beta Normalization Policy Optimization Policy gradient methods for reinforcement learning with function approximation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.875943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.875943Z digest=sha256:e80a1305697a622ea6654d84c937e5bf6e4e97720eb3f696c76d79e5a50b2efb

Observation 25fac9b9-e53f-42c0-9c0e-9378c7b3d8fc · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

BNPO: Beta Normalization Policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.879621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.879621Z digest=sha256:8c60779e2cb0db565dfdf46a31424c84af87b7852588801b6011316268a9e5f2

Observation 9c124978-7fbc-4b94-898d-225fdd9387cd · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

BNPO: Beta Normalization Policy Optimization Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.883608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.883608Z digest=sha256:0bc48279078bf0960c63e8fccfa211c0441b6d4a6ce56d9f5559c5bcc2ae270d

Observation 23278fa2-de70-44cc-a742-d0ec0fd671e4 · outbound

This paper cites Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines.

BNPO: Beta Normalization Policy Optimization Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.887365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.887365Z digest=sha256:d79dfccf353271d803b5e08d68e722b0c1f25137e7e46b2d0416491872c96d2e

Observation 4dc3a996-6083-4363-92a6-930f3c91fae4 · outbound

This paper cites Qwen2.5 Technical Report.

BNPO: Beta Normalization Policy Optimization Qwen2.5 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.892782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.892782Z digest=sha256:df5c5b60ce143f7c9cb42ad23c87a0a4f121a6140d4c62e0a1742f666af98872

Observation 3ea6da40-e637-4f7b-8601-bd66849da427 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

BNPO: Beta Normalization Policy Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.897193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.897193Z digest=sha256:c244172ae0a753c9c0a385a9de70760f5cba943e13613138f536eb2c1117802f

Pith citing papers

Observation be2d50b4-5f92-4b7b-9444-e13a3ae5bbe5 · inbound

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems cites this paper.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems BNPO: Beta Normalization Policy Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.612952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.612952Z digest=sha256:ed0a61a32be7af2c0c49ead769e4166be2f79a689eca9d2d7b07903a4704625a

Observation 340567cb-e213-49c1-938d-5e13ac4c09de · inbound

CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization cites this paper.

CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization BNPO: Beta Normalization Policy Optimization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:26.497930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T00:59:44.364491Z digest=sha256:ab3b94bb3b8ae5a4d810e9b1fbcd3ccb5027d115c6ca71717bb2ac88af932583

Observation 1e2cb5c6-f008-41d8-9cdf-d57deb0f46ce · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation BNPO: Beta Normalization Policy Optimization

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:12:22.759733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T06:08:28.855671Z digest=sha256:e0af9c86febcc8d9e63732a0698700eaffb09ef39b914cf998f86c69dfb8111d

Observation b9e69fa3-e40f-4cfb-8334-32dbe2cca0f6 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation BNPO: Beta Normalization Policy Optimization

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.160661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:96332a498400db353c19c8d9ba916badab2cb26077a6280b845d4bc79968bb42

Observation 64639df8-7838-4fb5-8962-e9ea3bb461ae · inbound

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes cites this paper.

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes BNPO: Beta Normalization Policy Optimization

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T11:53:23.678846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T11:51:01.769882Z digest=sha256:20ddf7894bfe08dce35de267192bd1c8d9521000ac1a3598458edfad0aa25508

Observation 7fc63523-d7e1-434f-b8d7-ad74c1ad31a5 · inbound

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes cites this paper.

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes BNPO: Beta Normalization Policy Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T12:58:10.962452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:58:10.962452Z digest=sha256:b0ea23a472ec70d95ff46428b63d77d0c1ef60f0c2b7d9255872118425f8369f

Observation 1ba334c9-db92-4101-be60-18993d05933c · inbound

GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards cites this paper.

GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards BNPO: Beta Normalization Policy Optimization

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:36:47.644321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T05:59:00.005336Z digest=sha256:97ac9b7c3f5c23ceb4b70e854facb937b46068f19131caa4e81e5bb79ee91071

Observation 66ffc8c9-8175-4f9d-a9e2-a24b01b993b5 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning BNPO: Beta Normalization Policy Optimization

Reference 232

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.715150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:9b2fe4f6af7d74ed1621ac28b2758af5d67d07a7daff42488ff9856b8f12e4e2

Observation f0105c2f-0ea7-4100-808a-26d1de6910b5 · inbound

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF cites this paper.

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF BNPO: Beta Normalization Policy Optimization

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:04:29.019093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-30T07:36:39.835616Z digest=sha256:ecd6fdc338f2cdeb1b89eeab1d6cc920bc2288ac14377836050fbd1960fdc043

Observation 93c0b0fa-9621-45d6-8aef-b56b0f3682c0 · inbound

Aligning Language Models with Selective Prediction cites this paper.

Aligning Language Models with Selective Prediction BNPO: Beta Normalization Policy Optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T01:51:25.883463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:51:25.883463Z digest=sha256:0a4e64e264244925e04c5985618b2e518029a38cfc4b0e1c4c4832c6676780a3

Observation 5d5f2df7-0534-44ca-88d3-6e66d96480bd · inbound

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works cites this paper.

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works BNPO: Beta Normalization Policy Optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T08:01:23.300594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:01:23.300594Z digest=sha256:4859a379d03481d85d7a9e808ce4a149b6c0ce429d22ee2641b8ade77be96c24