Pith. sign in

Paper Citation Record · LEDGER

BNPO: Beta Normalization Policy Optimization

As of 22 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 11 inbound Pith citation observations for arXiv:2506.02864.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02864 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:20:55.897193Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:49:07.612952Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.713683Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0294f35-78d3-4939-b50c-55a588b1b7b6 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

BNPO: Beta Normalization Policy Optimization Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.811097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.811097Z digest=sha256:e7119784a7d80bddfb41e161b3962dea184befdd8393fee8e100097f5085124e

Observation 5bb22eda-c1b8-48e3-bead-b82e5b41d5f3 · outbound

This paper cites Aime problems and solutions, 2025 a.

BNPO: Beta Normalization Policy Optimization Aime problems and solutions, 2025 a

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:56.129057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T11:20:55.816379Z digest=sha256:f4e62b75e72b7970c06c76443611e87a2fa679f6a43bb7d68d0986844be749b8

Observation 4cf3e0aa-173e-432d-9620-05a096a9706d · outbound

This paper cites Amc problems and solutions, 2025 b.

BNPO: Beta Normalization Policy Optimization Amc problems and solutions, 2025 b

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:56.115926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T11:20:55.821455Z digest=sha256:7544f3113f651b0e27833b382e704cb36f14071bb749d09ce6b5561f6ee7b6af

Observation 8aff3bcc-1730-4ccb-b064-900e97e15618 · outbound

This paper cites Reinforcement learning: An introduction.

BNPO: Beta Normalization Policy Optimization Reinforcement learning: An introduction

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:56.102198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T11:20:55.826091Z digest=sha256:88d6e822ac544b5f6aca79790bb1e880c0f509f770c44eb5b95a8762075688cb

Observation a3d766c9-e77f-4c17-9a74-f67746ab7989 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

BNPO: Beta Normalization Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.831118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.831118Z digest=sha256:8a5866d74e10b5e94f079dd7b64e591e5fb69828f28bdef2d33e9f4c361273e2

Observation 4e04a79b-5b7a-4e7b-97f0-cf4fb1bea277 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

BNPO: Beta Normalization Policy Optimization Measuring Mathematical Problem Solving With the MATH Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.836739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.836739Z digest=sha256:e573980d0d7cc8d33324a4bf5af3f76324be5debbaeddda5390b8b6b2b89416b

Observation a927a665-0fdc-42de-985a-abdf099b8613 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

BNPO: Beta Normalization Policy Optimization REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.841916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.841916Z digest=sha256:f0cd4b0911b4b0484780370263850f66a87096ca3ddd46f2f456d1f62aad2a89

Observation 5de13090-1d38-48ba-8fcc-448e26f3e692 · outbound

This paper cites Buy 4 REINFORCE samples, get a baseline for free!, 2019.

BNPO: Beta Normalization Policy Optimization Buy 4 REINFORCE samples, get a baseline for free!, 2019

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:56.090147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T11:20:55.846031Z digest=sha256:142c2c4f9fbeb4761401eca478e22723b185d80d6b612956770a57128a07a84f

Observation 4f0ba18f-0091-48a6-ad62-8ea5cfe1b05d · outbound

This paper cites Remax: a simple, effective, and efficient reinforcement learning method for aligning large language models.

BNPO: Beta Normalization Policy Optimization Remax: a simple, effective, and efficient reinforcement learning method for aligning large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:20:56.077493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T11:20:55.850087Z digest=sha256:d3fa0e6267d8c539039b596ca2f658ce95bc658f614d1f1de2cbf729f63cddcc

Observation ddf762d4-f6a2-4d57-9851-b2fbc931dd95 · outbound

This paper cites Let's verify step by step.

BNPO: Beta Normalization Policy Optimization Let's verify step by step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.854090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.854090Z digest=sha256:e0bf48f74a594d7a469485fcd3b085caa4982ad968085e7240a5779ec16e4788

Observation b169dba6-9989-489d-be39-f451faea0e9d · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

BNPO: Beta Normalization Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.857749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.857749Z digest=sha256:5a222a0efeed8df3092a95d943253848350b73265e5553ccb00055c076fbef1e

Observation 1744213d-6b24-47be-9fbf-aa287f875d3d · outbound

This paper cites Training language models to follow instructions with human feedback.

BNPO: Beta Normalization Policy Optimization Training language models to follow instructions with human feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.862927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.862927Z digest=sha256:ad868c5ca678248a42cc6ffd2237d3281760fbe08af26da8127d4459235d4399

Observation 00bdef7b-6df4-4c50-8fe0-14e33526e5ff · outbound

This paper cites Proximal Policy Optimization Algorithms.

BNPO: Beta Normalization Policy Optimization Proximal Policy Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.866604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.866604Z digest=sha256:95bc83389a8845ee3547eb5e147f2ed11d8dbf9adab30af674d6a241735dc4c0

Observation 1a3cc03e-4159-4fc3-bc86-9771e7e771c6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

BNPO: Beta Normalization Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.871831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.871831Z digest=sha256:c398fc439e5ae8576e6eeb9a911516d099b474d98690f4143e29bb65453b16a1

Observation 849dfdd8-d8fc-4af0-924d-a6f58e89a1c0 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

BNPO: Beta Normalization Policy Optimization Policy gradient methods for reinforcement learning with function approximation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.875943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.875943Z digest=sha256:e80a1305697a622ea6654d84c937e5bf6e4e97720eb3f696c76d79e5a50b2efb

Observation 25fac9b9-e53f-42c0-9c0e-9378c7b3d8fc · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

BNPO: Beta Normalization Policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.879621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.879621Z digest=sha256:8c60779e2cb0db565dfdf46a31424c84af87b7852588801b6011316268a9e5f2

Observation 9c124978-7fbc-4b94-898d-225fdd9387cd · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

BNPO: Beta Normalization Policy Optimization Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.883608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.883608Z digest=sha256:0bc48279078bf0960c63e8fccfa211c0441b6d4a6ce56d9f5559c5bcc2ae270d

Observation 23278fa2-de70-44cc-a742-d0ec0fd671e4 · outbound

This paper cites Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines.

BNPO: Beta Normalization Policy Optimization Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.887365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.887365Z digest=sha256:d79dfccf353271d803b5e08d68e722b0c1f25137e7e46b2d0416491872c96d2e

Observation 4dc3a996-6083-4363-92a6-930f3c91fae4 · outbound

This paper cites Qwen2.5 Technical Report.

BNPO: Beta Normalization Policy Optimization Qwen2.5 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.892782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.892782Z digest=sha256:df5c5b60ce143f7c9cb42ad23c87a0a4f121a6140d4c62e0a1742f666af98872

Observation 3ea6da40-e637-4f7b-8601-bd66849da427 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

BNPO: Beta Normalization Policy Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:55.897193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:55.897193Z digest=sha256:c244172ae0a753c9c0a385a9de70760f5cba943e13613138f536eb2c1117802f

Pith citing papers

Observation be2d50b4-5f92-4b7b-9444-e13a3ae5bbe5 · inbound

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems cites this paper.

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems BNPO: Beta Normalization Policy Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:07.612952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:07.612952Z digest=sha256:ed0a61a32be7af2c0c49ead769e4166be2f79a689eca9d2d7b07903a4704625a

Observation 340567cb-e213-49c1-938d-5e13ac4c09de · inbound

CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization cites this paper.

CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization BNPO: Beta Normalization Policy Optimization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:26.497930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T00:59:44.364491Z digest=sha256:6ddbb5e496a4aeca6cf0a82b12a25ef89f09f85126b5487b10132c1a34708b09

Observation 1e2cb5c6-f008-41d8-9cdf-d57deb0f46ce · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation BNPO: Beta Normalization Policy Optimization

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:12:22.759733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T06:08:28.855671Z digest=sha256:b824392d68add3c50c0e0fc59d40033bd1a6d123cdb098833a63bada7ae22ce5

Observation b9e69fa3-e40f-4cfb-8334-32dbe2cca0f6 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation BNPO: Beta Normalization Policy Optimization

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.160661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:930d615298061e854f52a761973b1de454682d3bbd9796d348c54ce3789ca66c

Observation 64639df8-7838-4fb5-8962-e9ea3bb461ae · inbound

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes cites this paper.

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes BNPO: Beta Normalization Policy Optimization

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T11:53:23.678846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T11:51:01.769882Z digest=sha256:19721bb2ab3c05fdbcd79b3ef5ff2caad55efc59434e8c40ce27260a49693a53

Observation 7fc63523-d7e1-434f-b8d7-ad74c1ad31a5 · inbound

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes cites this paper.

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes BNPO: Beta Normalization Policy Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T12:58:10.962452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:58:10.962452Z digest=sha256:b0ea23a472ec70d95ff46428b63d77d0c1ef60f0c2b7d9255872118425f8369f

Observation 1ba334c9-db92-4101-be60-18993d05933c · inbound

GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards cites this paper.

GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards BNPO: Beta Normalization Policy Optimization

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:36:47.644321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T05:59:00.005336Z digest=sha256:392f028cad54b2ca5416b7e31bdefa5d20b82ba5d74604fc5b8c12be7ab8853c

Observation 66ffc8c9-8175-4f9d-a9e2-a24b01b993b5 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning BNPO: Beta Normalization Policy Optimization

Reference 232

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.715150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:149dc69dd94489f0484f39d825629d870b495f717484486ff15c436a441f118c

Observation f0105c2f-0ea7-4100-808a-26d1de6910b5 · inbound

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF cites this paper.

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF BNPO: Beta Normalization Policy Optimization

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:04:29.019093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T07:36:39.835616Z digest=sha256:c81272b8bc1651986351d0a8b607f17f6dabc26f68aced28294ff2015e9cb32a

Observation 93c0b0fa-9621-45d6-8aef-b56b0f3682c0 · inbound

Aligning Language Models with Selective Prediction cites this paper.

Aligning Language Models with Selective Prediction BNPO: Beta Normalization Policy Optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T01:51:25.883463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:51:25.883463Z digest=sha256:0a4e64e264244925e04c5985618b2e518029a38cfc4b0e1c4c4832c6676780a3

Observation 5d5f2df7-0534-44ca-88d3-6e66d96480bd · inbound

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works cites this paper.

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works BNPO: Beta Normalization Policy Optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T08:01:23.300594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:01:23.300594Z digest=sha256:4859a379d03481d85d7a9e808ce4a149b6c0ce429d22ee2641b8ade77be96c24