Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:20:55.897193Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 11 inbound Pith citation observations for arXiv:2506.02864.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:20:55.897193Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:49:07.612952Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:09:40.713683Z
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d0294f35-78d3-4939-b50c-55a588b1b7b6 · outbound
BNPO: Beta Normalization Policy Optimization Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bb22eda-c1b8-48e3-bead-b82e5b41d5f3 · outbound
BNPO: Beta Normalization Policy Optimization Aime problems and solutions, 2025 a
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4cf3e0aa-173e-432d-9620-05a096a9706d · outbound
BNPO: Beta Normalization Policy Optimization Amc problems and solutions, 2025 b
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8aff3bcc-1730-4ccb-b064-900e97e15618 · outbound
BNPO: Beta Normalization Policy Optimization Reinforcement learning: An introduction
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a3d766c9-e77f-4c17-9a74-f67746ab7989 · outbound
BNPO: Beta Normalization Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e04a79b-5b7a-4e7b-97f0-cf4fb1bea277 · outbound
BNPO: Beta Normalization Policy Optimization Measuring Mathematical Problem Solving With the MATH Dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a927a665-0fdc-42de-985a-abdf099b8613 · outbound
BNPO: Beta Normalization Policy Optimization REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5de13090-1d38-48ba-8fcc-448e26f3e692 · outbound
BNPO: Beta Normalization Policy Optimization Buy 4 REINFORCE samples, get a baseline for free!, 2019
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4f0ba18f-0091-48a6-ad62-8ea5cfe1b05d · outbound
BNPO: Beta Normalization Policy Optimization Remax: a simple, effective, and efficient reinforcement learning method for aligning large language models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ddf762d4-f6a2-4d57-9851-b2fbc931dd95 · outbound
BNPO: Beta Normalization Policy Optimization Let's verify step by step
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b169dba6-9989-489d-be39-f451faea0e9d · outbound
BNPO: Beta Normalization Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1744213d-6b24-47be-9fbf-aa287f875d3d · outbound
BNPO: Beta Normalization Policy Optimization Training language models to follow instructions with human feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00bdef7b-6df4-4c50-8fe0-14e33526e5ff · outbound
BNPO: Beta Normalization Policy Optimization Proximal Policy Optimization Algorithms
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a3cc03e-4159-4fc3-bc86-9771e7e771c6 · outbound
BNPO: Beta Normalization Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 849dfdd8-d8fc-4af0-924d-a6f58e89a1c0 · outbound
BNPO: Beta Normalization Policy Optimization Policy gradient methods for reinforcement learning with function approximation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25fac9b9-e53f-42c0-9c0e-9378c7b3d8fc · outbound
BNPO: Beta Normalization Policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c124978-7fbc-4b94-898d-225fdd9387cd · outbound
BNPO: Beta Normalization Policy Optimization Simple statistical gradient-following algorithms for connectionist reinforcement learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23278fa2-de70-44cc-a742-d0ec0fd671e4 · outbound
BNPO: Beta Normalization Policy Optimization Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dc3a996-6083-4363-92a6-930f3c91fae4 · outbound
BNPO: Beta Normalization Policy Optimization Qwen2.5 Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ea6da40-e637-4f7b-8601-bd66849da427 · outbound
BNPO: Beta Normalization Policy Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be2d50b4-5f92-4b7b-9444-e13a3ae5bbe5 · inbound
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems BNPO: Beta Normalization Policy Optimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 340567cb-e213-49c1-938d-5e13ac4c09de · inbound
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization BNPO: Beta Normalization Policy Optimization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1e2cb5c6-f008-41d8-9cdf-d57deb0f46ce · inbound
Holder Policy Optimisation BNPO: Beta Normalization Policy Optimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b9e69fa3-e40f-4cfb-8334-32dbe2cca0f6 · inbound
Holder Policy Optimisation BNPO: Beta Normalization Policy Optimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 64639df8-7838-4fb5-8962-e9ea3bb461ae · inbound
DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes BNPO: Beta Normalization Policy Optimization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7fc63523-d7e1-434f-b8d7-ad74c1ad31a5 · inbound
DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes BNPO: Beta Normalization Policy Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ba334c9-db92-4101-be60-18993d05933c · inbound
GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards BNPO: Beta Normalization Policy Optimization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 66ffc8c9-8175-4f9d-a9e2-a24b01b993b5 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning BNPO: Beta Normalization Policy Optimization
Reference 232
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f0105c2f-0ea7-4100-808a-26d1de6910b5 · inbound
PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF BNPO: Beta Normalization Policy Optimization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 93c0b0fa-9621-45d6-8aef-b56b0f3682c0 · inbound
Aligning Language Models with Selective Prediction BNPO: Beta Normalization Policy Optimization
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d5f2df7-0534-44ca-88d3-6e66d96480bd · inbound
The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works BNPO: Beta Normalization Policy Optimization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.