Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T18:10:53.458422Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2502.05773.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T18:10:53.458422Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a869fe09-2ea4-4a9f-a322-c1eff8008cb4 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Learning from negative feedback, or positive feedback or both
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db981569-bbae-4bbb-9a6c-4b523ed8decb · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation AlphaMath Almost Zero: Process Supervision without Process
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c8db598-fd29-4a54-ac1d-4ffa02a74aa8 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation The Llama 3 Herd of Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb4e99f-bfdf-4464-8ff4-4f5022b07e01 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation A density estimation perspective on learning from pairwise human preferences
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42ec5816-5684-40d8-8f3b-ddce6bd3e19d · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation KTO: Model Alignment as Prospect Theoretic Optimization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1c35814-267b-4b9e-9abe-6ac178bfb733 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb67e40f-4cf7-42ae-8179-a6522be88ea6 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0267fe8-794e-4583-8a94-005c3ab57fbd · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Measuring Mathematical Problem Solving With the MATH Dataset
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae9d5c95-2e19-4a76-95f1-22213100fe36 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation LoRA: Low-Rank Adaptation of Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbbde740-5f70-4ffc-9d45-157e1e11084d · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 184611e2-ca00-49f0-9320-c3fe399d533d · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation A Distributional Approach to Controlled Text Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c620067a-5469-4982-bbf5-665206007a51 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Let's Verify Step by Step
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 633c2580-ae6d-44b5-8df8-b43e7dba58a0 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb198fad-5d20-4ba4-8b02-351a13ba207f · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01f88fa3-c77b-44a4-9d1a-f45e326b74e3 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52924448-f3bd-467e-863b-c2a89593b9da · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9b4352dd-93a6-40ec-bb58-8b885156257f · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Iterative Reasoning Preference Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c873275-e653-42f2-a8b7-2bd1fb2da73f · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Distributional Reinforcement Learning for Energy-Based Sequential Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6ca434a0-5e97-4251-bae6-9b0c3d111987 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99303e61-f71a-423f-bb96-2215a3635f4d · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Proximal Policy Optimization Algorithms
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38788a5f-748a-40cb-8c14-7866e2068cf4 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Generalized Preference Optimization: A Unified Approach to Offline Alignment
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b599f16-a305-4fc8-9721-66e3ce66ec3c · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Offline Reinforcement Learning for LLM Multi-Step Reasoning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4169deec-8b4d-4d56-b262-25b42c7f9643 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 404700b7-9a2b-4d46-9a65-bf8dc02287f0 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Token-level Direct Preference Optimization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c915cf3-30ac-4f5a-a29b-5b1223555aed · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0247b814-5dab-4ed2-8f03-997e2df12d94 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c916166-c222-4c93-a4fe-a1c463add787 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50182318-2c61-450f-9ba8-3b7aaa1fd591 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Treating the sequences as a whole, the original DPO loss is given by LDPO(x, y+, y−, c+, c−) =− log σ X t rt(x, y+) − X t rt(x, y−) !
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7d34b587-4f78-4d23-ab71-4222eaa0414b · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa0fa2ab-c8a0-4e6d-b5be-63ace50658b6 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation UI-TARS: Pioneering Automated GUI Interaction with Native Agents
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9192cd6e-6866-4b1a-8956-9477bd06f053 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation RLHF Workflow: From Reward Modeling to Online RLHF
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f4307c2-cea9-4983-a12e-5a00a1d79f35 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb563930-792b-4ab2-b5c8-1eef0d77164f · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d47ed88-ee14-4d31-a401-0e3114e2f2d0 · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation GPT-4 Technical Report
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5eb9dc4-4f7a-4afb-96da-16520c492d9f · outbound
PIPA: Preference Alignment as Prior-Informed Statistical Estimation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.