Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2309.06657.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:05.396920Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T01:17:30.660177Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation c33e5668-b295-4b15-9dba-830b5acee434 · inbound
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Statistical Rejection Sampling Improves Preference Optimization
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 10740686-15de-458a-8a08-82ef54fc1c73 · inbound
Learning a Pessimistic Reward Model in RLHF Statistical Rejection Sampling Improves Preference Optimization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d2fe920-b1a1-4ac1-acf1-4f5e463db2ba · inbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Statistical Rejection Sampling Improves Preference Optimization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41e00320-9fdb-472b-a391-488796e832a0 · inbound
Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy Statistical Rejection Sampling Improves Preference Optimization
Reference 104
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e9a7c7f-6d8b-48bb-a13c-f3dd10c93cb9 · inbound
Thompson Sampling in Online RLHF with General Function Approximation Statistical Rejection Sampling Improves Preference Optimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d9f8622-e1c9-452a-a71b-742cf211ddd4 · inbound
Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences Statistical Rejection Sampling Improves Preference Optimization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c055b56-2701-4ae8-8fbb-2fc9eba799b1 · inbound
ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization Statistical Rejection Sampling Improves Preference Optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d68b4f96-d8a1-4c78-a8f5-74ec1ec203de · inbound
Bridging Offline and Online Reinforcement Learning for LLMs Statistical Rejection Sampling Improves Preference Optimization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8345db02-524c-42cc-9e93-99dfba1bf963 · inbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Statistical Rejection Sampling Improves Preference Optimization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 092382b6-2b8c-42d5-9650-e3796259d9ac · inbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Statistical Rejection Sampling Improves Preference Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8b226be-80c6-4ec8-90d4-8fd3d7beda10 · inbound
PKG-DPO: Optimizing Domain-Specific AI systems with Physics Knowledge Graphs and Direct Preference Optimization Statistical Rejection Sampling Improves Preference Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72979ff6-f349-44ed-a7d6-d56854c786b7 · inbound
Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving Statistical Rejection Sampling Improves Preference Optimization
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ea8dbbf-0c3d-4d0c-b1c7-fb18f039d4e7 · inbound
SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance Statistical Rejection Sampling Improves Preference Optimization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0fdf0e4a-ba7e-46c6-9975-7e81f73d13b9 · inbound
Beyond Importance Sampling: Rejection-Gated Policy Optimization Statistical Rejection Sampling Improves Preference Optimization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2a2cfc4-c3ad-45dd-89c0-3094a38f3b29 · inbound
Reasoning Structure Matters for Safety Alignment of Reasoning Models Statistical Rejection Sampling Improves Preference Optimization
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 37f56f6a-d22a-466c-b5aa-bcf3033f0566 · inbound
HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs Statistical Rejection Sampling Improves Preference Optimization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50a5ae45-660c-4b1b-b769-58288a06f74b · inbound
Supplement Generation Training for Enhancing Agentic Task Performance Statistical Rejection Sampling Improves Preference Optimization
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f1a64bc-0b13-4256-9573-312652e6000e · inbound
Efficient Preference Poisoning Attack on Offline RLHF Statistical Rejection Sampling Improves Preference Optimization
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 359e0d48-73cc-4d37-a616-09c67fa5c344 · inbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Statistical Rejection Sampling Improves Preference Optimization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8070190b-b727-4a74-b125-30ae1a237423 · inbound
Response Time Enhances Alignment with Heterogeneous Preferences Statistical Rejection Sampling Improves Preference Optimization
Reference 166
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 562b6676-0d06-48a8-beef-5182bb4a7395 · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Statistical Rejection Sampling Improves Preference Optimization
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8df7b4cb-a5d7-49bb-8496-79b847b3a2a6 · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Statistical Rejection Sampling Improves Preference Optimization
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83ac0f78-88ac-4817-9194-7682d793bd13 · inbound
Gradient-Guided Reward Optimization for Inference-time Alignment Statistical Rejection Sampling Improves Preference Optimization
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83c5d7a1-e3bf-44bd-afd3-774f0fea94e4 · inbound
BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Statistical Rejection Sampling Improves Preference Optimization
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d44c5aba-3f08-4bb6-b0d3-ca20f5df1148 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Statistical Rejection Sampling Improves Preference Optimization
Reference 261
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c2d654c-71e1-4dc5-b084-4cc6390e2569 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Statistical Rejection Sampling Improves Preference Optimization
Reference 262
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bdf339d-0117-4b99-ad3b-f719cd719c69 · inbound
Test-Time Scaling via Error Localization Statistical Rejection Sampling Improves Preference Optimization
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.