Pith. sign in

Paper Citation Record · LEDGER

Statistical Rejection Sampling Improves Preference Optimization

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2309.06657.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.06657 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:05.396920Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:17:30.660177Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c33e5668-b295-4b15-9dba-830b5acee434 · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Statistical Rejection Sampling Improves Preference Optimization

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.289989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:5d10d11e9d3b003e94841fd90a4f0a28a6b1907776105835f7ec41d8434efdb2

Observation 10740686-15de-458a-8a08-82ef54fc1c73 · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Statistical Rejection Sampling Improves Preference Optimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.396920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.396920Z digest=sha256:1310bed9b1267f232abe9080a8b32bf83cbed355c1c5cb629fac9c40c148f1e8

Observation 9d2fe920-b1a1-4ac1-acf1-4f5e463db2ba · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Statistical Rejection Sampling Improves Preference Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.881471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.881471Z digest=sha256:abf25b1906ad4124a1bdcbd80068b99a27775bf2e5d1333532070d2fc6e0f464

Observation 41e00320-9fdb-472b-a391-488796e832a0 · inbound

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy cites this paper.

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy Statistical Rejection Sampling Improves Preference Optimization

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:30.648329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:30.648329Z digest=sha256:2e6a24ba1335bb857a86dec9eaf28d4891bb0c58579c6ae9df78034c9f238460

Observation 5e9a7c7f-6d8b-48bb-a13c-f3dd10c93cb9 · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation Statistical Rejection Sampling Improves Preference Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.431444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.431444Z digest=sha256:f87db01697fba5a46da1c10219220f66d6f09f9475bf368763a07f91e72cd1cb

Observation 8d9f8622-e1c9-452a-a71b-742cf211ddd4 · inbound

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences cites this paper.

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences Statistical Rejection Sampling Improves Preference Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:57.308215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:57.308215Z digest=sha256:bdfb1286c942dab10f5f7929482f32c39feb012298360eae8445f468bb89ee39

Observation 4c055b56-2701-4ae8-8fbb-2fc9eba799b1 · inbound

ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization cites this paper.

ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization Statistical Rejection Sampling Improves Preference Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:02.421271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:02.421271Z digest=sha256:0598244fce4655815e63ca342be0e904b30a8118f1ffb7d68700303d5adba2a0

Observation d68b4f96-d8a1-4c78-a8f5-74ec1ec203de · inbound

Bridging Offline and Online Reinforcement Learning for LLMs cites this paper.

Bridging Offline and Online Reinforcement Learning for LLMs Statistical Rejection Sampling Improves Preference Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.163268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.163268Z digest=sha256:1a2ca4bfad1b841de9da37a31c73a52585201c40fb891e55d2acbe6da4fbb46a

Observation 8345db02-524c-42cc-9e93-99dfba1bf963 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Statistical Rejection Sampling Improves Preference Optimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.092262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.092262Z digest=sha256:43692a0b1f85fd0fbcfcabcaaa5be896a5cb6237dadb31e133d81ea86a5c38e6

Observation 092382b6-2b8c-42d5-9650-e3796259d9ac · inbound

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future cites this paper.

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Statistical Rejection Sampling Improves Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:49.645169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:06:49.645169Z digest=sha256:db1ab553c44fb7091a11bdb038f9d3fe21a7e37eface278620340e8bf7667899

Observation b8b226be-80c6-4ec8-90d4-8fd3d7beda10 · inbound

PKG-DPO: Optimizing Domain-Specific AI systems with Physics Knowledge Graphs and Direct Preference Optimization cites this paper.

PKG-DPO: Optimizing Domain-Specific AI systems with Physics Knowledge Graphs and Direct Preference Optimization Statistical Rejection Sampling Improves Preference Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T16:30:13.422573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:30:13.422573Z digest=sha256:41de33a392d797108e291020b9a84646b5b32528d9d8c2f18d46469193c2d7db

Observation 72979ff6-f349-44ed-a7d6-d56854c786b7 · inbound

Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving cites this paper.

Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving Statistical Rejection Sampling Improves Preference Optimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T11:30:39.841578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:30:39.841578Z digest=sha256:b3a6ebf827fa2c42686619f5e4d16eab7de6abbbaee334ae07e8debb88aefe26

Observation 0ea8dbbf-0c3d-4d0c-b1c7-fb18f039d4e7 · inbound

SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance cites this paper.

SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance Statistical Rejection Sampling Improves Preference Optimization

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:36:11.230544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T09:33:57.037949Z digest=sha256:94b94758ff9ba6cfb3fd7949e77e9ddcb90f9e6bf6e5d4b814f91b225c52be1b

Observation 0fdf0e4a-ba7e-46c6-9975-7e81f73d13b9 · inbound

Beyond Importance Sampling: Rejection-Gated Policy Optimization cites this paper.

Beyond Importance Sampling: Rejection-Gated Policy Optimization Statistical Rejection Sampling Improves Preference Optimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:30:18.778593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:27:31.464362Z digest=sha256:55ca6fab68d603e3f3f6125d71590600265f1f32fa06b848700d92303834c054

Observation a2a2cfc4-c3ad-45dd-89c0-3094a38f3b29 · inbound

Reasoning Structure Matters for Safety Alignment of Reasoning Models cites this paper.

Reasoning Structure Matters for Safety Alignment of Reasoning Models Statistical Rejection Sampling Improves Preference Optimization

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:56:04.115722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T02:36:57.093584Z digest=sha256:d6250fbe03cb7b65fcb17c418eb08842cd4c2768e5dac81ef1813acbeaea76c8

Observation 37f56f6a-d22a-466c-b5aa-bcf3033f0566 · inbound

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs cites this paper.

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs Statistical Rejection Sampling Improves Preference Optimization

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.584130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:51:37.506096Z digest=sha256:5b19b9828b7e2a302a8804e8975c0343be40edd5c7211c3d8cca876f7cf0e50c

Observation 50a5ae45-660c-4b1b-b769-58288a06f74b · inbound

Supplement Generation Training for Enhancing Agentic Task Performance cites this paper.

Supplement Generation Training for Enhancing Agentic Task Performance Statistical Rejection Sampling Improves Preference Optimization

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:59:49.531092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T00:58:27.655909Z digest=sha256:4b9b66e054029c7763b63df6620fa7e24738a88134c20bc599bcb36867e1d9da

Observation 8f1a64bc-0b13-4256-9573-312652e6000e · inbound

Efficient Preference Poisoning Attack on Offline RLHF cites this paper.

Efficient Preference Poisoning Attack on Offline RLHF Statistical Rejection Sampling Improves Preference Optimization

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:50:27.173052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T19:29:25.000361Z digest=sha256:b9da027b2ff82d930256185e9e96173425e2390e07ef8114c006a377161ce325

Observation 359e0d48-73cc-4d37-a616-09c67fa5c344 · inbound

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients cites this paper.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Statistical Rejection Sampling Improves Preference Optimization

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:16:08.547916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:be3eb2bb81cd63011455290b22cc8858be8069edafa98680e5b202aaaac9997f

Observation 8070190b-b727-4a74-b125-30ae1a237423 · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences Statistical Rejection Sampling Improves Preference Optimization

Reference 166

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:45:59.782062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:cf3a5a3d8baa4d9e6827297acda8f7db4c6300bb268485132618b1f2d6a39fec

Observation 562b6676-0d06-48a8-beef-5182bb4a7395 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Statistical Rejection Sampling Improves Preference Optimization

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.284521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:97accfc5116812d49a29aca56a797b5ba62da1e97ef284379958367d07a5e5db

Observation 8df7b4cb-a5d7-49bb-8496-79b847b3a2a6 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Statistical Rejection Sampling Improves Preference Optimization

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:45:06.667802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:f21bb357a288ae5d4d6183668a9931252383d97abc1e01b0c4f048320ea1d5c4

Observation 83ac0f78-88ac-4817-9194-7682d793bd13 · inbound

Gradient-Guided Reward Optimization for Inference-time Alignment cites this paper.

Gradient-Guided Reward Optimization for Inference-time Alignment Statistical Rejection Sampling Improves Preference Optimization

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:17:30.661524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T16:43:04.984866Z digest=sha256:161e58f76be04d76b0da066fadb25fbd1ea0626e11e48e148555fc91a6b62677

Observation 83c5d7a1-e3bf-44bd-afd3-774f0fea94e4 · inbound

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards cites this paper.

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Statistical Rejection Sampling Improves Preference Optimization

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:44:40.158482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T10:07:39.554999Z digest=sha256:a0362d7fbe406262b1d9e8aa3d9ceda50c28168fceda21f2a8a97eb76444510e

Observation d44c5aba-3f08-4bb6-b0d3-ca20f5df1148 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Statistical Rejection Sampling Improves Preference Optimization

Reference 261

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:6bfee3b284912618781016a356984e009ba3df56665f3d521695fc3746387d68

Observation 8c2d654c-71e1-4dc5-b084-4cc6390e2569 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Statistical Rejection Sampling Improves Preference Optimization

Reference 262

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:02.899448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:02.899448Z digest=sha256:d32889a09002855cf7cdd86235953f93748f7a2af682ee204a121a3a2d2ff2ff

Observation 4bdf339d-0117-4b99-ad3b-f719cd719c69 · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization Statistical Rejection Sampling Improves Preference Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:20.944852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:20.944852Z digest=sha256:7bd38c2c5fc63a27dae4ffc4bdb131d2c7f54123fa6906adf65c78a6adae38df