Pith. sign in

Paper Citation Record · LEDGER

RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2402.10038.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.10038 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:47:39.035241Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T04:12:02.944973Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aab445be-73d4-48e6-bcc1-45aa677de7e8 · inbound

Enhancing RAG with Active Learning on Conversation Records: Reject Incapables and Answer Capables cites this paper.

Enhancing RAG with Active Learning on Conversation Records: Reject Incapables and Answer Capables RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T22:47:39.035241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:47:39.035241Z digest=sha256:28aa49c0ab1aa1df975a39a802383d71327e5b218978525baccef0dc962b506e

Observation 846f9a8e-399f-4bad-bf6b-ea1336025286 · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.235610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.235610Z digest=sha256:948c7073d2a7caeee7fe7d199164b99bcbde6d8abbd5380e535e9fae795b41e0

Observation adf0dcf7-981a-4893-86d8-c3b53b99b4a7 · inbound

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy cites this paper.

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:30.734672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:30.734672Z digest=sha256:724d1e151e1e6e3f6023aa23cc91220b921393d9f2a2ba70746e0c23fd076787

Observation c8114cdb-25c0-47e5-a745-3d238a0a2864 · inbound

Scaling-up Perceptual Video Quality Assessment cites this paper.

Scaling-up Perceptual Video Quality Assessment RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:07.261600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:07.261600Z digest=sha256:8742750aace1ee6d7077306fbaa44cd89bfbabf1e7a7ac7fa99020bf790ac6f2

Observation 40a5a8fe-fa25-41ed-98bc-6e4737b931c2 · inbound

Knowledge-Aware Diverse Reranking for Cross-Source Question Answering cites this paper.

Knowledge-Aware Diverse Reranking for Cross-Source Question Answering RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:07.603300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:52:07.603300Z digest=sha256:5e97454faa5899bce0305999cefd773d96769f7317a181eef26b5450224553e8

Observation 0cc3a2af-fe41-4746-a08a-ebc37cc62b27 · inbound

MetaLint: Easy-to-Hard Generalization for Code Linting cites this paper.

MetaLint: Easy-to-Hard Generalization for Code Linting RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:12:02.947633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T04:07:31.283348Z digest=sha256:8497316e9f8784e23ce3a0e4b9da6e17a92961c4ffbb0e93227babf3b8e3b1c6

Observation 8b2b6b88-1bdb-4c4a-abe2-553e0ab4164e · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.802584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.802584Z digest=sha256:be6ec9cc30fc61c9ad15281de8536ba4e9993364398ac9139998c219e4267319

Observation 06676987-a8c8-4a8d-be70-93735bf3525f · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:06.443976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:85299c66c8c021da22991f252b2800e7075b11a09d20e5734cc1eb7e655b15de

Observation 3ba40e30-fde9-4350-b537-9f7b970a3f53 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:58.868410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:0a2a55c6002c7ace9fe12645530d13da2ac11b9115a9fdce58b4f5b1862a54ac

Observation 73fd6499-e6b4-4df9-99e5-61aa69576854 · inbound

Exploring Post-Training Alignment of Small Language Models for Biomedical Data-to-Text Generation: A Case Study of Medication Leaflet cites this paper.

Exploring Post-Training Alignment of Small Language Models for Biomedical Data-to-Text Generation: A Case Study of Medication Leaflet RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T05:16:31.076390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:16:31.076390Z digest=sha256:844ed88695f201e30d8357f235309e9e4082e32b006de0a55f06a1c1c278d4a9