Pith. sign in

Paper Citation Record · LEDGER

BPO: Revisiting Preference Modeling in Direct Preference Optimization

As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2506.03557.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03557 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:07:35.625982Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved14
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e5950ee1-ad35-435a-b45e-7accb3e99048 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

BPO: Revisiting Preference Modeling in Direct Preference Optimization A general theoretical paradigm to understand learning from human preferences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.485605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.485605Z digest=sha256:329706f163172f3c94c7d798fea699c613822c7b44f59835ab4e2ea40578370e

Observation f3974c4a-bf03-4ccf-9eb7-ad3e028f6862 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.490833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.490833Z digest=sha256:2d137b9bea3270f22af3c14f300911b931ff9e0ae95ab20e48422299b1bd531f

Observation e15453b3-d4dd-4897-8f13-766193e0a113 · outbound

This paper cites Bradley and Milton E.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Bradley and Milton E

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:36.133739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:07:35.496675Z digest=sha256:0b083f4b42c4cc57839df94f690aec61decf0488d90e2797fffbe449cc8f1405

Observation 8e2dfe17-bf77-4f5d-84bf-98628c0f0223 · outbound

This paper cites Christiano, Jan Leike, Tom B.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Christiano, Jan Leike, Tom B

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:36.117449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:07:35.501795Z digest=sha256:9896a30b549bfef4cbb0795f4df39267f36846f7efd494fdfade7ed1deba66e2

Observation 206d22a5-23d9-499f-845a-c3e075efd8a1 · outbound

This paper cites The Llama 3 Herd of Models.

BPO: Revisiting Preference Modeling in Direct Preference Optimization The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.511097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.511097Z digest=sha256:237610eab3fe7454cf43786ee33270bec42cb916c68903c51ab39f47a3f057c1

Observation 7d895142-cc54-4c2f-9968-50909c9aae6e · outbound

This paper cites Model alignment as prospect theoretic optimization.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Model alignment as prospect theoretic optimization

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:36.101078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:07:35.516121Z digest=sha256:6f12c7c43db1a923e1740df96b08162bb2f74f0aa55060454e17ed34b863b183

Observation e5742135-7f18-4600-b5fa-a099ec757e78 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Olympiadbench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:36.084841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:07:35.521120Z digest=sha256:8a65d28fbbc4f0634f49150eda509105fd59407da6ac11fa92a6a19d6d8dfdaf

Observation 9d2fa7ae-b1fa-4363-862e-47de82b4b06a · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Measuring mathematical problem solving with the MATH dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.526189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.526189Z digest=sha256:7efaab21b56812e4d3956744d31f5c3707824de4e8c09c843547e355821ee0a8

Observation 0c5d7277-ac09-46ff-ba0b-e9ce2e6f52af · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.531217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.531217Z digest=sha256:c92ce34d40ed222b2571c8678917fabe017ce3e091923783d9a0f95f055b42c5

Observation 9b54130d-a35a-43f2-8238-0200e750e39f · outbound

This paper cites GPT-4o System Card.

BPO: Revisiting Preference Modeling in Direct Preference Optimization GPT-4o System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.536884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.536884Z digest=sha256:d3bba2d6462aa9b6989d2d1bac17c574d0feae992154c0dfb7721e8b5263b94e

Observation 42713f56-de29-4ef6-88d9-65a1bbef4e6b · outbound

This paper cites Smith, Yejin Choi, and Hanna Hajishirzi.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Smith, Yejin Choi, and Hanna Hajishirzi

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:36.057918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:07:35.541847Z digest=sha256:7d03df6b45b9afa5b1a698afffce50b46b2cd763de81376021c63a96fd8faf90

Observation f48ca749-2555-4439-8d9b-b906c9031993 · outbound

This paper cites Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:36.040876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:07:35.546416Z digest=sha256:a776b368dad6c34346b8f5862993250a6e1d01d66999b37d0cd5805b180c433d

Observation 473d34a4-8559-4511-b02f-13f0fb36e1f7 · outbound

This paper cites Mitigating the alignment tax of RLHF.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Mitigating the alignment tax of RLHF

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:36.017881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:07:35.550615Z digest=sha256:4b3ec66076345679de9a7081f0afbaf6526347593607c8a218d7e3fe9275836e

Observation d9fca5fb-4421-44c9-b46f-ce16871d4c3e · outbound

This paper cites Liu, and Jialu Liu.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Liu, and Jialu Liu

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:36.001899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:07:35.554964Z digest=sha256:91ebbdfc6826a8fd97ccb3d87ec26c8dd491d03506209036cba37d61867a8d16

Observation cd2b8c06-7a06-490c-919e-87236fc078f3 · outbound

This paper cites Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.563139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.563139Z digest=sha256:3ec4ddb8d46025d1bc67474acc9f4fc8b56ce53378ba530ca1879b54b5a63225

Observation 076167b2-c41a-4115-a719-d5d8db6f08a1 · outbound

This paper cites Mankowitz, Doina Precup, and Bilal Piot.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Mankowitz, Doina Precup, and Bilal Piot

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:35.970965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:07:35.567906Z digest=sha256:4bfb1259f32f89b8a126a8f5aab365e92630951144cd246076e1d3f4193254dd

Observation 3600ca7a-a6cc-4e69-9115-59b0436bb9d7 · outbound

This paper cites Training language models to follow instructions with human feedback.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Training language models to follow instructions with human feedback

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:35.956335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:07:35.572108Z digest=sha256:dd761f6413d2c4a48bb88868c1320d119eec09ac9c153a84cc799b81c2369ab4

Observation e2924f18-9bb9-4ea1-a250-7dba680a2c2a · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.576714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.576714Z digest=sha256:0a12f1e20829ce1b7cb6435cbb33660706ae192ead1dcb53c246fbf93658d750

Observation f05344bb-5d1b-4041-ad1d-4162e15c0862 · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Manning, Stefano Ermon, and Chelsea Finn

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.581612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.581612Z digest=sha256:16b46e79e3d2f380f33cc35088a494ca115aa306605437b5bb25d9ad611c7884

Observation 631117fe-0c0c-415f-9f38-53bf08418ae9 · outbound

This paper cites Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.586376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.586376Z digest=sha256:88a641b5c2127be430b07c28956bcdb8ddfaff5872dfb268ca40000fa72bf899

Observation 2e60a435-eee6-4934-9705-3076e36ab609 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul F.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul F

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.591108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.591108Z digest=sha256:5fe9e6f537413a0624674a255bcbc3d942a02c31fc299c130394cb0f07ff4ee6

Observation f0838878-0a6b-45c3-a985-f6dd2b9d34f1 · outbound

This paper cites Generalized preference optimization: A unified approach to offline alignment.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Generalized preference optimization: A unified approach to offline alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.595929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.595929Z digest=sha256:d07d868850bd1cfa05fb5ccda43cc34169ff4db67901ce8d81064ee5140754eb

Observation cb73d625-993f-4164-9794-65d8dca20cb4 · outbound

This paper cites an unresolved cited work.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.600673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.600673Z digest=sha256:fb9d8e664be6318a2d0d22f2c8b8fb1b85eac4a6487f3f3a198fc11ae65e671f

Observation df8e9244-c3c7-46ba-a09c-fdefdbd65d6f · outbound

This paper cites Contrastive preference optimization: Pushing the bound- aries of LLM performance in machine translation.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Contrastive preference optimization: Pushing the bound- aries of LLM performance in machine translation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:35.889562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:07:35.605409Z digest=sha256:c846927ff9f04ae31113957321c057b85766824d5a2ad3e2097cc7b6c57fa74a

Observation 30d500de-f698-43bd-a336-8e2b867414d7 · outbound

This paper cites Is DPO superior to PPO for LLM alignment? A comprehensive study.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Is DPO superior to PPO for LLM alignment? A comprehensive study

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:35.873245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:07:35.609927Z digest=sha256:1cc1becb2a131cbcfa9b23cdc7276f37d9de2c1d42619b7c8028a43074d0e389

Observation ee9aeee2-2e93-44d4-981d-6d58fae6594e · outbound

This paper cites Qwen2.5 Technical Report.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Qwen2.5 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:35.615063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.615063Z digest=sha256:7dab8d4750dd681ca956f55ce26dcafe0a21ef8b35eb574ed55d5d375691c818

Observation e928c41e-609e-4942-909e-de9dc75b6c28 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

BPO: Revisiting Preference Modeling in Direct Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 28

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:07:35.619682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.619682Z digest=sha256:7ee5ba3ce0e567462e7d9a2dd77e8e28f6b2e4e828e3eacd74d43cb32392831b

Observation 7aa043a6-3a0b-4006-8b71-1094c7bd69dd · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:07:35.855307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:07:35.625982Z digest=sha256:7072e4e8167b0eaef3ffd12dfc40ca4dcc888889084d909dca4efd9d65ec6002

Observation d9bb6598-2471-4399-a008-d5c62c8262ff · outbound

This paper cites an unresolved cited work.

BPO: Revisiting Preference Modeling in Direct Preference Optimization Unresolved cited work

Reference 2024

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T11:07:35.986800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:07:35.558965Z digest=sha256:758c4b4cd6cdd4d1cddd56e3d18d5e5242d3dbee83bcf08a74dee8eb2ab8d57b

Pith citing papers

No inbound Pith citation observations are available.