Pith. sign in

Paper Citation Record · LEDGER

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable)

As of 7 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2507.07855.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07855 v4

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:52:38.826438Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 965346be-dd5c-4541-a403-540a244512fa · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.129033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.515815Z digest=sha256:464e1bdaaa044d1c36ea70575ac344ec2825df25696653cd472c756b739e709e

Observation 81fccb27-e677-4950-bdec-22a72fff3fde · outbound

This paper cites Alfano, S.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Alfano, S

Reference 2

Resolution
verified exact
raw_fallback, observed 2026-08-06T18:52:39.375745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.577577Z digest=sha256:4611f1654d5f03c090c35eaf9147aad9f3ca9103a1f78e296962a31a73f5fcc6

Observation 80ac216e-2012-4d87-b9ef-03a0217c83c2 · outbound

This paper cites Amari and H.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Amari and H

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:40.116234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.693550Z digest=sha256:1d1fd2c15ac4f70cf5da5d0d405bee2d368ddf78918365df82c01a2148dd75b9

Observation 9a063b77-0699-426f-9c4c-ff02860b399c · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.103357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.787534Z digest=sha256:f702375b816b729797b2786af0498a9a4626e490ea155d0c754f0d0c75e43806

Observation d8a5099f-8521-45ab-8fd5-6f6d36949e01 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.091265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.863589Z digest=sha256:bb76e30b7636d087189b50e60d7693f5d01dab845130ec8e4abb33bd6aaf0069

Observation 5acc9787-cfeb-4a03-a9ef-73a41d7caafc · outbound

This paper cites Bao and N.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Bao and N

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:40.078926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.934402Z digest=sha256:24b3128c04b80acad4e8ec4e1a03adc64647d8d8df1832784f097cf4504dadf3

Observation 75ca5948-f075-4972-aad3-1adc2626e9d7 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.065896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.039406Z digest=sha256:fe6493fb39bdbd5101e149af63081372bd032d90d3ace13fb9a9b85d8f2005c6

Observation 02724bf2-6657-422a-991a-34a97ae0e295 · outbound

This paper cites Blondel, A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Blondel, A

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:40.053423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.197979Z digest=sha256:94204c095ff74ee26ca77c2be25b64ce4530d80f6c78992c89147f0c6158da6b

Observation 9f798169-da66-4976-988e-1f5427eabfa9 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.040610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.263682Z digest=sha256:a34b052c2d98cc843a3e6f41c80af69fcb743675fb8813352348c35039f4dff6

Observation 02b8a3f9-e2c3-4d3b-9b5f-1ec9f1231fd8 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.027737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.366115Z digest=sha256:491d1833dff1b9b82dfd4f1e7095eddfd5b734a51422e3124625025914ecc72a

Observation 98eb9345-e399-4045-bbdb-d7f2cdf0b977 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.014204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.469748Z digest=sha256:41ce21626b8307662a3d921fabb400bfbfd4c6e9c0dca106e15fda995e2329fb

Observation ec68244f-d308-4ea8-8d49-23025093a4f6 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.999814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.542796Z digest=sha256:1435981b334ff34649b75f1a7ba6d6f0b9f5090de56640c5bcf8dc3430a29c71

Observation c8a8fa34-7bb3-4ca3-9565-c24ec8214238 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.985297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.639454Z digest=sha256:dc5e79103692a314af22b9b524c5bd89fc9861375b546ad0d34666e00977d25d

Observation d5c96902-2571-47f1-997d-aba9d37f84e3 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.971992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.793576Z digest=sha256:050230f44c3b74a22230e0599f35223cbb3433226cac1f7ed0542ece372a33a4

Observation 3aba22c1-3715-4e1a-8acc-44cf78ad5ff3 · outbound

This paper cites Doignon and J.-C.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Doignon and J.-C

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.958070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.870271Z digest=sha256:26b1b8965c31724aa5960e35f65520571beb96d95d56da84a58a6f4653139179

Observation e6e2ce3d-05c8-46d8-99fd-a7221203e38f · outbound

This paper cites Ethayarajh, W.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Ethayarajh, W

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.945039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.946745Z digest=sha256:336461c1b94ed438d5b8b9f5bf746601ae39072108f18665a8a99f843f36aa68

Observation b4cd4805-d636-4f9b-999f-2bf556a9bfbe · outbound

This paper cites Gneiting and A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Gneiting and A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.931857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.057746Z digest=sha256:005df7bc08d426fa0d3184a3e3425cf9a2cee31207989ffbaf7eec36dfa2a269

Observation 32cf51ff-128e-4429-903c-9e18fb93f4dc · outbound

This paper cites The Llama 3 Herd of Models.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.156728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:35.156728Z digest=sha256:d49609ff8b7eb51ac6ac688872592e8ffcc4aad9c949409272183b248704424c

Observation 0b4a9192-b9bf-4986-bc94-d7c986731020 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.231062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:35.231062Z digest=sha256:635bfacd157d570c11cc7ee63164b038dd1e2d30202e3b94fe1dc614baa9028f

Observation 8a966e47-8557-4927-a89b-76ffe185a115 · outbound

This paper cites AlphaPO: Reward Shape Matters for LLM Alignment.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) AlphaPO: Reward Shape Matters for LLM Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.347941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:35.347941Z digest=sha256:8f9b2d71a0b6026064bbec60961346218f96665f209b4ef5fc8a62d761093e42

Observation 0b463851-738e-4987-a460-caae629ac756 · outbound

This paper cites Hastie, R.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Hastie, R

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.909162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.450714Z digest=sha256:051f2d23aacc7f6d02c24a9fbf603c8e201f5db7f5ddd94497ef062c7da9853e

Observation 5229e8a3-3a7e-4622-99b7-61e3fb4701fb · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.896574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.600101Z digest=sha256:a3d1b4b1f9b7b9cfecef495dcbb4646941e903f91fd4c51b17de9cec18a4ab90

Observation 8a5b238e-5927-42cc-bb41-8ba86e955dda · outbound

This paper cites Huang, W.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Huang, W

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.883467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.670121Z digest=sha256:e77fb02420bf600f614e5d1a19c4a8a0a20cd72021e52640c18226f4749e4684

Observation f4ce977c-d379-4f66-ae69-9b6c4f9dd52b · outbound

This paper cites Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.793726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:35.793726Z digest=sha256:6c6b85f310fd13f6fbc63f978d67f1246d7bd486ef1d336a8919c1a12f5a1b94

Observation 54ec5f14-ff72-42b4-be43-e7ced6156b14 · outbound

This paper cites Kakade, A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Kakade, A

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.870296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.884856Z digest=sha256:5f28128ee9417cc8d3f3ff515ff02e4977d84c0933f4d50922ac36d207468aea

Observation 207c34d9-a568-4712-a475-d5c6144f41cf · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.857109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.950676Z digest=sha256:f7d4facc76770cd77b970a4a211608c6475f8589f00a2ec90f399b838381f28c

Observation 1a545f81-5c72-4731-a87e-71a3aea4a00b · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.843801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.013640Z digest=sha256:9969fcc96958e077954c3cf76bbf6a94b5e5d96965baf2cd08645566403f3bc9

Observation d2c9642c-e399-49e0-952e-8cf8db96b42a · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.830449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.081485Z digest=sha256:e6ff3bc039ade33fcb0e9758271a8c0c255b5c24f64ea2e9c7c4927a12146e59

Observation d31023e5-7897-444b-a6ca-3cbc8878500c · outbound

This paper cites Direct Preference Knowledge Distillation for Large Language Models.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Direct Preference Knowledge Distillation for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:36.169261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:36.169261Z digest=sha256:3231944512192d2300cd9af55da53bfeb040decbcf49f216fd864fe03c0e3b77

Observation 90904036-3fb3-48b8-8851-53fe7c53e6a8 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.817390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.262670Z digest=sha256:3c92ffc04c6f89f6744639a24f03ee190672f7b50dee528e25848d6d908172a3

Observation 8a726a5b-64ff-4d3e-b2b1-3d1b9561d547 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.804388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.334781Z digest=sha256:15b50f4cfccbcda201b0e18de7a470a1a3ad175b642672f3edceb8bcd4409e21

Observation 84697ba7-17f3-401a-adc4-40948e8236d8 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.791682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.408635Z digest=sha256:dd6393f15d0f1f359ea3b1e2f8dba72ca69536d7ef10a0b48b173a95ef512072

Observation 8ba6d603-4793-466b-9b3d-f8b415ccc5aa · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.778771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.471253Z digest=sha256:8f11fd9f7f688011b698d5eb3d5499d814dc15dc4df78e7912f3ec837292b4be

Observation 076ddc84-144c-4949-81a0-2359698b150f · outbound

This paper cites McCarthy.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) McCarthy

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.765690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.561030Z digest=sha256:153f5e48d453e6c9d8acfa46d687c7fd01d7dac87d09e7bf8053590b124bb2c8

Observation 30c1683d-dab5-4f38-9829-ae7326836ebf · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.752867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.626119Z digest=sha256:6bbedbbd0122b9acbe94006e812f7f286196680cc07c86ccc064b35e1266e3ef

Observation 2d8a6005-dce8-457c-94ce-b0e00586c136 · outbound

This paper cites Mitchell.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Mitchell

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.740328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.721352Z digest=sha256:f8a07b23704cea700c8040996275b38e5840bc821787234fd9737e6a6539081f

Observation 640d0667-d9b3-4cf8-a373-16a7155fd115 · outbound

This paper cites Nock and A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Nock and A

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.726540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.788286Z digest=sha256:6bf6fa3999ec64ae7699140eb0cbc2c32c7350cc833e193873ddf621e0d4eb43

Observation 47acef93-28e0-467b-8706-75b236f882bc · outbound

This paper cites Nock and F.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Nock and F

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.713336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.854967Z digest=sha256:ce821481dc13fa8ca9c47e8dc93d39853947007188b1170f6e3ded419dab992e

Observation 19057d04-3f49-4fe8-a5dd-af41e8ea68d9 · outbound

This paper cites Nock and F.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Nock and F

Reference 39

Resolution
verified exact
doi, observed 2026-08-06T18:52:38.889685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.915265Z digest=sha256:10909edb11fe4e2ebada01a49a56e1a226ff55a97ea05a08c11ea348db84773a

Observation c59ab0c2-4ffc-48e9-ba54-f641f9820612 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.699922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.001364Z digest=sha256:c804d0aaf660b163380d2965e898cb9ca74d107368efc430d2ea00833b881d1e

Observation d48ec375-fdcf-4214-8efa-b92bb651f903 · outbound

This paper cites Nemotron-4 340B Technical Report.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Nemotron-4 340B Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.079112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:37.079112Z digest=sha256:28136127ba40c2d4a945187a2d94a1fc3c5a7aa507169076cefc87be5ca1e0d3

Observation 6d9a245f-f9fd-455f-8b32-49a6a3e425fb · outbound

This paper cites Ouyang, J.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Ouyang, J

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.687241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.152941Z digest=sha256:1c3ecb83fae51c04c4dcc9703e51926d50de0bc60291fc7f749554b7630297f1

Observation b96596a0-c6c6-45de-8988-40fcf3df7371 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.224508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:37.224508Z digest=sha256:b60b34a73c2a97ec0037c73059ecf529620882426f8b244da0f98579f7758c14

Observation f56d7776-02e4-497b-9014-63d4212ff29a · outbound

This paper cites Rafailov, A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Rafailov, A

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.673035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.280470Z digest=sha256:147a93ce25ae1fd510606ec0251943b239c182ff433512982d6dafa7c680bd9b

Observation c9a5e42b-2936-4a46-93f6-1f6c17925fc3 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.659448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.389498Z digest=sha256:ffe1d164139279bbbaa3ab1ac77665ec883fa1c013428418b2585ab2a029356d

Observation 4d5ac30a-a30f-43e0-915a-d8c703a5fb74 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.471398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:37.471398Z digest=sha256:e1ac5beab5034747a1cb39b32ffe687db12ecb8b9d251ed8099ce4074a244921

Observation cbb500e1-6af3-4c61-9eea-1934bd815d86 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.535852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:37.535852Z digest=sha256:0053b2f3ce02af52b6eefcc3ffe5f80917c6f91b530db23598533e5667adc749

Observation 0d563e8f-888e-405b-992a-05b6de337c58 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.646073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.599101Z digest=sha256:c8b3ccc89738244f6cf5e8460f7c7fc45f3ce3f7aca2bec99f608b617ee109a0

Observation ac251eed-27f3-4bb2-8ba0-6b38c10425e9 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.633197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.665105Z digest=sha256:75e1564a1e2e1b10c7d836bbcf541ac1d2be25dcb66fa30816961115cf67d94f

Observation 3ce4ba05-615a-4030-be25-2691ab56a28b · outbound

This paper cites Slocum, A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Slocum, A

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.619757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.723880Z digest=sha256:be05a950bb05f8b39265d2f23f36230d6c5f8d9c333e1b350196df428eedb30d

Observation 8812f057-e68f-4d3e-b688-476d9e86ca73 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.607057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.830259Z digest=sha256:9622fe09adc2a423c333a8777ddd1825778948efc3ea086522026bbc6cc144ad

Observation bd040ad2-ddcd-4ea5-8881-49e96b1d3e82 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.893399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:37.893399Z digest=sha256:d1f31f002090a652499cf1ccf5a32b28e700637797ce454ff9bbac3e7e760733

Observation dac90df9-7ad8-44c4-9c78-95413555e016 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.594541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.954555Z digest=sha256:7f1382fe8d5404c7a912376b5fcc40a663236cc768134be1369de1eb7c5415d3

Observation 7f341fc4-d89b-4a78-bcab-887fe94274b5 · outbound

This paper cites Sypherd, R.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Sypherd, R

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.580769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.019369Z digest=sha256:dd1b4508f8f58e12dd480615408c426e481212054d9da4ae87158b40e3918968

Observation 9daebe3b-71e2-4a3b-bbe4-99ffe76e8547 · outbound

This paper cites Tunstall, E.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Tunstall, E

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.566845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.147360Z digest=sha256:fe741dae8fbdd11bb809d326d83c59f222a7c67e011b294ed2166c098331630c

Observation 88ca329d-0891-469e-9bfd-9ba8e26cdf17 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.553032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.268692Z digest=sha256:6a054f6e29236c81d5e9556ff9ff328b384a92e0370d85e1bdd08a1fd2d2f1e9

Observation 50fd11cb-35ba-45bb-bacb-9fb4b9a50010 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.538286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.437209Z digest=sha256:b492e6be610441c27c818e1d21ecb238293b949a7d067b9a724e5cfae9d7a77e

Observation bfc3f260-911f-4edc-ac4c-2c45d148e19b · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.524264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.563817Z digest=sha256:913a589b77838b76941d212a08684ef6f98c919f1e8a3642af01be5c5625c9de

Observation adc61e49-10df-49f2-96e6-7f0df74fd2f1 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.510989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.682104Z digest=sha256:9893a6a4de1f1fb4bfbb87c24c71fa377a7460c83f6e05147cbbdc6a96431ba1

Observation e22ddd89-11b9-4ec6-8db6-9870ea80c6b6 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.496283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.778837Z digest=sha256:cd8d0242a88ff024a93ce3e9bb2bcb4da2d51a600ba0b13568b53057467ca4a4

Observation 08a406d6-ea26-4385-ad56-1768411060d0 · outbound

This paper cites Qwen2.5-Omni Technical Report.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Qwen2.5-Omni Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.783043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.783043Z digest=sha256:9ac0e7d5cd1eedfece7350265df06414cac49e629dbcededefeb858ca12f54d6

Observation 87803603-1cad-408c-b14e-b59054fc6f47 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.480895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.787672Z digest=sha256:bd2017a177be20e432abcea66d0af21fa4e6ff704bee0851fddf6eae14de4740

Observation 0881ed7b-d103-4605-bd91-21ebd84de778 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.465569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.791473Z digest=sha256:8ffdf03a963bfaa8648184b010db26816a7c1ae63b72ec464765a35e065ded4c

Observation ec96c192-f2de-446e-a537-58949b0618cb · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.451379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.795690Z digest=sha256:5462a262f222dc7544a710de18f4dbca9660dbfc7480ca4cbc7d26493855985c

Observation 87ebc61f-baca-466e-855a-47b090309761 · outbound

This paper cites Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.799899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.799899Z digest=sha256:ecd0c59ee6984ed4ce795d046a5a5789277395e54f2a8842ec43382883cc1b6b

Observation 57dd9b03-2a44-46d3-b077-a2ffcc84e3d3 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.436613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.804815Z digest=sha256:c41852495d206c1e9ecaedad171286cb79426ec3ff404e7f8e6e0b5bad0e88b1

Observation 6b833df4-92a0-4977-88b7-29c850e08e7f · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.808890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.808890Z digest=sha256:87cae93730f4c063596fd6e509635f2e17e4cbf5fd20c44130ac7c6c98a9bfd5

Observation 0e55d8d0-b384-45ba-917f-38e4578e5718 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.421914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.813440Z digest=sha256:93885767d992bca386fbdc1838eb3defc2023ef868cd3622febdddd88f44fbcd

Observation c8c75fc2-f3d9-4982-935b-29a743119ac3 · outbound

This paper cites @esa (Ref.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) @esa (Ref

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.817370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.817370Z digest=sha256:b871cc150735a77b7f911f9a467a98a84f604c81e02128e4f0b7e3f3dee4e822

Observation acb13461-2315-4188-8636-15f03c2cd226 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.822165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.822165Z digest=sha256:cf745325e184f86e01d75468d7b464502802b2921adf6a5da14f9e498410ac4e

Observation df42e405-82d6-4b7c-92ba-e6aeddbf3765 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.826438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.826438Z digest=sha256:01314ff44310c80d819ea34f15eb9fe5b2958db1cf1bebb1c91c1e86efaa942d

Pith citing papers

No inbound Pith citation observations are available.