Pith. sign in

Paper Citation Record · LEDGER

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable)

As of 7 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2507.07855.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07855 v4

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:52:38.826438Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 965346be-dd5c-4541-a403-540a244512fa · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.129033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.515815Z digest=sha256:7254931e49df8a277f9864ed8492009068e0c72bc9d49e6a3f426725c3d4cbf8

Observation 81fccb27-e677-4950-bdec-22a72fff3fde · outbound

This paper cites Alfano, S.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Alfano, S

Reference 2

Resolution
verified exact
raw_fallback, observed 2026-08-06T18:52:39.375745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.577577Z digest=sha256:00b4773a01164a33035fa5bbc56fbcc60761a7695106b562ff842e0f4c2278f7

Observation 80ac216e-2012-4d87-b9ef-03a0217c83c2 · outbound

This paper cites Amari and H.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Amari and H

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:40.116234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.693550Z digest=sha256:921a42c20a75beb6df5374edc763a11d4f2bfb9abdc671819ffd9866ed4cc728

Observation 9a063b77-0699-426f-9c4c-ff02860b399c · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.103357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.787534Z digest=sha256:16a048bfb46f8dde33559b492de89dc523f1a940a61837d7494d855b62275534

Observation d8a5099f-8521-45ab-8fd5-6f6d36949e01 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.091265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.863589Z digest=sha256:89caa15caec3a41fd67f6660498af921b958c3894efdd725ea3473316e3d6002

Observation 5acc9787-cfeb-4a03-a9ef-73a41d7caafc · outbound

This paper cites Bao and N.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Bao and N

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:40.078926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:33.934402Z digest=sha256:d82f4539f81f72656f95d849520ad11a0568127552f57e070870db7fa7555d43

Observation 75ca5948-f075-4972-aad3-1adc2626e9d7 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.065896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.039406Z digest=sha256:94957507e7be5130ce7f34674b43da3150f98e3f136eb0ace0c2feef0444256c

Observation 02724bf2-6657-422a-991a-34a97ae0e295 · outbound

This paper cites Blondel, A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Blondel, A

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:40.053423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.197979Z digest=sha256:33ff19b168ae58d118c7362689aafef0796ab5bd12aefdde25aade83b9c2121c

Observation 9f798169-da66-4976-988e-1f5427eabfa9 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.040610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.263682Z digest=sha256:9c1a12a15754e153ab11e76b663956dbcbe409b89ea4d47431e9f191faa7622f

Observation 02b8a3f9-e2c3-4d3b-9b5f-1ec9f1231fd8 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.027737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.366115Z digest=sha256:c57938237d65948301a052ece9f1406e4107525b7453c805202f4772916673c3

Observation 98eb9345-e399-4045-bbdb-d7f2cdf0b977 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:40.014204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.469748Z digest=sha256:b1280dec7c898b0175127551e6667cad36b7e13c7cde8311b0e98be222bb644a

Observation ec68244f-d308-4ea8-8d49-23025093a4f6 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.999814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.542796Z digest=sha256:04001a5f78754aadf81b5916501aa3e8bea26e9f8d6ca2f885f8058135fcb434

Observation c8a8fa34-7bb3-4ca3-9565-c24ec8214238 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.985297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.639454Z digest=sha256:e0e1b9641195bb2fa9d6e96c3970ec74a9891eff013a1ee1e1833762a9ed3843

Observation d5c96902-2571-47f1-997d-aba9d37f84e3 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.971992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.793576Z digest=sha256:fecbb0d8bfb44342ef4e638f739c85d4321810d406e7b17c89705db1fc28ab24

Observation 3aba22c1-3715-4e1a-8acc-44cf78ad5ff3 · outbound

This paper cites Doignon and J.-C.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Doignon and J.-C

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.958070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.870271Z digest=sha256:3549e31bc50ff4f2d61e86c2af350d2b0656d5f68a4438421114c139aac307b8

Observation e6e2ce3d-05c8-46d8-99fd-a7221203e38f · outbound

This paper cites Ethayarajh, W.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Ethayarajh, W

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.945039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:34.946745Z digest=sha256:5ef5390e6b1f5ac276b6cbf78d7da77928328736876f6f5252ca445d98e7c2c4

Observation b4cd4805-d636-4f9b-999f-2bf556a9bfbe · outbound

This paper cites Gneiting and A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Gneiting and A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.931857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.057746Z digest=sha256:b29cc67eab141ec60a93f854c8a36b954c474051521560fa39343e4acc1519c4

Observation 32cf51ff-128e-4429-903c-9e18fb93f4dc · outbound

This paper cites The Llama 3 Herd of Models.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.156728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:35.156728Z digest=sha256:d49609ff8b7eb51ac6ac688872592e8ffcc4aad9c949409272183b248704424c

Observation 0b4a9192-b9bf-4986-bc94-d7c986731020 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.231062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:35.231062Z digest=sha256:635bfacd157d570c11cc7ee63164b038dd1e2d30202e3b94fe1dc614baa9028f

Observation 8a966e47-8557-4927-a89b-76ffe185a115 · outbound

This paper cites AlphaPO: Reward Shape Matters for LLM Alignment.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) AlphaPO: Reward Shape Matters for LLM Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.347941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:35.347941Z digest=sha256:8f9b2d71a0b6026064bbec60961346218f96665f209b4ef5fc8a62d761093e42

Observation 0b463851-738e-4987-a460-caae629ac756 · outbound

This paper cites Hastie, R.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Hastie, R

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.909162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.450714Z digest=sha256:0cdba965ac69d59e38907cc9ef76ac40062b9cdf9f55304382e76bf672c649de

Observation 5229e8a3-3a7e-4622-99b7-61e3fb4701fb · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.896574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.600101Z digest=sha256:6c0fb8a6fa94c47535b9eee43a830480ce8bb4c50a7f2ea18f576a70a6d4f542

Observation 8a5b238e-5927-42cc-bb41-8ba86e955dda · outbound

This paper cites Huang, W.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Huang, W

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.883467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.670121Z digest=sha256:d21254a4331868034485fa096c8410669d73a653fe7a714ea8a3bf9665cec88a

Observation f4ce977c-d379-4f66-ae69-9b6c4f9dd52b · outbound

This paper cites Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.793726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:35.793726Z digest=sha256:6c6b85f310fd13f6fbc63f978d67f1246d7bd486ef1d336a8919c1a12f5a1b94

Observation 54ec5f14-ff72-42b4-be43-e7ced6156b14 · outbound

This paper cites Kakade, A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Kakade, A

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.870296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.884856Z digest=sha256:aaf77f147996f4d5dfc85a02c6d3f5499f5d25716af0200556822dff89f69638

Observation 207c34d9-a568-4712-a475-d5c6144f41cf · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.857109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:35.950676Z digest=sha256:1b725b2058ffcef00e025da1a7e907e67beb5b00cb6ff4dc9d3941bc7a2a3b95

Observation 1a545f81-5c72-4731-a87e-71a3aea4a00b · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.843801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.013640Z digest=sha256:7bc217146dbccc88dc8fbe7694bbccb3ba9a809b8d04c714f71fb4379bf78f80

Observation d2c9642c-e399-49e0-952e-8cf8db96b42a · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.830449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.081485Z digest=sha256:609aa44a189b2c3aa82747c4ee3d4eec21db99c44c4574f77c98479b918fa867

Observation d31023e5-7897-444b-a6ca-3cbc8878500c · outbound

This paper cites Direct Preference Knowledge Distillation for Large Language Models.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Direct Preference Knowledge Distillation for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:36.169261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:36.169261Z digest=sha256:3231944512192d2300cd9af55da53bfeb040decbcf49f216fd864fe03c0e3b77

Observation 90904036-3fb3-48b8-8851-53fe7c53e6a8 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.817390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.262670Z digest=sha256:f7a6fc58e6303ac28cd7fa8448faa2bf2269e87289cf992f54a16c77c7277051

Observation 8a726a5b-64ff-4d3e-b2b1-3d1b9561d547 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.804388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.334781Z digest=sha256:e813949af9b83552e4a58c250f9f81d889f11b5ae6f1ef44e826da3e5037840a

Observation 84697ba7-17f3-401a-adc4-40948e8236d8 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.791682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.408635Z digest=sha256:ff2c5f18a5808d020d087a43bccb6fe552f0e720fc4d725ba8c2dfc4bd15fef6

Observation 8ba6d603-4793-466b-9b3d-f8b415ccc5aa · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.778771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.471253Z digest=sha256:8c73ab064950940829dc66443409668c7a194667ae0033b53e005cb61746ec4a

Observation 076ddc84-144c-4949-81a0-2359698b150f · outbound

This paper cites McCarthy.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) McCarthy

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.765690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.561030Z digest=sha256:fdbec039cdbd0b25111d53348408957ac1e7db144226fc3aa4f6a7d375b0479b

Observation 30c1683d-dab5-4f38-9829-ae7326836ebf · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.752867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.626119Z digest=sha256:cdd9154ffdc3c8a62c85dcc5745117cfa89aa9020374277540d2dda5ede53ffb

Observation 2d8a6005-dce8-457c-94ce-b0e00586c136 · outbound

This paper cites Mitchell.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Mitchell

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.740328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.721352Z digest=sha256:52ebedd67bf53cd409561729d8c603189baa424be8f206d842026eac8f1ec0e8

Observation 640d0667-d9b3-4cf8-a373-16a7155fd115 · outbound

This paper cites Nock and A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Nock and A

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.726540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.788286Z digest=sha256:c4fcc38b25c39c6ab5d200067669e38cf62a76bd10b483357f3b85d72b20a756

Observation 47acef93-28e0-467b-8706-75b236f882bc · outbound

This paper cites Nock and F.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Nock and F

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.713336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.854967Z digest=sha256:54ed1c96adea0b5de383f6d1e9a171fb47084a99df73371402e809c8eddf8650

Observation 19057d04-3f49-4fe8-a5dd-af41e8ea68d9 · outbound

This paper cites Nock and F.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Nock and F

Reference 39

Resolution
verified exact
doi, observed 2026-08-06T18:52:38.889685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:36.915265Z digest=sha256:690a0dc4d7b7631a98eac09d7827360a6fecd8c3cd5bd724ff1fbc9d8ab3c97b

Observation c59ab0c2-4ffc-48e9-ba54-f641f9820612 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.699922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.001364Z digest=sha256:13b52ace33f04e2c086b9370fb25bc71ce780e69932015ae204a803bcaeffc23

Observation d48ec375-fdcf-4214-8efa-b92bb651f903 · outbound

This paper cites Nemotron-4 340B Technical Report.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Nemotron-4 340B Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.079112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:37.079112Z digest=sha256:28136127ba40c2d4a945187a2d94a1fc3c5a7aa507169076cefc87be5ca1e0d3

Observation 6d9a245f-f9fd-455f-8b32-49a6a3e425fb · outbound

This paper cites Ouyang, J.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Ouyang, J

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.687241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.152941Z digest=sha256:25c22c471e00c65097574f41eb165c357831110ae4de1e8d176977d0ba8afa57

Observation b96596a0-c6c6-45de-8988-40fcf3df7371 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.224508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:37.224508Z digest=sha256:b60b34a73c2a97ec0037c73059ecf529620882426f8b244da0f98579f7758c14

Observation f56d7776-02e4-497b-9014-63d4212ff29a · outbound

This paper cites Rafailov, A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Rafailov, A

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.673035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.280470Z digest=sha256:bb96a244b15ab5c4d912191d32d39cc2a423abcbcbcd2e249dfa3b654bf37f9c

Observation c9a5e42b-2936-4a46-93f6-1f6c17925fc3 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.659448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.389498Z digest=sha256:a9112662c8e93fbb4a8a32ca2168a26a1d6fa724493fc2e7c35ae5b9b994d32c

Observation 4d5ac30a-a30f-43e0-915a-d8c703a5fb74 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.471398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:37.471398Z digest=sha256:e1ac5beab5034747a1cb39b32ffe687db12ecb8b9d251ed8099ce4074a244921

Observation cbb500e1-6af3-4c61-9eea-1934bd815d86 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.535852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:37.535852Z digest=sha256:0053b2f3ce02af52b6eefcc3ffe5f80917c6f91b530db23598533e5667adc749

Observation 0d563e8f-888e-405b-992a-05b6de337c58 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.646073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.599101Z digest=sha256:1e1e597ec4689766577c30a9fb02ab4f7c2c166441afa08cd1d3f6dd2cd664e1

Observation ac251eed-27f3-4bb2-8ba0-6b38c10425e9 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.633197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.665105Z digest=sha256:9d7532bba1f3b4fa5950c6ff4264e0bd8e27e9584c19bdfd86338f65402bbf87

Observation 3ce4ba05-615a-4030-be25-2691ab56a28b · outbound

This paper cites Slocum, A.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Slocum, A

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.619757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.723880Z digest=sha256:8444a6fd095cec92174e765530cb156b182eb67f5e66e35d6a1a0881de0c1056

Observation 8812f057-e68f-4d3e-b688-476d9e86ca73 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.607057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.830259Z digest=sha256:be5ed53c65a5c322ffb29823e63250355cf094bec0210675d82d1ab12d9a6899

Observation bd040ad2-ddcd-4ea5-8881-49e96b1d3e82 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.893399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:37.893399Z digest=sha256:d1f31f002090a652499cf1ccf5a32b28e700637797ce454ff9bbac3e7e760733

Observation dac90df9-7ad8-44c4-9c78-95413555e016 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.594541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:37.954555Z digest=sha256:b7186250033a52557b338350a56a770b352088a311d6a0d561c36b127f95b13e

Observation 7f341fc4-d89b-4a78-bcab-887fe94274b5 · outbound

This paper cites Sypherd, R.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Sypherd, R

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.580769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.019369Z digest=sha256:95f924009f9ff9d3812a5d753d893c7ce8b54dd3324b411ad21630254f3661bb

Observation 9daebe3b-71e2-4a3b-bbe4-99ffe76e8547 · outbound

This paper cites Tunstall, E.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Tunstall, E

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.566845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.147360Z digest=sha256:b556da1516d54d4f7150bb76d6a73f4d1799292c885ab3e5de1da09026711039

Observation 88ca329d-0891-469e-9bfd-9ba8e26cdf17 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.553032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.268692Z digest=sha256:b8a8e9cf5fda8e186b67d93c787cc0ed3c8b5fec3ad70326d1e35a24ceeba088

Observation 50fd11cb-35ba-45bb-bacb-9fb4b9a50010 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.538286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.437209Z digest=sha256:2fb1e0763be1628f8dc02efb84ebc2d3b4ee18e0d0457c9efc120e86e0cb4a9f

Observation bfc3f260-911f-4edc-ac4c-2c45d148e19b · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.524264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.563817Z digest=sha256:1214cf7d1541772fdb30932c28dbdb8f84df90aff171b68e25e5153edece1e70

Observation adc61e49-10df-49f2-96e6-7f0df74fd2f1 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.510989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.682104Z digest=sha256:147240ee35d8cd9a56049302d86b9a9bcae8f0783db6aa3b3d940eb83d5bf4d9

Observation e22ddd89-11b9-4ec6-8db6-9870ea80c6b6 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.496283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.778837Z digest=sha256:a66ec11b9892c218a43cbec131b287951f69c07cd28d3db5bba906c8474f3d56

Observation 08a406d6-ea26-4385-ad56-1768411060d0 · outbound

This paper cites Qwen2.5-Omni Technical Report.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Qwen2.5-Omni Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.783043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.783043Z digest=sha256:9ac0e7d5cd1eedfece7350265df06414cac49e629dbcededefeb858ca12f54d6

Observation 87803603-1cad-408c-b14e-b59054fc6f47 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.480895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.787672Z digest=sha256:a34c971fbf39cb063954d81479a6e7f12abddb8d807a33e0a4aa99f2b1fee5a2

Observation 0881ed7b-d103-4605-bd91-21ebd84de778 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.465569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.791473Z digest=sha256:b2821ab83659cfde0fea9bf2ebcd36b3b379df75fac06f78aa6e3367dce70ab1

Observation ec96c192-f2de-446e-a537-58949b0618cb · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.451379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.795690Z digest=sha256:32137984ff9bb44a6e6af06d58aad7682ebeb29dc37f24091aa72b1f81594ccd

Observation 87ebc61f-baca-466e-855a-47b090309761 · outbound

This paper cites Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.799899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.799899Z digest=sha256:ecd0c59ee6984ed4ce795d046a5a5789277395e54f2a8842ec43382883cc1b6b

Observation 57dd9b03-2a44-46d3-b077-a2ffcc84e3d3 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.436613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.804815Z digest=sha256:3aaede3f4d4a640f27502274f2db4eb80aff0f86da0f1f2d946a12310f3dd88f

Observation 6b833df4-92a0-4977-88b7-29c850e08e7f · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.808890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.808890Z digest=sha256:87cae93730f4c063596fd6e509635f2e17e4cbf5fd20c44130ac7c6c98a9bfd5

Observation 0e55d8d0-b384-45ba-917f-38e4578e5718 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:52:39.421914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:52:38.813440Z digest=sha256:4cfebcaa54666679444765c763d95a6186866af9f106c71df73ff9606bbd8bf9

Observation c8c75fc2-f3d9-4982-935b-29a743119ac3 · outbound

This paper cites @esa (Ref.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) @esa (Ref

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.817370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.817370Z digest=sha256:b871cc150735a77b7f911f9a467a98a84f604c81e02128e4f0b7e3f3dee4e822

Observation acb13461-2315-4188-8636-15f03c2cd226 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.822165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.822165Z digest=sha256:cf745325e184f86e01d75468d7b464502802b2921adf6a5da14f9e498410ac4e

Observation df42e405-82d6-4b7c-92ba-e6aeddbf3765 · outbound

This paper cites an unresolved cited work.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.826438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.826438Z digest=sha256:01314ff44310c80d819ea34f15eb9fe5b2958db1cf1bebb1c91c1e86efaa942d

Pith citing papers

No inbound Pith citation observations are available.