Pith. sign in

Paper Citation Record · LEDGER

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

As of 8 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 5 inbound Pith citation observations for arXiv:2505.23749.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23749 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:49:12.618400Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:50:33.210790Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T10:16:28.587559Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact3
  • verified fuzzy33
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7deaf369-0115-498f-9e6f-35896dbed747 · outbound

This paper cites URL https://incidentdatabase.ai/.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? URL https://incidentdatabase.ai/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:21.735414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:04.050224Z digest=sha256:1ca3dbfe0988e7f51c5d8982a44fa3f0a4b48a578718e904410703e0c55297f0

Observation f17035f9-7e11-455c-a298-7639db870ecf · outbound

This paper cites Statistical methods for ranking data, volume 1341.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Statistical methods for ranking data, volume 1341

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:04.218743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:04.218743Z digest=sha256:5623116c4de99b22d4a85ac7fca7d377b1fee998a6221bd74579f407e51bb613

Observation f044d052-40b6-424c-b74d-4b9bf0b44fd6 · outbound

This paper cites Approximating optimal social choice under metric preferences.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Approximating optimal social choice under metric preferences

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:21.448648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:04.359091Z digest=sha256:a28077db99ed57db9997946b6a490ec9dd694ff0c82896ae715db74780ba725a

Observation 39a3ad23-7acb-4a70-aaec-668482b00c9e · outbound

This paper cites Voudouris.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Voudouris

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:21.210185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:04.481803Z digest=sha256:102e3a7d325124490d9a9beaa9026bdc552ca50cb8c7da6bd261a6c16110c921

Observation 281a5242-cd7a-4be5-aa31-39f775883da9 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? A general theoretical paradigm to understand learning from human preferences

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:21.040602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:04.663339Z digest=sha256:79ea7f759c832f7404c961834e80fc7dfecb04db7885b07be47048b06ed8c842

Observation 933601e4-587f-4584-a255-27fc293941ac · outbound

This paper cites A statistical decision-theoretic framework for social choice.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? A statistical decision-theoretic framework for social choice

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:20.920213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:04.811896Z digest=sha256:58fdc59ee6aa02b524417b08f09a0561c184511a685b024eff0ae0046c40d182

Observation 6e2ad6e4-b9b9-4bbb-8bfb-0f2cabdaf924 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:04.928905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:04.928905Z digest=sha256:ef23f8808a1cc8707ca357d15a38e871467b390b52a8029edf4b141399d6b4e4

Observation 124a96fa-0c60-489c-8166-cfaa13d61c76 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Constitutional AI: Harmlessness from AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:05.079571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:05.079571Z digest=sha256:6c93a6d8912a7080f246a0380f05dcf089ddb135876d0cddeec34f01a86bb4a1

Observation 4d3a6849-baec-456b-9c58-9dea8b789c5b · outbound

This paper cites Procaccia, and Nisarg Shah.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Procaccia, and Nisarg Shah

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:20.856174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:05.240930Z digest=sha256:e30b51c775acfec7ba4173d2d401d88fb30b63360f30b81604d444a46745f502

Observation a53979b3-731c-4b6b-9285-733542d78c26 · outbound

This paper cites Procaccia, and Or Sheffet.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Procaccia, and Or Sheffet

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:20.739548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:05.355996Z digest=sha256:24900615d2c4914289eec37c77491394206764a386d22272de1c1a3de4bbfeac

Observation 135d9ea8-8faf-4c42-88a4-8e136d720380 · outbound

This paper cites Human alignment of large language models through online preference optimisation.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Human alignment of large language models through online preference optimisation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:20.633947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:05.484886Z digest=sha256:8abbe93d092b621a7bd1fb64859e1ace607574ca7c04dc474a880dc161a3631b

Observation 74a561e3-88ef-4dfb-a4f0-09abee07dfa2 · outbound

This paper cites Procaccia.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Procaccia

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:20.517125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:05.608842Z digest=sha256:5b8ffe2ae80654a517ca78a767913d7b602b8adff2531e7be09cea6d00090b29

Observation ceee655c-706a-4bed-9b54-bddb236a2c17 · outbound

This paper cites MaxMin-RLHF: Alignment with Diverse Human Preferences.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:05.774124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:05.774124Z digest=sha256:c786249c0dc465db58f37b710f7f9568a550d13f9e965b78f8b717adad708aaf

Observation 1505f2ee-acb4-4ea2-b434-c085a5c3cf15 · outbound

This paper cites Breaking the Metric Voting Distortion Barrier.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Breaking the Metric Voting Distortion Barrier

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:20.338470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:05.937891Z digest=sha256:7cce69b95c5b593769fe048bc13c218ad8f884109b8e45820f4a0e7bae2f1975

Observation 657b84e8-a411-47bf-b3d9-913419ba6b05 · outbound

This paper cites PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:06.059537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:06.059537Z digest=sha256:07f200bb6426f4cb90b93327d5af0bc68f3a223edf48eae6b57108f8dfe696c3

Observation b656a40b-4940-4c9c-9452-a3602d222e2b · outbound

This paper cites Chatbot arena: An open platform for evaluating llms by human preference.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Chatbot arena: An open platform for evaluating llms by human preference

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:06.193799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:06.193799Z digest=sha256:bf4be51f54bc9673df142dc04ff1691fe8cedb2d17ab4849f43929de7650ab6b

Observation bba5fc7f-0794-4c31-95f9-2e38d7b9a17d · outbound

This paper cites Direct preference optimization with unobserved preference heterogeneity.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Direct preference optimization with unobserved preference heterogeneity

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:06.339837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:06.339837Z digest=sha256:1c6c584e0214edc4eb5460f9811a9f77de10e07033e1895b3868c16fcc635b8c

Observation f3287c2d-66b1-4ba7-8719-2bad0b0111cd · outbound

This paper cites Deep reinforcement learning from human preferences.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Deep reinforcement learning from human preferences

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:06.486162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:06.486162Z digest=sha256:241e5a22e3295b732fba34491f17eb336a9f0c4a01aacc554afeaf2e3de3b0c3

Observation 3f635759-00d9-4f10-a602-8e4f40f0544c · outbound

This paper cites Common voting rules as maximum likelihood estimators.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Common voting rules as maximum likelihood estimators

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:20.085504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:06.622030Z digest=sha256:e8e4948acc14ee59600d64243a3cf9059d25197e7b36495c5912cd61a620335d

Observation 53035af2-5446-4811-8be4-e76807208457 · outbound

This paper cites Position: social choice should guide ai alignment in dealing with diverse human feedback.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Position: social choice should guide ai alignment in dealing with diverse human feedback

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:19.828790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:06.755545Z digest=sha256:fb7fbd72f35f1f09983cc5a8cbdf9746c7d0eb2be0b6c27a68084b585f3896bb

Observation 8255786f-6863-4923-b588-b353e24419ae · outbound

This paper cites Mapping social choice theory to RLHF.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Mapping social choice theory to RLHF

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:19.625540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:06.899719Z digest=sha256:701749c1b3576e7ebbd941e31d2f19216cdfdb1e982ec2f05ccd8fa9e4a45c5f

Observation f930d716-8fae-4698-a4a6-979d1ee73a3c · outbound

This paper cites Metric distortion with elicited pairwise comparisons.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Metric distortion with elicited pairwise comparisons

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:19.389157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:07.049863Z digest=sha256:33912615c2df7790551200b6401fb839c44c7b756fc7a6db296eb2646f331698

Observation f241ac89-4777-4b40-9923-727fd10053f3 · outbound

This paper cites Optimized Distortion and Proportional Fairness in Voting.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Optimized Distortion and Proportional Fairness in Voting

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:19.136531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:07.201291Z digest=sha256:1ce9716daadf0f6257ade09d58a0a456fb3115e96f2e881e5726e09bf8df5c28

Observation ed957c40-ce0e-4abd-8715-ee0b64d970e2 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? KTO: Model Alignment as Prospect Theoretic Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:07.327822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:07.327822Z digest=sha256:8a61a426e8907653ee86eb3231fe75c418de3525a799e76994f4dd5e2afeac90

Observation c4d7ac5d-9d6d-4b59-b981-cf5524b77920 · outbound

This paper cites Fishburn.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Fishburn

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:18.855523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:07.517409Z digest=sha256:9dbff19a0b4a831466ef4a7742fca9eecc73f074e8a54d9c7981e3d273709574

Observation 31fb267e-8df1-492f-8210-f97849efe89a · outbound

This paper cites Distortion Under Public-Spirited Voting.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Distortion Under Public-Spirited Voting

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:18.619592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:07.658976Z digest=sha256:815ae338450047d4633e032d7775a8e4b4630f488838bae0e122219c14fcaab1

Observation 4c3300b7-4f7e-47ce-81db-41a4f3551c64 · outbound

This paper cites Axioms for AI Alignment from Human Feedback.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Axioms for AI Alignment from Human Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:07.791066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:07.791066Z digest=sha256:0c6d7cf4506df4cc2075e7d2e7b851bb8b0624ff3e44a961164953f2fccdba43

Observation 443802cf-84bc-46c4-a923-a1cc58f2e459 · outbound

This paper cites Resolving the optimal metric distortion conjecture.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Resolving the optimal metric distortion conjecture

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:18.318623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:07.927462Z digest=sha256:d4408d13afc01a01ac3bd89efd0a570acd5126547f227714605bf05458ab898a

Observation d98f4e0b-79c1-45e8-af3a-b78a486267f6 · outbound

This paper cites Metric distortion Under Probabilistic Voting.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Metric distortion Under Probabilistic Voting

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:49:13.559350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:08.065470Z digest=sha256:5b660760b7c2caceba6795c7a1cf01ea97aed252f86a098ec92765fc317cfe8c

Observation b55ea79f-9f8e-422d-bb49-586c347b5e54 · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:08.217635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:08.217635Z digest=sha256:6f1d0aa978f40ecc0a27dd4c5e3fe472b9a707ea2480d7573a78be5be8d91130

Observation 187a5caa-f303-427a-ab77-a8f27d4a248c · outbound

This paper cites Plurality Veto : A Simple Voting Rule Achieving Optimal Metric Distortion.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Plurality Veto : A Simple Voting Rule Achieving Optimal Metric Distortion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:17.996032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:08.372790Z digest=sha256:92a8f185458765d690f9022ba517491059c38fea0a652f7d62d1bb1f24b3b0a8

Observation af056ee5-f7c8-4972-b936-4e9d118cbfdf · outbound

This paper cites Provably mitigating overoptimization in RLHF : Your SFT loss is implicitly an adversarial regularizer.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Provably mitigating overoptimization in RLHF : Your SFT loss is implicitly an adversarial regularizer

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:17.708990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:08.506337Z digest=sha256:69f27be39e16195f7bef7cfc483478eb273ca2e87dd1fea1a69b55c622563457

Observation 34e55da6-758a-4f5d-8540-f8d6d76bcb01 · outbound

This paper cites Jackpot! Alignment as a Maximal Lottery.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Jackpot! Alignment as a Maximal Lottery

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:08.695717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:08.695717Z digest=sha256:808577147dc2d3836c335044a05a802540d33032810f204acbf988ad75e15994

Observation 380fc395-09fe-4f14-86a5-85e5e85f348c · outbound

This paper cites SimPO : Simple preference optimization with a reference-free reward.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? SimPO : Simple preference optimization with a reference-free reward

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:17.358461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:08.835075Z digest=sha256:cb7107ffe3385ed703421a8ed1a8e765888a41e3519de223e78158024d755331

Observation eec2a541-f21b-427a-832c-569c9656b9d6 · outbound

This paper cites AI Alignment and Social Choice: Fundamental Limitations and Policy Implications.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? AI Alignment and Social Choice: Fundamental Limitations and Policy Implications

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:08.964242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:08.964242Z digest=sha256:50a6bfd8675333934f0bfc7f19a4b74e46a983c229da9bb477ca642a231da299

Observation 0cdd7021-dfb3-48b7-beed-bb8e171de671 · outbound

This paper cites Nash learning from human feedback.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Nash learning from human feedback

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:17.072272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:09.099448Z digest=sha256:90ab0ec954cfdd86efb1bdb3a93d13507f0c5624846f44741e3daf0c8093aaff

Observation ec4ebb3b-c40d-4cb6-a5b3-2407794f695d · outbound

This paper cites Axioms for learning from pairwise comparisons.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Axioms for learning from pairwise comparisons

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:16.748259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:09.270503Z digest=sha256:f76e1062c03b9928ec88d84462097b4a604c6672c4fb352682dc6eb649754e21

Observation d3e44c7f-c6df-4086-ba1e-66ccb2066cb7 · outbound

This paper cites Training language models to follow instructions with human feedback.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Training language models to follow instructions with human feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:09.400159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:09.400159Z digest=sha256:725efbaf39440f6578eec5a827d63239309ff4ddc0217b3efb108f3910d631fd

Observation a3775ccd-ef83-4694-be14-831abd0f9ccf · outbound

This paper cites RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:09.533133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:09.533133Z digest=sha256:d7dd385b46668b39cc4152574f65c2bbec0aa88ba834d88ad647a650e973ac83

Observation 3dfc14ac-9814-4f7f-aaf3-284253afea04 · outbound

This paper cites Personalizing reinforcement learning from human feedback with variational preference learning.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Personalizing reinforcement learning from human feedback with variational preference learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:16.459944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:09.655174Z digest=sha256:def3723503dd1199923cde63341abde53f97efdf57a658473fdac7458044af5d

Observation f0803d3b-e7d3-4d2e-b241-01c425f39311 · outbound

This paper cites Procaccia and Jeffrey S.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Procaccia and Jeffrey S

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:16.197248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:09.778185Z digest=sha256:710b091f7732a74bb51a45ab110434d2a1bc1d4c2b7c5c608017b79a089e9f59

Observation f9158f3a-d4a9-4615-8655-ff963d98d703 · outbound

This paper cites Clone-Robust AI Alignment.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Clone-Robust AI Alignment

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:49:13.263112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:09.960375Z digest=sha256:d9addef394e23649911866ed98e9305d7e6cb7b303f2392feb6d721ba85db9e6

Observation a9b96989-7746-4bd0-a0a4-71dd9b38c846 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Direct preference optimization: Your language model is secretly a reward model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:10.086522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:10.086522Z digest=sha256:dfe6f5f54e0933bdd2ad1acaacd000484ff63cd443736049311a870da7f4e0ac

Observation 32346169-d22e-45e7-964e-082a9f96a582 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Proximal Policy Optimization Algorithms

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:10.234824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:10.234824Z digest=sha256:25dace3929549b8e7ba1918ae9b15e2c4c7b149f7ce40ed2bd23526ed823cc45

Observation 55c70479-b0f0-423f-9c06-9d3c7fb55913 · outbound

This paper cites Direct Alignment with Heterogeneous Preferences.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Direct Alignment with Heterogeneous Preferences

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:10.365118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:10.365118Z digest=sha256:6514d71bda9ddbe2e518b8bc961f1c14894ebd1c09cd8cdd429d745160983b56

Observation 5d19d2f2-1c7b-4921-a9fd-08f0ea002ba9 · outbound

This paper cites Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:10.543390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:10.543390Z digest=sha256:3f45b5fbb816c1cf90667fa65f46cdec2c205f169ebd01f9c95f66fd7f26964a

Observation 85be41e3-bb23-401c-9146-ccc3194f7a51 · outbound

This paper cites Position: a roadmap to pluralistic alignment.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Position: a roadmap to pluralistic alignment

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:15.958217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:10.662526Z digest=sha256:4ecae268bb5a93ab722072bcf5ceb8eb497c6e0f594a71afb4824dd1ec128ae3

Observation cd7f7b4b-28b6-43e8-b9aa-3ae3317ff2c8 · outbound

This paper cites A Minimaximalist Approach to Reinforcement Learning from Human Feedback.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:10.847272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:10.847272Z digest=sha256:567f4e3f88069ec819be692207907342fd5c7c926df9a5490d81658e3075fe87

Observation d8c9c893-73b3-4dc6-afdf-e447d80452e0 · outbound

This paper cites Learning populations of preferences via pairwise comparison queries.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Learning populations of preferences via pairwise comparison queries

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:15.702859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:11.025830Z digest=sha256:0abb8fe8df5fe0485f181737de707ed829c06f6d1507525ad76575807d63efe5

Observation 2389c91d-76bf-4b95-a76a-2daf0a194396 · outbound

This paper cites Is RLHF more difficult than standard RL ? a theoretical perspective.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Is RLHF more difficult than standard RL ? a theoretical perspective

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:15.379770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:11.141013Z digest=sha256:6f4aafce130a8fffb2d9e2a309df15346848e15460bd5ad68448847b9ae852e5

Observation 73729af5-8349-48e5-b23e-46786efbb058 · outbound

This paper cites Metric learning from limited pairwise preference comparisons.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Metric learning from limited pairwise preference comparisons

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:15.107506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:11.304347Z digest=sha256:edb0605c9b64917df329e8178a5d38b5ca9c26fb5156b72c9b501153f39b9c44

Observation f4cbf904-d1a0-4470-bc7c-2ff692979b15 · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Self-Play Preference Optimization for Language Model Alignment

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:11.460500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:11.460500Z digest=sha256:7e1ce051aa8f62d3e0304b2ffa3dd015f86c833ec669523d1cad1f7bc7e9fcfc

Observation 72cd5d9d-8c3e-4ba5-b8b9-c87ae757715d · outbound

This paper cites Bayesian estimators as voting rules.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Bayesian estimators as voting rules

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:14.853647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:11.576108Z digest=sha256:3dd87a544f71dcc9c8b2765a47498163a126be0213772b0c9f2105fedceb8e65

Observation 95b897aa-fda4-43fc-9d1e-beac3deecb61 · outbound

This paper cites Learning and decision-making from rank data.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Learning and decision-making from rank data

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:14.606243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:11.774245Z digest=sha256:fc7a2cec108ce9801897104ab9bacd073968d05cf2b7c44a5bf93125963a3e85

Observation 4bdc156b-1a8c-44a1-927b-62721605be01 · outbound

This paper cites On the identifiability of mixtures of ranking models.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? On the identifiability of mixtures of ranking models

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:49:12.965798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:11.935144Z digest=sha256:54e9393d388483b8c13ade9ff3883e5808543e6f71070d669f16aed9d0e1aab2

Observation 2295f788-56e7-4694-a0bc-4acd973fa490 · outbound

This paper cites Learning mixtures of plackett-luce models from structured partial orders.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Learning mixtures of plackett-luce models from structured partial orders

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:14.337018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:12.133491Z digest=sha256:e33902e44d8d29799a7b26ca0c4d29c0ed4f0b8423c3485e3cc603c5d0011682

Observation a108ef03-f17e-45ac-a714-4404501048aa · outbound

This paper cites Learning mixtures of plackett-luce models.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Learning mixtures of plackett-luce models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:14.038561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:49:12.283875Z digest=sha256:e9539f072c24617bcf77eaf08289c1ed8aa1829b2e60f64209f39d89b1968e4f

Observation 86890663-deca-4538-9ae6-381b1a419079 · outbound

This paper cites Provable Multi-Party Reinforcement Learning with Diverse Human Feedback.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Provable Multi-Party Reinforcement Learning with Diverse Human Feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:12.458231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:12.458231Z digest=sha256:b556a4447e993fc1c37c1274e95e18f12a9bb237fc8ee7e9e965e9f98493a1a3

Observation 23373edc-8627-46d1-a076-7e35ca23979c · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Fine-Tuning Language Models from Human Preferences

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:12.618400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:12.618400Z digest=sha256:9286bebf9a93c21ba4b9ddbd283b0d72a1d8e0e604086d6bf31f9d8aa2f7d86e

Pith citing papers

Observation 0c2bb61c-f6a6-4a06-8322-dafc461ed06d · inbound

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers cites this paper.

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-07T22:50:33.210790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:50:33.210790Z digest=sha256:6da4290e65556d426d34061a5093c83920b237bb4b349abd500c734ff5d684d8

Observation 551547b7-8d69-4119-a622-2bb0160cf755 · inbound

Power and Limitations of Aggregation in Compound AI Systems cites this paper.

Power and Limitations of Aggregation in Compound AI Systems Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T21:09:50.286275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:09:50.286275Z digest=sha256:e04f70a059fe38e5ebfb88de45b8a1f0e5d28f8120e7997ad1d58a2d75d8058a

Observation f45ffd9c-70ac-4c7b-857c-6846556304ef · inbound

Mind the Gap: Structure-Aware Consistency in Preference Learning cites this paper.

Mind the Gap: Structure-Aware Consistency in Preference Learning Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.589799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-07T06:41:18.311986Z digest=sha256:f19812f38afc3feb11b6384aa71f30d4b6cbf69e2e5091a26d191e6b8c76eaa0

Observation 1da2f02c-27f8-440b-8fa5-c2ad82146aba · inbound

Internal Pluralism and the Limits of Pairwise Comparisons cites this paper.

Internal Pluralism and the Limits of Pairwise Comparisons Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

Reference 133

Resolution
unresolved
no resolver link, observed 2026-07-12T07:49:57.204875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T07:49:57.204875Z digest=sha256:61bcc659f1c07290db1665a9e631809613005adf352a40da3785dca37bbbf35e

Observation 0da2bd54-9479-4abd-a04d-c364a1e6eb54 · inbound

Internal Pluralism and the Limits of Pairwise Comparisons cites this paper.

Internal Pluralism and the Limits of Pairwise Comparisons Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

Reference 222

Resolution
unresolved
no resolver link, observed 2026-08-02T09:03:15.212226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:03:15.212226Z digest=sha256:da11f3fea8ddcc1efae21369b43c3416ddddaca8aa94fea0bb3f723e119187d4