Pith. sign in

Paper Citation Record · LEDGER

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

As of 23 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 5 inbound Pith citation observations for arXiv:2505.23749.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23749 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:49:12.618400Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:50:33.210790Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T10:16:28.587559Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact3
  • verified fuzzy33
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7deaf369-0115-498f-9e6f-35896dbed747 · outbound

This paper cites URL https://incidentdatabase.ai/.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? URL https://incidentdatabase.ai/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:21.735414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:04.050224Z digest=sha256:ce32af429eb8ac2cde2c7780f270c52327d17a25b2b67c8f9671bb3616581b6f

Observation f17035f9-7e11-455c-a298-7639db870ecf · outbound

This paper cites Statistical methods for ranking data, volume 1341.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Statistical methods for ranking data, volume 1341

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:04.218743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:04.218743Z digest=sha256:cee39d981ab27c68e82d640f83e853a2afc1073ce8d9d00c6ed3a503cd4594e9

Observation f044d052-40b6-424c-b74d-4b9bf0b44fd6 · outbound

This paper cites Approximating optimal social choice under metric preferences.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Approximating optimal social choice under metric preferences

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:21.448648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:04.359091Z digest=sha256:0df4971153948c9af8ceeee4f9c3e6eeaf157eb8d79c86f7339d2cc83f90e80e

Observation 39a3ad23-7acb-4a70-aaec-668482b00c9e · outbound

This paper cites Voudouris.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Voudouris

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:21.210185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:04.481803Z digest=sha256:01c238b4b66b6e54183d8b78810253659fe399c1818faca6bfc9b61dd81a7dc1

Observation 281a5242-cd7a-4be5-aa31-39f775883da9 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? A general theoretical paradigm to understand learning from human preferences

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:21.040602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:04.663339Z digest=sha256:4923dae182cb9c3b2b0f92f67aa39d9839a48ad9f5d2867a833223196b50d2ac

Observation 933601e4-587f-4584-a255-27fc293941ac · outbound

This paper cites A statistical decision-theoretic framework for social choice.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? A statistical decision-theoretic framework for social choice

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:20.920213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:04.811896Z digest=sha256:2318ee6e4523961a85b4bbb7db2edcfe86d1ccdd903bcf3429eafdbb38cb484a

Observation 6e2ad6e4-b9b9-4bbb-8bfb-0f2cabdaf924 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:04.928905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:04.928905Z digest=sha256:21b3f37644e5a24e722d3681e9b470100f4de9d7e2b671cd62ec7ac11efd7404

Observation 124a96fa-0c60-489c-8166-cfaa13d61c76 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Constitutional AI: Harmlessness from AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:05.079571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:05.079571Z digest=sha256:3c7ccba92532995513f4f60359fd39b5cf47c09c82666ca55e9ecff5a93c1ed0

Observation 4d3a6849-baec-456b-9c58-9dea8b789c5b · outbound

This paper cites Procaccia, and Nisarg Shah.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Procaccia, and Nisarg Shah

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:20.856174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:05.240930Z digest=sha256:d045c1d21fa8ee08694bb20a6c8074f2c9f25a9576e8a4341ae91612a11e9bde

Observation a53979b3-731c-4b6b-9285-733542d78c26 · outbound

This paper cites Procaccia, and Or Sheffet.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Procaccia, and Or Sheffet

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:20.739548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:05.355996Z digest=sha256:86539769eab81fd88e89c1335fb78fc724176500e71f867e76b1bb706bda5311

Observation 135d9ea8-8faf-4c42-88a4-8e136d720380 · outbound

This paper cites Human alignment of large language models through online preference optimisation.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Human alignment of large language models through online preference optimisation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:20.633947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:05.484886Z digest=sha256:93e5a34234e75687d441f4ce13223eb33f21a9b2307bc3a4578e5c90fb4c4960

Observation 74a561e3-88ef-4dfb-a4f0-09abee07dfa2 · outbound

This paper cites Procaccia.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Procaccia

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:20.517125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:05.608842Z digest=sha256:3da4cc9b7031e525423ea86697608a87802dfef3e1126d5ac05069827e67e7e0

Observation ceee655c-706a-4bed-9b54-bddb236a2c17 · outbound

This paper cites MaxMin-RLHF: Alignment with Diverse Human Preferences.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:05.774124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:05.774124Z digest=sha256:1d3666a17ec76147561db52767a3b9ad3144e8fd11efd3f303ebac82ba56ca03

Observation 1505f2ee-acb4-4ea2-b434-c085a5c3cf15 · outbound

This paper cites Breaking the Metric Voting Distortion Barrier.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Breaking the Metric Voting Distortion Barrier

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:20.338470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:05.937891Z digest=sha256:6ff52cd21aa90ba28f1f80bd595985d2574ce455090f3df0d65d583548a8e645

Observation 657b84e8-a411-47bf-b3d9-913419ba6b05 · outbound

This paper cites PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:06.059537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:06.059537Z digest=sha256:223a5788f4813d9a229c88e8ae143461d704b832fab5590972411186d8df93eb

Observation b656a40b-4940-4c9c-9452-a3602d222e2b · outbound

This paper cites Chatbot arena: An open platform for evaluating llms by human preference.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Chatbot arena: An open platform for evaluating llms by human preference

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:06.193799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:06.193799Z digest=sha256:e6afe4967a364da5e806ce11905a865e6659cd358225f0358cb1ecc008b9e241

Observation bba5fc7f-0794-4c31-95f9-2e38d7b9a17d · outbound

This paper cites Direct preference optimization with unobserved preference heterogeneity.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Direct preference optimization with unobserved preference heterogeneity

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:06.339837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:06.339837Z digest=sha256:6a997a7e1dbf55e83411e62a619171b0bcc6a407d5584a327c539d825d26f0fa

Observation f3287c2d-66b1-4ba7-8719-2bad0b0111cd · outbound

This paper cites Deep reinforcement learning from human preferences.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Deep reinforcement learning from human preferences

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:06.486162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:06.486162Z digest=sha256:a2dd80bffafee82cee7744014cb757a042ca51684ca9a548247bc7f4eb9a4613

Observation 3f635759-00d9-4f10-a602-8e4f40f0544c · outbound

This paper cites Common voting rules as maximum likelihood estimators.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Common voting rules as maximum likelihood estimators

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:20.085504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:06.622030Z digest=sha256:6de459b95a93be117187c5eb9f4ae388f0431a473e1ba0ff1e505f1d70d417bc

Observation 53035af2-5446-4811-8be4-e76807208457 · outbound

This paper cites Position: social choice should guide ai alignment in dealing with diverse human feedback.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Position: social choice should guide ai alignment in dealing with diverse human feedback

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:19.828790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:06.755545Z digest=sha256:29c650a1fef8c30c1ae3e3dbeb38dd2c82058c3c23c44972c724a06a693b813f

Observation 8255786f-6863-4923-b588-b353e24419ae · outbound

This paper cites Mapping social choice theory to RLHF.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Mapping social choice theory to RLHF

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:19.625540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:06.899719Z digest=sha256:3af26db35bf149ebac345c36df1b474daef6cc83c2895e48e06e32d9b8807afc

Observation f930d716-8fae-4698-a4a6-979d1ee73a3c · outbound

This paper cites Metric distortion with elicited pairwise comparisons.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Metric distortion with elicited pairwise comparisons

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:19.389157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:07.049863Z digest=sha256:28b631eeb9d24a0c707b849698936dd622c44c4ea796c506216f119cb103e266

Observation f241ac89-4777-4b40-9923-727fd10053f3 · outbound

This paper cites Optimized Distortion and Proportional Fairness in Voting.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Optimized Distortion and Proportional Fairness in Voting

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:19.136531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:07.201291Z digest=sha256:4d31d7f3f9ccc577716c3d8f3dfd95d982dac14be41ebfc8127531ffe0d8a37a

Observation ed957c40-ce0e-4abd-8715-ee0b64d970e2 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? KTO: Model Alignment as Prospect Theoretic Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:07.327822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:07.327822Z digest=sha256:0d11daaec3565c4105e1548622645b57412c12480ed4773bce5771e3207dcd94

Observation c4d7ac5d-9d6d-4b59-b981-cf5524b77920 · outbound

This paper cites Fishburn.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Fishburn

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:18.855523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:07.517409Z digest=sha256:00b2269dba61b64559049b8858478df958e60d77ed38e463195824fc456feeb8

Observation 31fb267e-8df1-492f-8210-f97849efe89a · outbound

This paper cites Distortion Under Public-Spirited Voting.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Distortion Under Public-Spirited Voting

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:18.619592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:07.658976Z digest=sha256:da810a4dc02cd16ef410147da5a764879792a76b6e4f83cf0c347c7ae3b775b2

Observation 4c3300b7-4f7e-47ce-81db-41a4f3551c64 · outbound

This paper cites Axioms for AI Alignment from Human Feedback.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Axioms for AI Alignment from Human Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:07.791066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:07.791066Z digest=sha256:cafc34862b8e62202c0a23228b852d7886c6d5f46ac1f661fdaba87749a2fef9

Observation 443802cf-84bc-46c4-a923-a1cc58f2e459 · outbound

This paper cites Resolving the optimal metric distortion conjecture.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Resolving the optimal metric distortion conjecture

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:18.318623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:07.927462Z digest=sha256:6525638193c6a03297e93f0d4da620c4fd6a07597cbf13ec0d880e3edf01ca51

Observation d98f4e0b-79c1-45e8-af3a-b78a486267f6 · outbound

This paper cites Metric distortion Under Probabilistic Voting.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Metric distortion Under Probabilistic Voting

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:49:13.559350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:08.065470Z digest=sha256:97ec9ed8f405996a6a1e7cc9daa330224efff74440a9e8d88f1e40df7ee0dc76

Observation b55ea79f-9f8e-422d-bb49-586c347b5e54 · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:08.217635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:08.217635Z digest=sha256:925e59fc859e4bfbbcaa62ccc9941a759c38508f697071240c6e402cac567a3c

Observation 187a5caa-f303-427a-ab77-a8f27d4a248c · outbound

This paper cites Plurality Veto : A Simple Voting Rule Achieving Optimal Metric Distortion.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Plurality Veto : A Simple Voting Rule Achieving Optimal Metric Distortion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:17.996032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:08.372790Z digest=sha256:9736b278c98c11e4ea96e42b98e2b5c4a7494a14969ecd0bc27d2d478d664de3

Observation af056ee5-f7c8-4972-b936-4e9d118cbfdf · outbound

This paper cites Provably mitigating overoptimization in RLHF : Your SFT loss is implicitly an adversarial regularizer.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Provably mitigating overoptimization in RLHF : Your SFT loss is implicitly an adversarial regularizer

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:17.708990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:08.506337Z digest=sha256:c0f3bdd3d7704f814f85bec55c8f6e49ca9a3a4e9f41cef6b809c840e186d595

Observation 34e55da6-758a-4f5d-8540-f8d6d76bcb01 · outbound

This paper cites Jackpot! Alignment as a Maximal Lottery.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Jackpot! Alignment as a Maximal Lottery

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:08.695717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:08.695717Z digest=sha256:2f75b0510dfac504f6025a8f4068a1a3b257977409739b964d68959e2f5af363

Observation 380fc395-09fe-4f14-86a5-85e5e85f348c · outbound

This paper cites SimPO : Simple preference optimization with a reference-free reward.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? SimPO : Simple preference optimization with a reference-free reward

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:17.358461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:08.835075Z digest=sha256:655a080962782cae0a098567631e109cb932fa63aca5c0c99b64861534ca4235

Observation eec2a541-f21b-427a-832c-569c9656b9d6 · outbound

This paper cites AI Alignment and Social Choice: Fundamental Limitations and Policy Implications.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? AI Alignment and Social Choice: Fundamental Limitations and Policy Implications

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:08.964242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:08.964242Z digest=sha256:451bbf1052b021363dff0223ee9285cccb08d9f088192032f905d186802a923a

Observation 0cdd7021-dfb3-48b7-beed-bb8e171de671 · outbound

This paper cites Nash learning from human feedback.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Nash learning from human feedback

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:17.072272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:09.099448Z digest=sha256:c5d3faf938ad7d3e7ee939309ae107ff562dcdce2a7a598a0938128395acc882

Observation ec4ebb3b-c40d-4cb6-a5b3-2407794f695d · outbound

This paper cites Axioms for learning from pairwise comparisons.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Axioms for learning from pairwise comparisons

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:16.748259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:09.270503Z digest=sha256:895bccafb2bce263887489460df6705bcf458f656621794e93ab0d21b4758e95

Observation d3e44c7f-c6df-4086-ba1e-66ccb2066cb7 · outbound

This paper cites Training language models to follow instructions with human feedback.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Training language models to follow instructions with human feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:09.400159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:09.400159Z digest=sha256:f3077c441d2b68298e8500b1a21fa6f965fe6b755a4a1d6f00d30d27f8a68eda

Observation a3775ccd-ef83-4694-be14-831abd0f9ccf · outbound

This paper cites RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:09.533133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:09.533133Z digest=sha256:d7881ea4620d92c9ce57dff81afbec4f2aa888b56673c2f7e5b8c07e9121c6bf

Observation 3dfc14ac-9814-4f7f-aaf3-284253afea04 · outbound

This paper cites Personalizing reinforcement learning from human feedback with variational preference learning.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Personalizing reinforcement learning from human feedback with variational preference learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:16.459944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:09.655174Z digest=sha256:593dcba4aa78e81b5cef050b486c567cf05fec4996258dc935242cd6746a4857

Observation f0803d3b-e7d3-4d2e-b241-01c425f39311 · outbound

This paper cites Procaccia and Jeffrey S.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Procaccia and Jeffrey S

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:16.197248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:09.778185Z digest=sha256:191ddb71e68f4b6d289fbf9d019645218caa16773ee11951adef4dec2933e8c3

Observation f9158f3a-d4a9-4615-8655-ff963d98d703 · outbound

This paper cites Clone-Robust AI Alignment.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Clone-Robust AI Alignment

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:49:13.263112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:09.960375Z digest=sha256:dfadc16229b9fe363a594fd2ef80b0ab979ddf285d308c7a3364b96bffb74a5d

Observation a9b96989-7746-4bd0-a0a4-71dd9b38c846 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Direct preference optimization: Your language model is secretly a reward model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:10.086522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:10.086522Z digest=sha256:043448b65dcb574d442a01e9cbdb40f26323d8be1a099b6af9081473728942d2

Observation 32346169-d22e-45e7-964e-082a9f96a582 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Proximal Policy Optimization Algorithms

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:10.234824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:10.234824Z digest=sha256:38f06f4f3a729d68a084b256d8b45c1e1393001c03c8a73ca68174f3b90b93ac

Observation 55c70479-b0f0-423f-9c06-9d3c7fb55913 · outbound

This paper cites Direct Alignment with Heterogeneous Preferences.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Direct Alignment with Heterogeneous Preferences

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:10.365118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:10.365118Z digest=sha256:fc21d4006a0977da162ef7b928f71a56ec3f1e5e8afc579b1cee6b71f39b2e31

Observation 5d19d2f2-1c7b-4921-a9fd-08f0ea002ba9 · outbound

This paper cites Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:10.543390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:10.543390Z digest=sha256:647ae5612f20eab916c4e99ba949f2f41339b9178f2d006cb1f54e55f149e52b

Observation 85be41e3-bb23-401c-9146-ccc3194f7a51 · outbound

This paper cites Position: a roadmap to pluralistic alignment.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Position: a roadmap to pluralistic alignment

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:15.958217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:10.662526Z digest=sha256:44b0fd407a6e02ebbdc498177c3e95011c8f491bc0536938b007e684fa312639

Observation cd7f7b4b-28b6-43e8-b9aa-3ae3317ff2c8 · outbound

This paper cites A Minimaximalist Approach to Reinforcement Learning from Human Feedback.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:10.847272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:10.847272Z digest=sha256:f141f0de9af7608ebcda28bffa9b7ee3971427e4852a84b8c663d88625ddb42b

Observation d8c9c893-73b3-4dc6-afdf-e447d80452e0 · outbound

This paper cites Learning populations of preferences via pairwise comparison queries.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Learning populations of preferences via pairwise comparison queries

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:15.702859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:11.025830Z digest=sha256:e636b1f834e367d8210777965b1b26ccea1dd0e824515793ae022898856bd81b

Observation 2389c91d-76bf-4b95-a76a-2daf0a194396 · outbound

This paper cites Is RLHF more difficult than standard RL ? a theoretical perspective.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Is RLHF more difficult than standard RL ? a theoretical perspective

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:15.379770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:11.141013Z digest=sha256:35dd53f53b279a69d54c2b2db0f0a08db18901d554674c86e02f0d100c52c54f

Observation 73729af5-8349-48e5-b23e-46786efbb058 · outbound

This paper cites Metric learning from limited pairwise preference comparisons.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Metric learning from limited pairwise preference comparisons

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:15.107506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:11.304347Z digest=sha256:31fd32311442dddad217ce0598580b66a293a8e317287e41cbd6c8172d60ff8c

Observation f4cbf904-d1a0-4470-bc7c-2ff692979b15 · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Self-Play Preference Optimization for Language Model Alignment

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:11.460500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:11.460500Z digest=sha256:74b47f330195bcd11bc2a531355b6b84577bce0a3f04747223e36626225387e8

Observation 72cd5d9d-8c3e-4ba5-b8b9-c87ae757715d · outbound

This paper cites Bayesian estimators as voting rules.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Bayesian estimators as voting rules

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:14.853647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:11.576108Z digest=sha256:eade64275118f00ad3c3a4e283549262b9249c2bba97f196babc4ea99a91d8bb

Observation 95b897aa-fda4-43fc-9d1e-beac3deecb61 · outbound

This paper cites Learning and decision-making from rank data.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Learning and decision-making from rank data

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:14.606243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:11.774245Z digest=sha256:93af44aa1a984f3e9112cb2584fd72a2f8ec51ec67f45082928c8fca79549041

Observation 4bdc156b-1a8c-44a1-927b-62721605be01 · outbound

This paper cites On the identifiability of mixtures of ranking models.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? On the identifiability of mixtures of ranking models

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:49:12.965798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:11.935144Z digest=sha256:d6f39aaf96fa4503c7cfa9f8e9a58304effde673a2261ca4b47d85984d66223d

Observation 2295f788-56e7-4694-a0bc-4acd973fa490 · outbound

This paper cites Learning mixtures of plackett-luce models from structured partial orders.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Learning mixtures of plackett-luce models from structured partial orders

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:14.337018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:12.133491Z digest=sha256:9db973c659e07ec7eccac6ab05de29c17b1adeaba00b5c8264ab85d902eb414a

Observation a108ef03-f17e-45ac-a714-4404501048aa · outbound

This paper cites Learning mixtures of plackett-luce models.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Learning mixtures of plackett-luce models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:49:14.038561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:49:12.283875Z digest=sha256:1921f95f2aacac1e60d596774dff7e269a11f5d34e43cd1d34e131c4530e19cf

Observation 86890663-deca-4538-9ae6-381b1a419079 · outbound

This paper cites Provable Multi-Party Reinforcement Learning with Diverse Human Feedback.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Provable Multi-Party Reinforcement Learning with Diverse Human Feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:12.458231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:12.458231Z digest=sha256:0e7325f3336d132583827d5f0f7098e0cd34a107e437f4bba7733774d9437c87

Observation 23373edc-8627-46d1-a076-7e35ca23979c · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Fine-Tuning Language Models from Human Preferences

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:12.618400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:12.618400Z digest=sha256:dbf3940b9ec4b729e54ac13dbe05d841e818bfca06842cd89f3b527a544360d0

Pith citing papers

Observation 0c2bb61c-f6a6-4a06-8322-dafc461ed06d · inbound

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers cites this paper.

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-07T22:50:33.210790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:50:33.210790Z digest=sha256:85e12d397e06128e23492d6dc90cd1e08be03ad0d443054df5f28a5656e9c025

Observation 551547b7-8d69-4119-a622-2bb0160cf755 · inbound

Power and Limitations of Aggregation in Compound AI Systems cites this paper.

Power and Limitations of Aggregation in Compound AI Systems Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T21:09:50.286275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:09:50.286275Z digest=sha256:020bfb57b1bdf0558cac0cc3dd94f0d2c8a4a8c34302a92e4e0d767a93fc62ef

Observation f45ffd9c-70ac-4c7b-857c-6846556304ef · inbound

Mind the Gap: Structure-Aware Consistency in Preference Learning cites this paper.

Mind the Gap: Structure-Aware Consistency in Preference Learning Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.589799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-07T06:41:18.311986Z digest=sha256:4f90306ca3dff55982858692fad016da29e5c2a70131ae4824719539c1f0851c

Observation 1da2f02c-27f8-440b-8fa5-c2ad82146aba · inbound

Internal Pluralism and the Limits of Pairwise Comparisons cites this paper.

Internal Pluralism and the Limits of Pairwise Comparisons Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

Reference 133

Resolution
unresolved
no resolver link, observed 2026-07-12T07:49:57.204875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T07:49:57.204875Z digest=sha256:1d07f0837e4d8e85925c63afe0dea073526c68421a896e23ae4c9f068be054e7

Observation 0da2bd54-9479-4abd-a04d-c364a1e6eb54 · inbound

Internal Pluralism and the Limits of Pairwise Comparisons cites this paper.

Internal Pluralism and the Limits of Pairwise Comparisons Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

Reference 222

Resolution
unresolved
no resolver link, observed 2026-08-02T09:03:15.212226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:03:15.212226Z digest=sha256:36e196c931f3f1daaa14e1c1eb39be09b17823be5c58aa40f31cc151917aa506