Pith. sign in

REVIEW 4 major objections 4 minor 1 cited by

AI Alignment at Your Discretion

T0 review · 4 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read The paper claims that AI alignment hides an excessive, largely unexamined form of discretion: annotators and models, not principles, decide which outputs are 'better' or 'safer', and those choices are often arbitrary.

desk verdict A genuinely useful formalization of discretion in alignment, but the empirical headline numbers rest on a single-model oracle that also sits in the audited set. read the letter →

arxiv 2502.10441 v1 pith:A4ZUERIF submitted 2025-02-10 cs.AI cs.CYcs.LG

classification cs.AIcs.CYcs.LG
keywords AIalignmentsafetydiscretionhumanfeedbackRLHFprincipleprioritizationrewardmodelsjudicial
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

AI alignment is built on pairwise choices: annotators and models decide which response is 'better' or 'safer.' The paper argues that the latitude granted by those choices, which it calls alignment discretion, is excessive, mostly unexamined, and frequently arbitrary, since abstract principles routinely conflict or give no guidance in the cases that matter. To make the phenomenon measurable, the paper defines when discretion is required (principle conflict or indifference) and how it is exercised (whether it contradicts a principle consensus, and which principles win when they clash). Across two standard safety-preference datasets, human annotators disagreed with a unanimous verdict of all principles 28.9% of the time on one dataset and 14–20% on the other, while an RLHF-fine-tuned model's principle ranking diverged from human annotators' by up to 71.2%. If the claim holds, feedback-based alignment encodes unexamined value judgments and should not be treated as a transparent grounding for safety.

What carries the argument

The central object is alignment discretion, operationalized through four linked measurements built on ternary preference functions. A preference function returns +1, −1, or 0 for each response pair; principle-specific preference functions use an assumed oracle to score how well each response adheres to one principle at a time. Given all principle votes on a pair, the taxonomy classifies the pair as principle consensus (all non-indifferent principles agree), principle conflict (principles disagree), or principle indifference (all abstain). Discretion arbitrariness is the frequency with which an annotator picks the response opposed to a consensus; principle supremacy is the empirical probability that one principle wins over another when they clash; principle priority fits an ELO-style logistic model to those pairwise win frequencies to produce a single ranking per annotator; and discretion discrepancy is the normalized Kendall-tau distance between two annotators' rankings. This machinery turns the legal notion of discretion—when it is required, how it is exercised, and whether it is consistent across decision-makers—into numbers that can be computed for any preference dataset.

What would settle it

Take a random subsample of the same two preference datasets, have at least two independent oracle systems or a panel of human judges label which response better adheres to each of the 21 principles, and recompute discretion arbitrariness and discretion discrepancy. If the human 28.9% arbitrariness rate or the priority rankings move by more than the reported bootstrap standard errors, the paper's diagnosis is an artifact of the chosen oracle; if they are stable, the existence of large, arbitrary discretion is confirmed.

Watch

Extended reading notes

Core claim

On the authors' own terms, the central discovery is that alignment discretion is both necessary and currently out of control: annotators must exercise judgment precisely because principles conflict or are indecisive, yet the field has no systematic record of how that judgment is used. The paper formalizes the three situations principles can be in—consensus, conflict, and indifference—and then measures, per annotator, how often they contradict a consensus (discretion arbitrariness), which principles win when two conflict (principle supremacy), the one-dimensional priority ordering implied by those wins, and how far any two annotators' orderings are apart (discretion discrepancy). Human labels in the first dataset contradicted a unanimous principle verdict 28.9% of the time, and an RLHF-fine-tuned model diverged from human principle priorities by 71.2%; reward models fine-tuned on the same preference data stayed within roughly 15–20% ranking discrepancy, while off-the-shelf models sat between 16% and 53%. The authors conclude that there is currently an excessive amount of discretion in the hands of model developers and annotators, that principles alone underdetermine aligned behavior, and that algorithms develop their own forms of discretion rather than inheriting human discretion. They also note that the oracle model they use for principle judgments shows near-zero arbitrariness, a self-confirmation effect they flag rather than count as evidence of ideal alignment.

Load-bearing premise

The paper's measurements all rest on a single zero-shot LLM as the oracle that decides, for each principle, which response adheres to it better; if that oracle's principle-wise judgments are biased or noisy, every downstream number—consensus rates, arbitrariness, supremacy, priorities, and discrepancies—inherits the error.

Editorial extensions

If this is right

  • Human preference datasets already encode implicit hierarchies of principles, so reusing or fine-tuning on a dataset means inheriting those latent priorities along with the preference labels.
  • Reward models can partly learn human discretion, staying within roughly 15–20% ranking discrepancy after fine-tuning, but translating that discretion into an RLHF-tuned policy fails, with discrepancies rising to roughly 40–71%; transferring discretion from a reward model to an LLM is an open problem.
  • Because roughly 80–85% of response pairs in both datasets are consensus or indifference, most of what customizes an aligned model happens in the 15–20% of conflicted cases where principles underdetermine the answer.
  • Off-the-shelf models do not mirror human annotators' principle priorities, with discrepancies between 16% and 53%, so using them as de facto arbiters of what is 'better' shifts alignment away from the humans whose preferences the datasets record.
  • Alignment frameworks that declare a set of principles without documenting how conflicts are resolved will keep producing systems whose behavior is shaped by unrecorded, unreviewed discretion.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the paper's measurements are sound, the natural next product is a standard 'discretion card' for preference datasets and aligned models, reporting arbitrariness, conflict rate, and priority rankings alongside behavior-benchmark scores, so discretionary latitude becomes auditable instead of incidental.
  • A sensitivity test the authors did not run: replace the single principle-judging oracle with several independent judges, including human panels and different LLMs, on a random subsample; if arbitrariness and discrepancy rates shift with the oracle, part of what the paper labels 'discretion' is actually evaluator noise, and the field needs a disambiguation protocol.
  • The legal analogy suggests a mechanism the paper does not develop: appellate review. A practical implementation would record an annotator's principle-supremacy profile at annotation time and flag decisions that deviate from that annotator's own prior profile, much as courts check whether a decision departs from precedent.
  • One implicit consequence for pluralistic alignment is that if different communities genuinely rank principles differently, measured 'discrepancy' is not always a defect; the metric could double as a diagnostic for whose values a model is aligned to, turning a limitation into a governance tool.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces the concept of 'alignment discretion' to describe the latitude that annotators, human or algorithmic, have in deciding which model outputs are 'better' or 'safer' during preference-based AI alignment. Drawing on legal theory, the authors formalize principle-specific preferences and define several metrics: principle consensus/conflict/indifference, discretion arbitrariness, principle supremacy, ELO-style principle priority weights, and discretion discrepancy. They apply these metrics to the HH-RLHF and PKU-SafeRLHF datasets, using GPT-4o as an oracle for 21 principles from Collective Constitutional AI, and compare human, reward-model, and LLM annotators. The paper reports, among other results, 28.9% human discretion arbitrariness on HH, reward-model discrepancies around 14-20%, and RLHF-tuned LLM discrepancies up to 71.2%, and concludes that current alignment processes allow excessive and unexamined discretion.

Significance. The formal framework is a useful step: the definitions are crisp, largely principle-agnostic, and the paper carefully reports bootstrap standard errors, controls for positional bias by swapping response order, and uses separate train/test splits. The legal analogy and the proposed metrics could provide a common vocabulary for discussing annotator latitude and principle prioritization in alignment research. However, the empirical magnitudes that carry the paper's central claims are computed through a single GPT-4o oracle that is also one of the audited annotators, and the paper's own Sec. 8 acknowledges the resulting feedback loop. For that reason, the headline numbers should be read as conditional on the oracle's interpretation of the 21 selected principles, not as an independent measure of arbitrariness or of failure to transfer human discretion. The framework itself survives this concern, but the empirical support for the broad conclusions needs substantial reworking.

major comments (4)
  1. [Sec. 6.1 / Def. 5 / Sec. B.3 / Tab. 1] The load-bearing empirical quantities are not independent of the oracle. Def. 5 sets Prefc = Preforacle for every principle, and Sec. B.3 instantiates the oracle with GPT-4o; all consensus/conflict/indifference classifications (Fig. 3), arbitrariness rates (Tab. 1), supremacies (Fig. 4), priorities (Fig. 5), and discrepancies (Tab. 2) are computed from GPT-4o's judgments. Tab. 1 accordingly reports GPT-4o's arbitrariness as 0.65% on HH, which the paper itself attributes to the oracle/annotator overlap. This is not merely a residual limitation acknowledged in Sec. 8; it changes the interpretation of every headline magnitude. A human label that disagrees with the GPT-4o-based consensus is counted as 'arbitrary' even when it reflects a defensible interpretation of an abstract principle, and the paper concedes in Sec. 8 that even 'reject cruelty' requires discretion. I therefore cannot read 28.9% (HH human) or 71.2% (Llama-3 fine-tuned) as estimates of excessive discretion in alignment; they are estimates of disagreement with one model's interpretation of 21 fixed principles. The framework survives, but the empirical claims need to be reframed as conditional on the oracle, or better, re-run with independent oracles and/or human principle-specific labels; at minimum, the sensitivity of Fig. 3 and Tab. 1 to the choice of oracle should be reported.
  2. [Sec. 6.2 / Fig. 4 / Def. 8] The analysis treats the entire HH preference set as if it were produced by a single 'human annotator', even though HH labels come from different crowdworkers per item and the paper's own Sec. 2 cites Anthropic's reliance on crowdworker diversity. For principle supremacy and priority, the metric pools all these labels into one Bernoulli estimate per principle pair. This conflates inter-annotator pluralism with the discretion of a single decision-maker, which is exactly the distinction the paper needs to make to support claims about 'annotators' having excessive discretion. Reporting annotator-level analyses, or at least clarifying that the human baseline is the aggregated dataset preference, is necessary before the claim that human annotators frequently use their power of discretion arbitrarily (Sec. 7) is supported.
  3. [Sec. 6.1 / Def. 2-3 / Tab. 2] The claim that RLHF fails to transfer human discretion to LLMs is based on an apples-to-oranges comparison. Reward-model preferences are defined as the sign of a scalar reward difference (Def. 2), while LLM preferences are elicited through a textual template (Def. 3). Tab. 2 shows fine-tuned reward models with DD around 14-20% but the corresponding RLHF-tuned LLMs at 40-70%; before interpreting this gap as a fundamental limitation of RLHF, the authors need to control for the evaluation modality, for example by scoring the fine-tuned policy's outputs with the reward model used in training, or by eliciting reward-model preferences with the same textual template. With only two base models, one reward model per dataset, and no random-seed variation, the claim that translating human discretion from reward models to LLMs is an open problem (Sec. 6.2 and abstract) is stronger than the evidence.
  4. [Def. 9 / Fig. 4 / Tab. 2] Several priority estimates rest on very small conflict counts. In Fig. 4, many cells report conflict counts below 10, and the priority weights in Eq. (13) are fitted to such sparse supremacies. The bootstrap standard errors in Tab. 2 appear to capture resampling over dataset items only; they do not propagate oracle-instance variability or the choice of principle set, both of which are substantial because a different oracle can reclassify a pair as consensus versus conflict. Reporting the stability of w* and DD across oracle models and principle subsets would make the quantitative comparisons in Tab. 2 interpretable.
minor comments (4)
  1. [Sec. 3] The phrase 'human and algorihmic discretion' contains a typo ('algorihmic' should be 'algorithmic').
  2. [Fig. 2 caption] The caption refers to 'Def 5.1a', 'Def 5.1b', and 'Def 5.1c', which do not match any numbered definition; these should likely be Def. 6a, 6b, and 6c.
  3. [Figs. 4 and 14] Some cells display non-zero percentages with '(0)' conflict counts or otherwise inconsistent counts (e.g., '40% (0)' and '80% (0)'); these should be reconciled with the stated totals or the notation should be explained.
  4. [Appendix B.5] The RLHF training section mentions hyperparameter sweeps but does not report the chosen values for learning rate, batch size, KL coefficient, or number of PPO epochs; concrete configurations are needed for reproducibility, especially because the paper does not provide a code or data release link.

Circularity Check

1 steps flagged · score 6.0 of 10

GPT-4o's near-zero arbitrariness is a same-oracle artifact; the formal discretion framework is otherwise self-contained, but the headline empirical magnitudes inherit this feedback loop.

  1. self definitional [Sec. 4.3 (Def. 5), Sec. B.3, Sec. 6.2 (Tab. 1)]
    "Assuming the availability of an oracle to judge principle-specific preferences ≻c, we denote principle-specific preference functions Prefc ... Prefc(y1 ≻ y0 | x) ≜ Preforacle(y1 ≻c y0 | x). ... We use GPT-4o as an oracle ... Remarkably, GPT-4o’s arbitrariness is very low ( < 1%) ... This can be explained by noting that GPT-4o is also the model we use as the oracle for principle preferences; it is thus heavily biased towards agreeing with the consensus of its own preferences."

    The consensus standard in Def. 6 is built from Prefc, and in the experiments Prefc is instantiated as GPT-4o's own principle judgments (Sec. B.3). The audited annotator GPT-4o (Def. 3) is the same model. Therefore the reported 0.65% arbitrariness for GPT-4o does not measure disagreement with an independent principle ground truth; it measures self-consistency between GPT-4o's generic preference and GPT-4o's principle-specific judgments. The low value is an artifact of the oracle/annotator identity, not evidence that GPT-4o exercises little discretion. The paper explicitly acknowledges this bias but still presents the number as a measured arbitrariness rate.

full rationale

The formal definitions of consensus, conflict, indifference, arbitrariness, supremacy, priority, and discrepancy are principle-agnostic and internally consistent. No self-citation chain is load-bearing, and the framework itself does not reduce to its inputs. The circularity is confined to the empirical instantiation: the oracle that defines the principle-specific preferences (Def. 5) is the same GPT-4o model that is then audited as an annotator. This makes GPT-4o's near-zero arbitrariness in Tab. 1 a same-source feedback loop rather than an independent measurement. The paper's own Limitations section warns that this 'may create problematic feedback loops that prioritize mirroring its perspectives rather than intended human values,' and that even applying a single principle like 'reject cruelty' can require discretion. Because the headline human arbitrariness and discrepancy figures are all computed against GPT-4o's principle judgments, their magnitudes are oracle-dependent; however, they are not logically forced by the definitions, and the core proposal of measuring discretion remains viable with an improved oracle. Hence the central claim has independent content, but the empirical headline is partially circular.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The paper's central quantities are measurements derived from GPT-4o's judgments and from fitted ELO weights. The most important implicit assumptions are the existence of a reliable oracle and the choice of the principle set; both are acknowledged in the limitations. No new physical or ontological entities are introduced.

free parameters (1)
  • ELO-style principle priority weights w*_c(a) = Not tabulated; shown in Figs. 5 and 12
    Computed by maximum likelihood in Def. 9 (Eq. 13) to reduce pairwise supremacy frequencies to a one-dimensional priority scale per annotator. These fitted weights are the basis of the discretion discrepancy metric.
assumptions (5)
  • domain assumption An oracle exists that can perfectly judge principle-specific preferences (Def. 5).
    The classification into consensus, conflict, and indifference, and all derived metrics, assume this oracle. In experiments GPT-4o instantiates it; Sec. 8 acknowledges the oracle may itself require discretion.
  • domain assumption The 21 Collective Constitutional AI seed statements are an adequate principle set for auditing HH-RLHF and PKU-SafeRLHF.
    Borrowed from prior work [48] and interpreted broadly (e.g., 'support democracy' means avoiding subversion of government, Sec. B.1). No complete principle list is given by either dataset.
  • domain assumption Pairwise principle supremacy follows the logistic model PSc>c'(a) ≈ σ(w_c - w_c').
    Def. 9 assumes this ELO-like form to compute one-dimensional priorities; standard in rating systems, but an untested assumption about the data.
  • standard math Human preferences in RLHF follow the Bradley-Terry-Luce model (Eq. 1).
    Used to justify reward modeling; not needed for the discretion metrics but needed for the RLHF training part of the experiments.
  • domain assumption Kendall tau distance is an appropriate measure of discrepancy between priority rankings (Def. 10).
    The authors motivate it in Appendix C as a tie-aware ranking distance; other choices would change the reported discrepancy values.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI Alignment at Your Discretion." pith.science (2026). https://pith.science/paper/A4ZUERIF

@misc{pith2026250210441,
  author       = {Pith},
  title        = {Pith review of: AI Alignment at Your Discretion},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/A4ZUERIF}},
  note         = {Machine review of arXiv:2502.10441}
}
read the original abstract

In AI alignment, extensive latitude must be granted to annotators, either human or algorithmic, to judge which model outputs are `better' or `safer.' We refer to this latitude as alignment discretion. Such discretion remains largely unexamined, posing two risks: (i) annotators may use their power of discretion arbitrarily, and (ii) models may fail to mimic this discretion. To study this phenomenon, we draw on legal concepts of discretion that structure how decision-making authority is conferred and exercised, particularly in cases where principles conflict or their application is unclear or irrelevant. Extended to AI alignment, discretion is required when alignment principles and rules are (inevitably) conflicting or indecisive. We present a set of metrics to systematically analyze when and how discretion in AI alignment is exercised, such that both risks (i) and (ii) can be observed. Moreover, we distinguish between human and algorithmic discretion and analyze the discrepancy between them. By measuring both human and algorithmic discretion over safety alignment datasets, we reveal layers of discretion in the alignment process that were previously unaccounted for. Furthermore, we demonstrate how algorithms trained on these datasets develop their own forms of discretion in interpreting and applying these principles, which challenges the purpose of having any principles at all. Our paper presents the first step towards formalizing this core gap in current alignment processes, and we call on the community to further scrutinize and control alignment discretion.

Figures

Figures reproduced from arXiv: 2502.10441 by the authors.

Figure 1
Figure 1. Illustration of how different prioritizations of principles affect which AI model responses are preferred, [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Illustration of the three principle agreement cases in Def. 6. For each prompt, two candidate responses [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Principle agreement frequency (%) according to the three cases distinguished in Def. 6. [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Principle supremacy matrix for the human annotator of the HH-RLHF dataset. The [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 6
Figure 6. Figure 6: Prompt template for oracle. This prompt template was used to obtain the oracle’s principle-specific prefer [PITH_FULL_IMAGE:figures/full_fig_p024_6.png]
Figure 7
Figure 7. Figure 7: Screenshot of the most downloaded models in the Hugging Face platform that were trained in the HH [PITH_FULL_IMAGE:figures/full_fig_p025_7.png]
Figure 8
Figure 8. Figure 8: Screenshot of the most downloaded models in the Hugging Face platform that were trained in the PKU [PITH_FULL_IMAGE:figures/full_fig_p026_8.png]
Figure 9
Figure 9. Figure 9: Prompt template for LLM preferences. This prompt template was used to obtain an LLM’s preferences. [PITH_FULL_IMAGE:figures/full_fig_p026_9.png]
Figure 10
Figure 10. Figure 10: Principle-specific preferences (Def. 5) averaged over the Anthropic HH-RLHF dataset. The proportion [PITH_FULL_IMAGE:figures/full_fig_p031_10.png]
Figure 11
Figure 11. Figure 11: Principle-specific preferences (Def. 5) averaged over the PKU-SafeRLHF dataset. The proportion of [PITH_FULL_IMAGE:figures/full_fig_p032_11.png]
Figure 13
Figure 13. Figure 13: Comparison of the ranking of principles based on their priority weights across the HH-RLHF and PKU [PITH_FULL_IMAGE:figures/full_fig_p034_13.png]
Figure 14
Figure 14. Figure 14: Principle supremacy matrix for the human annotator of the PKU-SafeRLHF dataset. The [PITH_FULL_IMAGE:figures/full_fig_p035_14.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Statutory Construction and Interpretation for Artificial Intelligence

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Prompt-based legal canons and iterative rule refinement reduce disagreement among LLM judges about whether a response complies with natural-language rules.

Reference graph

Works this paper leans on

118 extracted references · 48 canonical work pages · cited by 1 Pith paper

  1. [1]

    Public constitutional ai

    Gilad Abiri. Public constitutional ai. Forthcoming in Georgia Law Review, Volume 59, 2024

  2. [2]

    Claude’s constitution

    Anthropic. Claude’s constitution. https://www.anthropic.com/news/claudes-constitution, 2024. Accessed: 2025-01-03

  3. [3]

    Model card addendum: Claude 3.5 haiku and upgraded claude 3.5 sonnet, 2024

    Anthropic. Model card addendum: Claude 3.5 haiku and upgraded claude 3.5 sonnet, 2024. https://www. anthropic.com/model-cards/claude-3.5

  4. [4]

    Which humans? PsyArXiv, 2023

    Mohammad Atari, Mona J Xue, Peter S Park, Dami ´an Blasi, and Joseph Henrich. Which humans? PsyArXiv, 2023

  5. [5]

    Training a helpful and harmless assistant with reinforcement learning from human feedback

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022. 15 AI Alignment at Your Discretion

  6. [6]

    Constitutional ai: Harmlessness from ai feedback

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022

  7. [7]

    Judicial Discretion

    Aharon Barak. Judicial Discretion. Yale University Press, 1989

  8. [8]

    The judge in a democracy

    Aharon Barak. The judge in a democracy. Princeton University Press, 2009

Show all 118 references
  1. [9]

    Anthropomorphism and ai hype

    Nicholas Barrow. Anthropomorphism and ai hype. AI and Ethics, pages 1–5, 2024

  2. [10]

    ‘two concepts of liberty’

    Isaiah Berlin. ‘two concepts of liberty’. In Reading Political Philosophy, pages 231–237. Routledge, 2014

  3. [11]

    From ethics washing to ethics bashing: a view on tech ethics from within moral philosophy

    Elettra Bietti. From ethics washing to ethics bashing: a view on tech ethics from within moral philosophy. In Proceedings of the 2020 conference on fairness, accountability, and transparency, pages 210–219, 2020

  4. [12]

    Stereotyping norwegian salmon: An inventory of pitfalls in fairness benchmark datasets

    Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. Stereotyping norwegian salmon: An inventory of pitfalls in fairness benchmark datasets. In Proceedings of the 59th Annual Meet- ing of the Association for Computational Linguistics and the 11th ...

  5. [13]

    Rank analysis of incomplete block designs: I

    Ralph Allan Bradley and Milton E Terry. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345, 1952

  6. [14]

    Toward a perspectivist turn in ground truthing for predictive computing

    Federico Cabitza, Andrea Campagner, and Valerio Basile. Toward a perspectivist turn in ground truthing for predictive computing. Proceedings of the AAAI Conference on Artificial Intelligence , 37(6):6860–6868, June 2023

  7. [15]

    Alignment as jurisprudence

    Nicholas Caputo. Alignment as jurisprudence. Yale Journal of Law and Technology (forthcoming), 2024

  8. [16]

    Open problems and fundamental limita- tions of reinforcement learning from human feedback

    Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, J ´er´emy Scheurer, Javier Rando Ramirez, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al. Open problems and fundamental limita- tions of reinforcement learning from human feedback. Transacti...

  9. [17]

    Chatbot arena: An open platform for evaluating llms by human preference

    Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E Gonzalez, et al. Chatbot arena: An open platform for evaluating llms by human preference. arXiv preprint arXiv:2403.04132, 2024

  10. [18]

    Deep reinforcement learning from human preferences

    Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017

  11. [19]

    A coefficient of agreement for nominal scales

    Jacob Cohen. A coefficient of agreement for nominal scales. Educational and Psychological Measurement , 20(1):37–46, 1960

  12. [20]

    Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit

    Jacob Cohen. Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit. Psychological Bulletin, 70(4):213–220, 1968

  13. [21]

    Position: Social choice should guide ai alignment in dealing with diverse human feedback

    Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H Holliday, Bob M Jacobs, Nathan Lambert, Milan Moss´e, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, et al. Position: Social choice should guide ai alignment in dealing with diverse human feedback. In Forty-first Inte...

  14. [22]

    Ultrafeedback: Boosting language models with scaled ai feedback

    Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Bingxiang He, Wei Zhu, Yuan Ni, Guotong Xie, Ruobing Xie, Yankai Lin, et al. Ultrafeedback: Boosting language models with scaled ai feedback. In Forty-first International Conference on Machine Learning, 2024

  15. [23]

    On extending the bradley-terry model to accommodate ties in paired comparison experi- ments

    Roger R Davidson. On extending the bradley-terry model to accommodate ties in paired comparison experi- ments. Journal of the American Statistical Association, 65(329):317–328, 1970

  16. [24]

    A bibliography on the method of paired comparisons

    Roger R Davidson and Peter H Farquhar. A bibliography on the method of paired comparisons. Biometrics, pages 241–252, 1976

  17. [25]

    ‘affordances’ for machine learning

    Jenny L Davis. ‘affordances’ for machine learning. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, pages 324–332, 2023

  18. [26]

    Discretionary Justice: A Preliminary Inquiry

    Kenneth Culp Davis. Discretionary Justice: A Preliminary Inquiry. Lousiana State University Press, 1969

  19. [27]

    Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

    DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei L...

  20. [28]

    RLHF workflow: From reward modeling to online RLHF

    Hanze Dong, Wei Xiong, Bo Pang, Haoxiang Wang, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caim- ing Xiong, and Tong Zhang. RLHF workflow: From reward modeling to online RLHF. Transactions on Machine Learning Research, 2024

  21. [29]

    Steerlm: Attribute conditioned sft as an (user-steerable) alternative to rlhf

    Yi Dong, Zhilin Wang, Makesh Narsimhan Sreedhar, Xianchao Wu, and Oleksii Kuchaiev. Steerlm: Attribute conditioned sft as an (user-steerable) alternative to rlhf. arXiv preprint arXiv:2310.05344, 2023

  22. [30]

    Alpacafarm: A simulation framework for methods that learn from human feedback

    Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto. Alpacafarm: A simulation framework for methods that learn from human feedback. Advances in Neural Information Processing Systems, 36, 2024

  23. [31]

    Law’s empire

    Ronald Dworkin. Law’s empire. Harvard University Press, 1986

  24. [32]

    Taking rights seriously

    Ronald Dworkin. Taking rights seriously. A&C Black, 2013

  25. [33]

    Crowdworksheets: Accounting for individual and collective identities underlying crowdsourced dataset annotation

    Mark D ´ıaz, Ian Kivlichan, Rachel Rosen, Dylan Baker, Razvan Amironesei, Vinodkumar Prabhakaran, and Emily Denton. Crowdworksheets: Accounting for individual and collective identities underlying crowdsourced dataset annotation. In 2022 ACM Conference on Fairness, Accountabili...

  26. [34]

    Elo-MMR: A rating system for massive multiplayer competitions

    Aram Ebtekar and Paul Liu. Elo-MMR: A rating system for massive multiplayer competitions. In Proceedings of the Web Conference 2021, pages 1772–1784, 2021

  27. [35]

    The rating of chessplayers: Past and present

    Arpad Emrick Elo. The rating of chessplayers: Past and present. Batsford Chess Books, 1978

  28. [36]

    GPT-4 Technical Report, March 2024

    OpenAI et al. GPT-4 Technical Report, March 2024

  29. [37]

    Kto: Model alignment as prospect theoretic optimization

    Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela. Kto: Model alignment as prospect theoretic optimization. arXiv preprint arXiv:2402.01306, 2024

  30. [38]

    Generalized bradley-terry models for score estimation from paired comparisons

    Julien Fageot, Sadegh Farhadkhani, Lˆe-Nguyˆen Hoang, and Oscar Villemaud. Generalized bradley-terry models for score estimation from paired comparisons. InProceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 20379–20386, 2024

  31. [39]

    Moral machine or tyranny of the majority? InProceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 5974–5982, 2023

    Michael Feffer, Hoda Heidari, and Zachary C Lipton. Moral machine or tyranny of the majority? InProceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 5974–5982, 2023

  32. [40]

    Modular pluralism: Pluralistic alignment via multi-llm collaboration

    Shangbin Feng, Taylor Sorensen, Yuhan Liu, Jillian Fisher, Chan Young Park, Yejin Choi, and Yulia Tsvetkov. Modular pluralism: Pluralistic alignment via multi-llm collaboration. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 41...

  33. [41]

    Inverse constitu- tional ai: Compressing preferences into principles

    Arduin Findeis, Timo Kaufmann, Eyke H ¨ullermeier, Samuel Albanie, and Robert Mullins. Inverse constitu- tional ai: Compressing preferences into principles. arXiv preprint arXiv:2406.06560, 2024

  34. [42]

    Measuring nominal scale agreement among many raters

    Joseph Fleiss. Measuring nominal scale agreement among many raters. Psychological Bulletin, 76:378–, 11 1971

  35. [43]

    Stuart Geiger, Kevin Yu, Yanlai Yang, Mindy Dai, Jie Qiu, Rebekah Tang, and Jenny Huang

    R. Stuart Geiger, Kevin Yu, Yanlai Yang, Mindy Dai, Jie Qiu, Rebekah Tang, and Jenny Huang. Garbage in, garbage out?: do machine learning application papers in social computing report where human-labeled training data comes from? In Proceedings of the 2020 Conference on Fairne...

  36. [44]

    Chatgpt outperforms crowd workers for text-annotation tasks

    Fabrizio Gilardi, Meysam Alizadeh, and Ma ¨el Kubli. Chatgpt outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences, 120(30), July 2023

  37. [45]

    The many routes to the ubiquitous bradley-terry model

    Ian Hamilton, Nick Tawn, and David Firth. The many routes to the ubiquitous bradley-terry model. arXiv preprint arXiv:2312.13619, 2023

  38. [46]

    The concept of law

    Herbert Lionel Adolphus Hart and Leslie Green. The concept of law. Oxford University Press, 2012

  39. [47]

    Hourly wages in crowdworking: A meta-analysis

    Lars Hornuf and Daniel Vrankar. Hourly wages in crowdworking: A meta-analysis. Business & Information Systems Engineering, 64(5):553–573, August 2022

  40. [48]

    Collective constitutional ai: Aligning a language model with public input

    Saffron Huang, Divya Siddarth, Liane Lovitt, Thomas I Liao, Esin Durmus, Alex Tamkin, and Deep Ganguli. Collective constitutional ai: Aligning a language model with public input. In The 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 1395–1417, 2024

  41. [49]

    Generalized bradley-terry models and multi-class probability estimates

    Tzu-Kuo Huang, Ruby C Weng, Chih-Jen Lin, and Greg Ridgeway. Generalized bradley-terry models and multi-class probability estimates. Journal of Machine Learning Research, 7(1), 2006

  42. [50]

    Trustllm: Trustworthiness in large language models

    Yue Huang, Lichao Sun, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, et al. Trustllm: Trustworthiness in large language models. arXiv preprint arXiv:2401.05561, 2024

  43. [51]

    Towards accountability for machine learning datasets: Practices from software engineering and infrastructure

    Ben Hutchinson, Andrew Smart, Alex Hanna, Emily Denton, Christina Greer, Oddur Kjartansson, Parker Barnes, and Margaret Mitchell. Towards accountability for machine learning datasets: Practices from software engineering and infrastructure. In Proceedings of the 2021 ACM Confer...

  44. [52]

    Nanna Inie, Stefania Druga, Peter Zukerman, and Emily M Bender. From ”AI” to Probabilistic Automation: How Does Anthropomorphization of Technical Systems Descriptions Influence Trust? In The 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 2322–2347, 2024

  45. [53]

    Algorithmic Pluralism: A Structural Approach To Equal Opportunity

    Shomik Jain, Vinith Suriyakumar, Kathleen Creel, and Ashia Wilson. Algorithmic Pluralism: A Structural Approach To Equal Opportunity. InThe 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 197–206, Rio de Janeiro Brazil, June 2024. ACM

  46. [54]

    Pku-saferlhf: Towards multi-level safety alignment for llms with human preference

    Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen, Josef Dai, Boren Zheng, Tianyi Qiu, Boxun Li, and Yaodong Yang. Pku-saferlhf: Towards multi-level safety alignment for llms with human preference. arXiv preprint arXiv:2406.15513, 2024

  47. [55]

    Ai alignment: A comprehensive survey

    Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, et al. Ai alignment: A comprehensive survey. arXiv preprint arXiv:2310.19852, 2023

  48. [56]

    Mistral-7b- instruct-v0.2

    Albert Jiang, Alexandre Sablayrolles, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chap- lot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, L ´elio Renard Lavaud, Louis Ternon, Lucile Saulnier, Marie-Ann...

  49. [57]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L ´elio Renard Lavaud, Marie- Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thom...

  50. [58]

    The prism alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models

    Hannah Rose Kirk, Alexander Whitefield, Paul R ¨ottger, Andrew Michael Bean, Katerina Margatina, Rafael Mosquera, Juan Manuel Ciro, Max Bartolo, Adina Williams, He He, et al. The prism alignment dataset: What participatory, representative and individualised human feedback reve...

  51. [59]

    Pluralistic alignment over time

    Toryn Q Klassen, Parand A Alamdari, and Sheila A McIlraith. Pluralistic alignment over time. In Pluralistic Alignment Workshop at NeurIPS 2024, 2024

  52. [60]

    What are human values, and how do we align ai to them? arXiv preprint arXiv:2404.10636, 2024

    Oliver Klingefjord, Ryan Lowe, and Joe Edelman. What are human values, and how do we align ai to them? arXiv preprint arXiv:2404.10636, 2024

  53. [61]

    A guideline of selecting and reporting intraclass correlation coefficients for relia- bility research

    Terry K Koo and Mae Y Li. A guideline of selecting and reporting intraclass correlation coefficients for relia- bility research. Journal of chiropractic medicine, 15(2):155–163, 2016

  54. [62]

    Generalized distances between rankings

    Ravi Kumar and Sergei Vassilvitskii. Generalized distances between rankings. In Proceedings of the 19th international conference on World wide web, pages 571–580, 2010. 18 AI Alignment at Your Discretion

  55. [63]

    Judicial deliberations: a comparative analysis of transparency and legitimacy

    Mitchel de S-O Lasser et al. Judicial deliberations: a comparative analysis of transparency and legitimacy . Oxford University Press, 2009

  56. [64]

    Rlaif: Scaling reinforcement learning from human feedback with ai feedback

    Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Ren Lu, Thomas Mesnard, Johan Ferret, Colton Bishop, Ethan Hall, Victor Carbune, and Abhinav Rastogi. Rlaif: Scaling reinforcement learning from human feedback with ai feedback. arXiv preprint arXiv:2309.00267, 2023

  57. [65]

    Dissecting human and llm preferences

    Junlong Li, Fan Zhou, Shichao Sun, Yikai Zhang, Hai Zhao, and Pengfei Liu. Dissecting human and llm preferences. arXiv preprint arXiv:2402.11296, 2024

  58. [66]

    Decompose and aggregate: A step-by-step interpretable evaluation framework

    Minzhi Li, Zhengyuan Liu, Shumin Deng, Shafiq Joty, Nancy F Chen, and Min-Yen Kan. Decompose and aggregate: A step-by-step interpretable evaluation framework. arXiv preprint arXiv:2405.15329, 2024

  59. [67]

    From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline

    Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E Gonzalez, and Ion Stoica. From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline. arXiv preprint arXiv:2406.11939, 2024

  60. [68]

    Align- ing with human judgement: The role of pairwise preference in large language model evaluators

    Yinhong Liu, Han Zhou, Zhijiang Guo, Ehsan Shareghi, Ivan Vulic, Anna Korhonen, and Nigel Collier. Align- ing with human judgement: The role of pairwise preference in large language model evaluators. arXiv preprint arXiv:2403.16950, 2024

  61. [69]

    Individual choice behavior, volume 4

    R Duncan Luce. Individual choice behavior, volume 4. Wiley New York, 1959

  62. [70]

    Beyond probabilities: Unveiling the misalignment in evaluating large language models, 2024

    Chenyang Lyu, Minghao Wu, and Alham Fikri Aji. Beyond probabilities: Unveiling the misalignment in evaluating large language models, 2024

  63. [71]

    Kangaroo Courts and the Rule of Law: The Legacy of Modernism

    Desmond Manderson. Kangaroo Courts and the Rule of Law: The Legacy of Modernism. Routledge, London, July 2012

  64. [72]

    Gemma Team Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, L. Sifre, Morgane Rivi `ere, Mihir Kale, J Christopher Love, Pouya Dehghani Tafti, L’eonard Hussenot, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambr...

  65. [73]

    Dealing with disagreements: Look- ing beyond the majority vote in subjective annotations

    Aida Mostafazadeh Davani, Mark D ´ıaz, and Vinodkumar Prabhakaran. Dealing with disagreements: Look- ing beyond the majority vote in subjective annotations. Transactions of the Association for Computational Linguistics, 10:92–110, 2022

  66. [74]

    Rule based rewards for fine-grained LLM safety

    Tong Mu, Alec Helyar, Johannes Heidecke, Joshua Achiam, Andrea Vallone, Ian D Kivlichan, Molly Lin, Alex Beutel, John Schulman, and Lilian Weng. Rule based rewards for fine-grained LLM safety. InICML 2024 Next Generation of AI Safety Workshop, 2024

  67. [75]

    three laws of robotics

    Randall Munroe. xkcd #1613: “three laws of robotics”. https://xkcd.com/1613/, Dec 2015. Accessed: 2025-01-18

  68. [76]

    John J. Nay. Law informs code: A legal informatics approach to aligning artificial intelligence with humans. SSRN Working Paper, 2024

  69. [77]

    Value imprint: A technique for auditing the human values embedded in rlhf datasets

    Ike Obi, Rohan Pant, Srishti Shekhar Agrawal, Maham Ghazanfar, and Aaron Basiletti. Value imprint: A technique for auditing the human values embedded in rlhf datasets. arXiv preprint arXiv:2411.11937, 2024

  70. [78]

    International covenant on civil and political rights

    Office of the United Nations High Commissioner for Human Rights (OHCHR). International covenant on civil and political rights. https://www.ohchr.org/en/instruments-mechanisms/instruments/ international-covenant-civil-and-political-rights . Accessed: 2025-01-22

  71. [79]

    Gpt-4o system card, 2024

    OpenAI. Gpt-4o system card, 2024. 19 AI Alignment at Your Discretion

  72. [80]

    Training language models to follow instructions with human feedback

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:2773...

  73. [81]

    Llm evaluators recognize and favor their own generations

    Arjun Panickssery, Samuel R Bowman, and Shi Feng. Llm evaluators recognize and favor their own generations. arXiv preprint arXiv:2404.13076, 2024

  74. [82]

    Don‘t blame the annotator: Bias already starts in the annotation instructions

    Mihir Parmar, Swaroop Mishra, Mor Geva, and Chitta Baral. Don‘t blame the annotator: Bias already starts in the annotation instructions. In Andreas Vlachos and Isabelle Augenstein, editors, Proceedings of the 17th Conference of the European Chapter of the Association for Compu...

  75. [83]

    Comparing Bayesian models of annotation

    Silviu Paun, Bob Carpenter, Jon Chamberlain, Dirk Hovy, Udo Kruschwitz, and Massimo Poesio. Comparing Bayesian models of annotation. Transactions of the Association for Computational Linguistics , 6:571–585, 2018

  76. [84]

    Karl Pearson. Vii. mathematical contributions to the theory of evolution.—iii. regression, heredity, and pan- mixia. Philosophical Transactions of the Royal Society of London. Series A, containing papers of a mathemat- ical or physical character, (187):253–318, 1896

  77. [85]

    Red teaming language models with language models

    Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. Red teaming language models with language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, page...

  78. [86]

    Improving context-aware preference mod- eling for language models

    Silviu Pitis, Ziang Xiao, Nicolas Le Roux, and Alessandro Sordoni. Improving context-aware preference mod- eling for language models. arXiv preprint arXiv:2407.14916, 2024

  79. [87]

    The analysis of permutations

    Robin L Plackett. The analysis of permutations. Journal of the Royal Statistical Society Series C: Applied Statistics, 24(2):193–202, 1975

  80. [88]

    The “problem” of human label variation: On ground truth in data, modeling and evaluation

    Barbara Plank. The “problem” of human label variation: On ground truth in data, modeling and evaluation. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors,Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10671–10682, Abu D...

  81. [89]

    On releasing annotator-level labels and information in datasets

    Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Diaz. On releasing annotator-level labels and information in datasets. In Claire Bonial and Nianwen Xue, editors, Proceedings of the Joint 15th Linguistic Annotation Workshop (LAW) and 3rd Designing Meaning Representat...

  82. [90]

    Di- rect preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2024

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Di- rect preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2024

  83. [91]

    Constructing domain-specific evalua- tion sets for llm-as-a-judge

    Ravi Raju, Swayambhoo Jain, Bo Li, Jonathan Li, and Urmish Thakker. Constructing domain-specific evalua- tion sets for llm-as-a-judge. arXiv preprint arXiv:2408.08808, 2024

  84. [92]

    Ethical reasoning over moral alignment: A case and framework for in-context ethical policies in LLMs

    Abhinav Sukumar Rao, Aditi Khandelwal, Kumar Tanmay, Utkarsh Agarwal, and Monojit Choudhury. Ethical reasoning over moral alignment: A case and framework for in-context ethical policies in LLMs. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Findings of the Association...

  85. [93]

    The authority of law: essays on law and morality

    Joseph Raz. The authority of law: essays on law and morality. Oxford University Press, 2009

  86. [94]

    Why don‘t you do it right? analysing annotators’ disagreement in subjective tasks

    Marta Sandri, Elisa Leonardelli, Sara Tonelli, and Elisabetta Jezek. Why don‘t you do it right? analysing annotators’ disagreement in subjective tasks. In Andreas Vlachos and Isabelle Augenstein, editors,Proceedings of the 17th Conference of the European Chapter of the Associa...

  87. [95]

    Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. Annotators with attitudes: How annotator beliefs and identities bias toxic language detection. In Marine Carpuat, Marie- Catherine de Marneffe, and Ivan Vladimir Meza Ruiz, editors, Proce...

  88. [96]

    The European convention on human rights: a commentary

    William A Schabas. The European convention on human rights: a commentary. Oxford University Press, 2015

  89. [97]

    Evaluating the Moral Beliefs Encoded in LLMs

    Nino Scherrer, Claudia Shi, Amir Feder, and David Blei. Evaluating the Moral Beliefs Encoded in LLMs. Advances in Neural Information Processing Systems, 36:51778–51809, December 2023. 20 AI Alignment at Your Discretion

  90. [98]

    Towards bidirectional human-ai alignment: A systematic review for clarifications, framework, and future directions

    Hua Shen, Tiffany Knearem, Reshmi Ghosh, Kenan Alkiek, Kundan Krishna, Yachuan Liu, Ziqiao Ma, Savvas Petridis, Yi-Hao Peng, Li Qiwei, et al. Towards bidirectional human-ai alignment: A systematic review for clarifications, framework, and future directions. arXiv preprint arXi...

  91. [99]

    Hwang, Sydney Levine, Valentina Pyatkin, Peter West, Nouha Dziri, Ximing Lu, Kavel Rao, Chandra Bhagavatula, Maarten Sap, John Tasioulas, and Yejin Choi

    Taylor Sorensen, Liwei Jiang, Jena D. Hwang, Sydney Levine, Valentina Pyatkin, Peter West, Nouha Dziri, Ximing Lu, Kavel Rao, Chandra Bhagavatula, Maarten Sap, John Tasioulas, and Yejin Choi. Value Kaleido- scope: Engaging AI with Pluralistic Human Values, Rights, and Duties. ...

  92. [100]

    A roadmap to pluralistic alignment

    Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, et al. A roadmap to pluralistic alignment. arXiv preprint arXiv:2402.05070, 2024

  93. [101]

    The proof and measurement of association between two things

    Charles Spearman. The proof and measurement of association between two things. Appleton-Century-Crofts, 1961

  94. [102]

    Learning to summarize with human feedback

    Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul F Christiano. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021, 2020

  95. [103]

    Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback

    Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. arXiv preprint a...

  96. [104]

    Context-dependent preferences

    Amos Tversky and Itamar Simonson. Context-dependent preferences. Manage. Sci., 39(10):1179–1189, Octo- ber 1993

  97. [105]

    Trl: Transformer reinforcement learning

    Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, Shengyi Huang, Kashif Rasul, and Quentin Gallou ´edec. Trl: Transformer reinforcement learning. GitHub repository, 2020

  98. [106]

    Aligning language models with human preferences via a bayesian approach

    Jiashuo Wang, Haozhao Wang, Shichao Sun, and Wenjie Li. Aligning language models with human preferences via a bayesian approach. Advances in Neural Information Processing Systems, 36, 2024

  99. [107]

    Large language models are not fair evaluators

    Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926, 2023

  100. [108]

    Systematic evaluation of llm-as-a-judge in llm alignment tasks: Explainable metrics and diverse prompt templates

    Hui Wei, Shenghua He, Tian Xia, Andy Wong, Jingyang Lin, and Mei Han. Systematic evaluation of llm-as-a-judge in llm alignment tasks: Explainable metrics and diverse prompt templates. arXiv preprint arXiv:2408.13006, 2024

  101. [109]

    A survey of preference-based reinforcement learning methods

    Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes F ¨urnkranz. A survey of preference-based reinforcement learning methods. Journal of Machine Learning Research, 18(136):1–46, 2017

  102. [110]

    Style over substance: Evaluation biases for large language models

    Minghao Wu and Alham Fikri Aji. Style over substance: Evaluation biases for large language models. arXiv preprint arXiv:2307.03025, 2023

  103. [111]

    Eunice Yiu, Eliza Kosoy, and Alison Gopnik. Transmission versus truth, imitation versus innovation: What children can do that large language and language-and-vision models cannot (yet).Perspectives on Psychological Science, 19(5):874–883, 2024

  104. [112]

    Diverging preferences: When do annotators disagree and do models know? arXiv preprint arXiv:2410.14632, 2024

    Michael JQ Zhang, Zhilin Wang, Jena D Hwang, Yi Dong, Olivier Delalleau, Yejin Choi, Eunsol Choi, Xiang Ren, and Valentina Pyatkin. Diverging preferences: When do annotators disagree and do models know? arXiv preprint arXiv:2410.14632, 2024

  105. [113]

    Judging llm-as-a-judge with mt-bench and chatbot arena

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuo- han Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36:46595–46623, 2023

  106. [114]

    Fine-tuning language models from human preferences

    Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019

  107. [115]

    Can large language models transform computational social science? Computational Linguistics, 50(1):237–291, 2024

    Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. Can large language models transform computational social science? Computational Linguistics, 50(1):237–291, 2024. 21 AI Alignment at Your Discretion A Overview of the Supplementary Material In thi...

  108. [117]

    Without a standardized framework, categorizing and interpreting the diverse factors that annotators may consider becomes exceedingly difficult

    Lack of a Universal Framework : There is no universally agreed-upon set of principles governing human preferences [92]. Without a standardized framework, categorizing and interpreting the diverse factors that annotators may consider becomes exceedingly difficult. Moreover, pre...

  109. [118]

    A” and “B

    Unclear Annotator Guidelines: The guidelines provided to annotators may be ambiguous or lack sufficient detail, leading to inconsistent interpretations of instructions [33, 43]. This issue is further exacerbated by the diverse backgrounds of annotators, who bring varying cultu...

  110. [2022]

    Association for Computational Linguistics

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.